A data mining method, device and computer-readable storage medium
By performing feature extraction and graph data processing on the data set, filtering and calculating the purity of the data cluster, the problem of indistinguishable category diversity in the data cluster is solved, and the accuracy and efficiency of data mining are improved.
Patent Information
- Application Number
- CN201910801360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-12-15
AI Technical Summary
In the prior art, due to the large diversity and differences in the data cluster in data mining, it is difficult to distinguish between bad files and normal data. The hit rate of manual search and simple calculation distance methods is low, and the purity of the data cluster cannot be effectively guaranteed.
By extracting the data set to be processed, constructing the feature space, generating graph data, filtering the data clusters corresponding to the nodes, and calculating the purity in the cluster. When the purity in the cluster is lower than the threshold, the corresponding data is obtained, and a graph convolutional neural network is used for feature extraction and classification to improve the purity of the data cluster.
It improves the hit rate of bad gears in data mining, realizes the rapid, efficient and accurate mining of bad gears in data, and reduces the dependence on feature representation.
Smart Images

Figure CN110598065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to a data mining method, apparatus, and computer-readable storage medium. Background Art
[0002] In data mining scenarios, whether it is image data or text data, pure data is required. However, limited by the representation ability of the model, the data generated by classification, clustering, etc. often cannot ensure purity within the cluster due to having bad cases. In the prior art, the purity of the data cluster is judged by manually searching and simply calculating the distance between pairwise features within the data cluster.
[0003] In the process of researching and practicing the prior art, the inventors of the present invention found that manually searching consumes a large amount of human costs, and simply calculating the distance between pairwise features within the data cluster makes it difficult to distinguish bad cases from normal data due to the large class diversity and differences of the data cluster, resulting in a low hit rate of bad cases. Summary of the Invention
[0004] Embodiments of the present invention provide a data mining method, apparatus, and computer-readable storage medium, which can improve the hit rate of bad cases in data mining.
[0005] A data mining method includes:
[0006] Performing feature extraction on a dataset to be processed to construct a feature space;
[0007] Extracting node features in the feature space to generate graph data of the dataset to be processed, where the graph data includes at least one node;
[0008] Screening out the data cluster corresponding to the node in the graph data;
[0009] Calculating the data purity of the data cluster to obtain the intra-cluster purity of the data cluster;
[0010] When the intra-cluster purity is lower than a preset intra-cluster purity threshold, obtaining the data corresponding to the node in the dataset to be processed to obtain the mined data.
[0011] Correspondingly, an embodiment of the present invention provides a data mining apparatus, including:
[0012] An extraction unit for performing feature extraction on a dataset to be processed to construct a feature space;
[0013] A generation unit for extracting node features in the feature space to generate graph data of the dataset to be processed, where the graph data includes at least one node;
[0014] A screening unit, configured to screen out the data clusters corresponding to the nodes in the graph data;
[0015] A calculation unit, configured to calculate the data purity of the data clusters to obtain the intra-cluster purity of the data clusters;
[0016] An acquisition unit, configured to, when the intra-cluster purity is lower than a preset purity threshold, acquire the data corresponding to the nodes in the to-be-processed data set to obtain the mined data.
[0017] Optionally, in some embodiments, the calculation unit is specifically configured to extract features from the data clusters by using a trained graph recognition model to obtain the data information of the data clusters, classify the data in the data clusters according to the data information, and calculate the data purity of the data clusters according to the classification result to obtain the intra-cluster purity of the data clusters.
[0018] Optionally, in some embodiments, the calculation unit is specifically configured to, according to the classification result, obtain the quantity of each category of data and the total quantity of the data of the data clusters in the data information, screen out the data with the largest quantity among the quantities of each category of data as the target data, and calculate the ratio of the target data to the total quantity of the data of the data clusters to obtain the intra-cluster purity of the data clusters.
[0019] Optionally, in some embodiments, the calculation unit is specifically configured to collect multiple data set samples, where the data set samples include data clusters with labeled cluster class purities, use a preset graph recognition model to predict the cluster class purities of the data set samples to obtain predicted cluster class purities, and converge the preset graph recognition model according to the predicted cluster class purities and the labeled cluster class purities to obtain a trained graph recognition model.
[0020] Optionally, in some embodiments, the acquisition unit is specifically configured to, when the cluster class purity is lower than a preset intra-cluster purity threshold, determine the target nodes corresponding to the data clusters, screen out the graph data corresponding to the target nodes in the graph data of the to-be-processed data set, and acquire the data corresponding to the nodes in the to-be-processed data set according to the graph data corresponding to the nodes, and use the data as the data to be mined in the to-be-processed data set.
[0021] Optionally, in some embodiments, the screening unit is specifically configured to search for the adjacent nodes corresponding to the nodes in the graph data, cluster the nodes and the corresponding adjacent nodes in the graph data to obtain the clustering graph of the nodes, and screen out the data clusters corresponding to the nodes in the clustering graph.
[0022] Optionally, in some embodiments, the extraction unit is specifically configured to extract node features in the feature space, classify the node features, and generate graph data of the to-be-processed data set according to the classification result.
[0023] Optionally, in one embodiment, the extraction unit is specifically configured to extract node information from the node features of each category according to the classification result, construct a relationship tree based on the node information, and generate graph data of the to-be-processed data set based on the constructed relationship tree.
[0024] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores an application program, and the processor is configured to run the application program in the memory to implement the data mining method provided by the embodiment of the present invention.
[0025] In addition, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the data mining methods provided by the embodiment of the present invention.
[0026] In the embodiment of the present invention, feature extraction is performed on the to-be-processed data set to construct a feature space, and node features are extracted in the feature space to generate graph data of the to-be-processed data set. The graph data includes at least one node. In the graph data, a data cluster corresponding to the node is screened out, the data purity of the data cluster is calculated to obtain the intra-cluster purity of the data cluster. When the intra-cluster purity is lower than a preset purity threshold, the data corresponding to the node in the to-be-processed data set is obtained to obtain the mined data. Since this solution not only examines all the feature information within the data cluster, but also evaluates bad cases through the intra-cluster purity within the data cluster, and then conducts bad case mining, reducing the over-reliance on feature representation, it can more quickly, efficiently, and accurately mine bad cases (Bad case) in the data, thereby improving the hit rate of bad cases in the data. Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 is a schematic diagram of the scenario of the data mining method provided by the embodiment of the present invention;
[0029] Figure 2 is a schematic flowchart of the data mining method provided by the embodiment of the present invention;
[0030] Figure 3 It is a schematic structural diagram of graph data provided by an embodiment of the present invention;
[0031] Figure 4 It is a schematic flowchart of calculating the intra-cluster purity of a data cluster provided by an embodiment of the present invention;
[0032] Figure 5 It is another schematic flowchart of a data mining method provided by an embodiment of the present invention;
[0033] Figure 6 It is a schematic structural diagram of a data mining device provided by an embodiment of the present invention;
[0034] Figure 7 It is a schematic structural diagram of an extraction unit of a data mining device provided by an embodiment of the present invention;
[0035] Figure 8 It is a schematic structural diagram of a generation unit of a data mining device provided by an embodiment of the present invention;
[0036] Figure 9 It is a schematic structural diagram of a screening unit of a data mining device provided by an embodiment of the present invention;
[0037] Figure 10 It is a schematic structural diagram of a calculation unit of a data mining device provided by an embodiment of the present invention;
[0038] Figure 11 It is another schematic structural diagram of a data mining device provided by an embodiment of the present invention;
[0039] Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0041] An embodiment of the present invention provides a data mining method, device, and computer-readable storage medium. Among them, the data mining device can be integrated in an electronic device, and the electronic device can be a server or a terminal device, etc.
[0042] The so-called data mining can be a process of extracting implicit, previously unknown but potentially useful information and knowledge from a large amount of incomplete, noisy, fuzzy, and random data, and can also be a process of discovering and extracting bad cases in the data from a vast amount of data. Among them, a bad case can include multiple different categories of data within a data cluster. Since a data cluster can only accommodate one or one type of file, when there are multiple different categories of data within the cluster, it will cause data chaos at this time. Therefore, when processing data, it is necessary to dig out the bad cases in the data. In the embodiments of the present invention, it mainly refers to digging out bad cases from a vast amount of data.
[0043] For example, referring to Figure 1 , taking the case where the data mining device is integrated in an electronic device as an example, the electronic device extracts features from the dataset to be processed to construct a feature space. Then, node features are extracted in the feature space to generate graph data of the dataset to be processed. The graph data includes at least one node. Then, the data clusters corresponding to the nodes are screened out in the graph data, and the data purity of the data clusters is calculated to obtain the intra-cluster purity of the data clusters. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the nodes in the dataset to be processed is obtained to obtain the mined data.
[0044] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0045] This embodiment will be described from the perspective of the data mining device. The text label generation device can be specifically integrated in an electronic device, which can be a server or a terminal device, etc.; among them, the terminal can include devices such as a tablet computer, a notebook computer, and a personal computer (PC).
[0046] A data mining method includes: extracting features from the dataset to be processed to construct a feature space, then extracting node features in the feature space to generate graph data of the dataset to be processed. The graph data includes at least one node. Then, the data clusters corresponding to the nodes are screened out in the graph data, and the data purity of the data clusters is calculated to obtain the intra-cluster purity of the data clusters. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the nodes in the dataset to be processed is obtained to obtain the mined data.
[0047] As Figure 2 shown, the specific process of the data mining method is as follows:
[0048] 101. Extract features from the dataset to be processed to construct a feature space.
[0049] Among them, the feature space may include the space where all feature vectors exist. All features in the dataset to be processed are stored in this space in the form of feature vectors, including the relationship attributes between features. For example, the relationship between features is represented by nodes, and it also includes the attributes of the features themselves.
[0050] (1) Obtain the dataset to be processed;
[0051] For example, there are multiple methods to obtain the dataset to be processed. For instance, data can be obtained from the Internet, such as downloading or collecting, and then forming a dataset. It can also include the user uploading data to the server, and the data mining device obtains the data uploaded by the user from the server to form a dataset. The dataset can include one type of data or multiple types of data.
[0052] (2) Extract features from the dataset to be processed to construct a feature space;
[0053] For example, there are multiple methods to extract features from the dataset to be processed. For instance, a deep residual network can be used to extract features from the dataset to be processed, and extract the feature information of the data in the dataset to be processed, such as the structure of the data, the relationship between data and data, and / or the type of data and other feature information. The extracted feature information is arranged and stored in the form of feature vectors to construct a feature space. For example, using the relationship between features, the overall structure of the feature space is constructed, where the connection positions in the overall structure can be the intersections or nodes between features. The feature space is improved or enriched using the attributes of the features themselves to form a feature space that includes all the extracted feature information, and all the feature information is stored in the feature space.
[0054] 102. Extract node features in the feature space to generate graph data of the dataset to be processed, and the graph data includes at least one node.
[0055] Among them, the node features can be the feature information of the intersections or nodes formed by the mutual relationship between features, and the node features can include the information of one or more nodes. The graph data can be a type in the data structure, also known as a graph, which can include nodes and edges. A node can have two or more adjacent elements, and the connection between two nodes is called an edge.
[0056] For example, node features are extracted in the feature space and classified. There are various classification methods. For example, the node features can be classified by hierarchical clustering, or the K-nearest neighbor algorithm can be used to classify the node features to obtain different types of node features. For example, according to the different positions of the node features in the feature space, they can be divided into head node features, middle node features, and tail node features.
[0057] According to the classification results, corresponding node information is extracted from each type of node features. For example, one or more node information of the head is extracted from the head node features, one or more node information of the middle is extracted from the middle node features, and one or more node information of the tail is extracted from the tail node features. A relationship tree is constructed based on the extracted node information of each type. For example, a relationship tree can be constructed according to the mutual relationship between the extracted node information. For instance, the mutual relationship among one or more node information of the tail is obtained, the node information of the root node is judged based on the mutual relationship, the root of the relationship tree is constructed based on the node information of the root node, the nodes in the trunk above the root node in the relationship tree are sequentially found according to the obtained information of the root node, and the nodes on the branches corresponding to the nodes on the trunk are searched from the remaining node information according to the node information on the main path of the trunk. These nodes are connected to each other to form the branches and the trunk of the relationship tree. Then, according to the node information on the branches and the trunk, the information of the leaf nodes on the branches is obtained to complete the construction of the relationship tree.
[0058] Based on the constructed relationship tree, the graph data of the dataset to be processed is generated. For example, the feature attributes on the nodes can be filled or fused into the relationship tree so that each node in the relationship tree includes one or more data, or the data in the features can be mapped to each connection line in the relationship tree to form the edges of the graph data. The generated graph data structures and visualizes the data in the dataset to be processed. The mutual relationships such as the positional relationship and the structural relationship between the data in the dataset can be intuitively reflected from the graph data, and it can also include the attribute information of the data itself. Common graph data structures are as Figure 3 shown. Each node or vertex can be each data in the dataset, and the relationships between each data can be represented by the edge lines. If the relationship between the data is added to the edge lines, the edge lines in the graph data can also represent some data in the dataset.
[0059] Among them, the relationship tree, also known as the tree structure, can be a data structure with a "one-to-many" tree relationship between data elements, which is an important type of non-linear data structure. In the tree structure, elements are connected through nodes to form a tree structure. Among them, the tree structure is compared to a tree. The root node has no precursor node, and each of the remaining nodes has exactly one precursor node. The leaf node has no successor node, and the number of successor nodes of each of the remaining nodes can be one or more.
[0060] 103. Filter the data clusters corresponding to the nodes in the graph data.
[0061] Among them, the data cluster, also known as the cluster, can be the smallest storage management unit in data storage. For example, a file is usually stored in one or more clusters, but it must at least occupy one "cluster" alone. That is to say, two files cannot be stored in the same cluster. Simply put, files or data exist in the data clusters in the computer system, and one or a class of files are stored in the same data cluster.
[0062] For example, randomly select a node in the graph data, use the selected node as the target node, and use the selected target node to search for the adjacent nodes corresponding to the target node. There are various methods to search for adjacent nodes. For example, the K adjacent nodes of the target node can be found through cosine similarity. K can be any value. Among them, the adjacent nodes can include the nodes directly adjacent to the target node, and can also include the nodes whose distance from the target node in the graph data is within a preset distance threshold. For example, in the graph data, the target node and the remaining other nodes are converted into vectors in space, and the cosine value of the angle between the vectors converted from the remaining other nodes and the target vector converted from the target node is used as a measure or judgment of the cosine distance between the remaining other nodes and the target node. By judging the cosine distance between the remaining other nodes and the target node and the preset distance threshold, the remaining nodes with a cosine distance within the preset distance threshold are used as the adjacent nodes of the target node.
[0063] Hierarchical clustering is performed on the target node and its corresponding adjacent nodes in the graph data to obtain a clustering graph of the target node. Among them, the merging algorithm of hierarchical clustering can be to calculate the similarity between two types of data points, combine the two most similar data points among all data points, and iterate this process repeatedly. Simply put, the merging algorithm of hierarchical clustering determines the similarity between them by calculating the distance between the data points of each category and all data points. The smaller the distance, the higher the similarity. Then, the two data points or categories with the closest distance are combined to generate a clustering graph, thereby completing the classification of the data. For example, in graph data, each node can be regarded as one or more data, and the obtained target node and its corresponding adjacent nodes can include multiple or multiple types of data. By calculating the similarity between the multiple data of the target node and its corresponding adjacent nodes, the two most similar data among all data are combined, and this process is iterated repeatedly. Finally, the data in the target node and its corresponding adjacent nodes are divided into two or more categories, and a clustering graph is generated.
[0064] Filter the data cluster corresponding to the target node in the clustering graph. For example, in the clustering graph, according to the target node, a clustering sub-graph of the target node is generated, and the clustering sub-graph is used as the data cluster corresponding to the target node. Among them, the clustering sub-graph can be regarded as a sub-topological graph composed of one or more data that are closest or most similar to the target node in the clustering graph. It can be seen that the data cluster contains multiple data, and the multiple data can be one type or multiple types.
[0065] 104. Calculate the data purity of the data cluster to obtain the intra-cluster purity of the data cluster.
[0066] Among them, the data purity can include the ratio between the data volume of various or all types of data in the data cluster and the total data volume. The intra-cluster purity can include the ratio between the data volume of the type of data with the largest number in the data cluster and the total data volume.
[0067] For example, use the trained graph recognition model to extract features from the data cluster to obtain the data information of the data cluster. For example, the features of the data cluster can be extracted through a Graph Convolutional Network (GCN) to obtain the data information of the data cluster. Specifically, it can be as follows:
[0068] Each node in the data cluster sends its own feature information to its respective adjacent nodes after transformation. At this time, the adjacent nodes include the adjacent nodes directly connected by the side lines. Each node aggregates the feature information sent by its respective adjacent nodes for feature information fusion to obtain the data information in the data cluster. Among them, the data information includes the total number of all data in the data cluster, and also includes the attribute information of all data.
[0069] Among them, the selected target node v j is used as the central vertex input to the GCN model. The GCN model takes as input a clustering subgraph (data cluster) composed of a vertex set associated with the visual features of the target node v j , extracts features from this clustering subgraph (data cluster), and obtains the data information of this clustering subgraph (data cluster). The calculation formula is as follows:
[0070]
[0071] Among them, A(P i ) is a clustering subgraph composed of a vertex set associated with the visual features of the target node v j , and are respectively the data information of the target node and any node associated with it in the clustering subgraph. is a diagonal matrix, I is the identity matrix, F l (P i ) is the feature representation of the l-th layer, W l is the feature mapping learned by the l-th layer GCN model, and σ is the activation function. It should be noted here that the activation function can be selected as ReLU (an activation function).
[0072] After the GCN model obtains the data information within the data cluster through feature extraction, it classifies the data within the data cluster. For example, it can classify according to the attribute information of the data, and can also classify according to the structure of the data. For instance, classifying according to the structure of the data can include classifying data with the same data structure into one category. According to the classification results, the number of data in each category and the total number of data in the data cluster are obtained from the data information. For example, according to the classification results, it can be divided into category A data, category B data, and category C data. Among them, category A data includes data 1 and data 2, category B data includes data 3 and data 4, category C data includes data 5 and data 6. The number of data 1 to data 6 is obtained from the data information. Based on the data of data 1 to data 6 obtained, the number of category A data is the sum of the numbers of data 1 and data 2. Similarly, the numbers of category B data and category C data can be obtained, and at the same time, the total number of all data within the data cluster can also be obtained.
[0073] Among the numbers of data in each category, the data with the largest number is selected as the target data. For example, as Figure 4 shown, squares represent category A data, circles represent category B data, and pentagrams represent category C data. Assuming the number of category A data is 100, the number of category B data is 20, and the number of category C data is 10, then among the three categories of data A, B, and C, category A data (i.e., the square in Figure 4 ) is selected as the data with the largest number, and category A data (i.e., Figure 4the square in it) as the target data, which can also be regarded as taking type-A data (i.e., Figure 4 the square in it) as the class represented by this data cluster. The within-cluster purity of this data cluster is the data purity of type-A data (i.e., Figure 4 the square in it). Calculate the ratio of the target data to the total number of data in the data cluster to obtain the within-cluster purity of the data cluster. As shown in Figure 4 , the within-cluster purity is the ratio of the number of type-A data to the total number of type-A, type-B, and type-C data in the data cluster. The calculation formula is as follows:
[0074]
[0075] where purity(P i , C gt ) is the within-cluster purity of the target data, w k is the total number of data in the data cluster, and C gt = c1, c2, …, c M is the result of the original classification.
[0076] For example, still taking the number of type-A data as 100, the number of type-B data as 20, and the number of type-C data as 10 as an example, the target data is type-A data. Then the within-cluster purity of this data cluster is the ratio of the number of type-A data, which is 100, to the total number of data in the data cluster, which is 130, approximately equal to 0.769.
[0077] Optionally, in addition to being pre-set by the operation and maintenance personnel, the trained graph recognition model can also be obtained by self-training of this data mining device. That is, before the step of "extracting features from the data cluster using the trained graph recognition model", this data mining method can further include:
[0078] (1) Collect multiple dataset samples, where the dataset samples include data clusters with labeled within-cluster purity.
[0079] For example, there are various ways to collect multiple dataset samples. For instance, data clusters can be formed by downloading data of known data types and quantities from the Internet, and these data clusters are formed into dataset samples. Then, the within-cluster purity of the data clusters is calculated according to the calculation formula and labeled. It is also possible to upload the dataset samples of known data types and data and the corresponding within-cluster purity of the dataset sample data clusters to this data mining device.
[0080] (2) Use a preset graph recognition model to predict the within-cluster purity of the dataset samples to obtain the predicted within-cluster purity.
[0081] For example, specifically, feature extraction can be performed on the dataset samples to construct a feature space, and node features can be extracted in the feature space to generate graph data of the dataset samples. The graph data includes at least one node. Data clusters corresponding to the nodes are filtered out from the graph data, and the data purity of the data clusters is calculated to obtain the predicted intra-cluster purity of the data clusters in the dataset.
[0082] (3) Converge the preset graph recognition model according to the predicted intra-cluster purity and the labeled intra-cluster purity of the clusters to obtain the trained graph recognition model.
[0083] In the embodiments of the present invention, the preset graph recognition model can be converged according to the intra-cluster purity of the labeled data clusters and the predicted intra-cluster purity in the dataset samples through an interpolation loss function to obtain the trained graph recognition model. For example, specifically, it can be as follows:
[0084] The Dice function (a loss function) is used to adjust the parameters for calculating the intra-cluster purity output in the graph recognition model according to the intra-cluster purity of the labeled data clusters and the predicted intra-cluster purity in the dataset samples, and the interpolation loss function is used to adjust the parameters for calculating the intra-cluster purity output in the graph recognition model according to the intra-cluster purity of the labeled data clusters and the predicted intra-cluster purity in the dataset samples to obtain the trained graph recognition model.
[0085] Optionally, in order to improve the accuracy of the context features, in addition to using the Dice function, other loss functions such as the cross-entropy loss function can also be used for convergence. Specifically, it can be as follows:
[0086] The cross-entropy loss function is used to adjust the parameters for calculating the intra-cluster purity output in the graph recognition model according to the intra-cluster purity of the labeled data clusters and the predicted intra-cluster purity in the dataset samples, and the interpolation loss function is used to adjust the parameters for calculating the intra-cluster purity output in the graph recognition model according to the intra-cluster purity of the labeled data clusters and the predicted intra-cluster purity in the dataset samples to obtain the trained graph recognition model.
[0087] 105. When the intra-cluster purity is lower than the preset purity threshold, obtain the data corresponding to the node in the to-be-processed dataset to obtain the mined data.
[0088] (1) When the intra-cluster purity is lower than the preset purity threshold, obtain the data corresponding to the node in the to-be-processed dataset to obtain the mined data;
[0089] For example, when the intra-cluster purity is lower than the preset purity threshold, determine the target node corresponding to the data. For example, calculate that the intra-cluster purity of the data cluster corresponding to node A is 0.769, and the preset intra-cluster purity is 0.8. Then, the intra-cluster purity of the data cluster corresponding to node A is lower than the preset intra-cluster purity threshold, indicating that node A is the target node corresponding to the data to be mined.
[0090] Filter the graph data corresponding to the target node in the graph data of the dataset to be processed. For example, after determining that the target node is node A, obtain the position of node A in the clustering graph of the graph data according to the clustering subgraph of node A. Based on the position of node A in the clustering graph of the graph data, the target area of node A in the graph data can be further obtained. Obtain the data corresponding to node A in the dataset according to the target area of the graph data, and use this data as the data to be mined in the dataset to be processed, that is, this data is a Bad case to be mined in the dataset. After the mining is completed, continue to calculate the intra-cluster purity of the data cluster corresponding to the next node until the intra-cluster purity of all the data clusters corresponding to the nodes in the graph data is calculated.
[0091] (2) When the intra-cluster purity is not lower than the preset intra-cluster purity threshold, continue to calculate the intra-cluster purity of the data cluster corresponding to the next node.
[0092] For example, when the intra-cluster purity is not lower than the preset intra-cluster purity threshold, obtain the data cluster corresponding to the next target node. For example, the intra-cluster purity of the data cluster corresponding to node A is 0.9, and the preset intra-cluster purity threshold is 0.8. Then the data cluster corresponding to node A does not contain Bad cases. Obtain the data cluster corresponding to node B in the graph data, calculate the intra-cluster purity of the data cluster corresponding to node B, and process the remaining nodes in the graph data corresponding to the dataset to be processed in turn until the intra-cluster purity of all the data clusters corresponding to the nodes is calculated.
[0093] As can be seen from the above, the embodiment of the present invention performs feature extraction on the dataset to be processed to construct a feature space, extracts node features in the feature space to generate the graph data of the dataset to be processed. The graph data includes at least one node. Filter out the data clusters corresponding to the nodes in the graph data, calculate the data purity of the data clusters, and obtain the intra-cluster purity of the data clusters. When the intra-cluster purity is lower than the preset purity threshold, obtain the data corresponding to the node in the dataset to be processed to obtain the mined data; since this solution not only examines all the feature information within the data cluster, but also evaluates the bad cases through the intra-cluster purity within the data cluster, and then conducts bad case mining, reducing the excessive dependence on feature representation, it can more quickly, efficiently, and accurately mine the bad cases (Bad cases) in the data, thereby improving the hit rate of bad cases in the data.
[0094] According to the method described in the above embodiment, the following will be further described in detail with examples.
[0095] In this embodiment, it will be described by taking the data mining device specifically integrated in an electronic device as an example.
[0096] (1) Training of the graph recognition model
[0097] First, the electronic device collects multiple dataset samples. For example, it can download data of known data types and quantities from the Internet to form data clusters, and then form these data clusters into a dataset sample. Next, it calculates the within-cluster purity of the data clusters according to a calculation formula and performs annotation. It can also upload the dataset sample of known data types and data and the corresponding within-cluster purity of the data clusters in the dataset sample to the data mining device.
[0098] Secondly, the electronic device can input the dataset sample into a preset graph recognition model. By extracting features from the dataset sample, it constructs a feature space, extracts node features in the feature space to generate graph data of the dataset sample. The graph data includes at least one node. It filters out the data clusters corresponding to the nodes in the graph data, calculates the data purity of the data clusters, and obtains the predicted within-cluster purity of the data clusters in the dataset.
[0099] Furthermore, the electronic device converges the preset graph recognition model according to the predicted within-cluster purity and the annotated cluster purity to obtain a trained graph recognition model. For example, specifically, the Dice function (a loss function) can be used to adjust the parameters for calculating the output of the within-cluster purity in the graph recognition model according to the within-cluster purity of the annotated data clusters and the predicted within-cluster purity in the dataset sample, and through an interpolation loss function, the parameters for calculating the output of the within-cluster purity in the graph recognition model are adjusted according to the within-cluster purity of the annotated data clusters and the predicted within-cluster purity in the dataset sample to obtain a trained graph recognition model.
[0100] Optionally, to improve the accuracy of context features, in addition to using the Dice function, other loss functions such as the cross-entropy loss function can also be used for convergence. Specifically, it can be as follows:
[0101] Using the cross-entropy loss function, the parameters for calculating the output of the within-cluster purity in the graph recognition model are adjusted according to the within-cluster purity of the annotated data clusters and the predicted within-cluster purity in the dataset sample, and through an interpolation loss function, the parameters for calculating the output of the within-cluster purity in the graph recognition model are adjusted according to the within-cluster purity of the annotated data clusters and the predicted within-cluster purity in the dataset sample to obtain a trained graph recognition model.
[0102] (2) Through the trained graph recognition model, the within-cluster purity of the data clusters corresponding to the nodes in the graph data of the dataset to be processed can be calculated. When the within-cluster purity of the data cluster is lower than the preset purity threshold, the data corresponding to the node in the dataset to be processed is obtained to get the mined data.
[0103] As Figure 5 shown, a data mining method has the following specific process:
[0104] 201. The electronic device obtains the dataset to be processed.
[0105] For example, an electronic device can obtain data from the Internet, such as downloading or collecting, and form a data set. It can also include a user uploading data to a server, and the data mining device obtains the data uploaded by the user from the server to form a data set. The data set can include one type of data or multiple types of data.
[0106] 202. The electronic device performs feature extraction on the data set to be processed to construct a feature space.
[0107] For example, the electronic device can use a deep residual network to perform feature extraction on the data set to be processed, and extract the feature information of the data in the data set to be processed. For example, feature information such as the structure of the data, the relationship between data and data, and / or the type of data. After arranging the extracted feature information in the form of feature vectors and storing them, a feature space is constructed. For example, using the relationship between features, the overall structure of the feature space is constructed, where the connection positions in the overall structure can be the intersections or nodes between features. The feature space is improved or enriched using the attributes of the features themselves to form a feature space that includes all the extracted feature information, and all the feature information is stored in the feature space.
[0108] 203. The electronic device extracts node features in the feature space to generate graph data of the data set to be processed, and the graph data includes at least one node.
[0109] For example, when the electronic device extracts node features in the feature space, it can classify the node features by hierarchical clustering or use the K-nearest neighbor algorithm to classify the node features to obtain different types of node features. For example, according to the different positions of the node features in the feature space, they can be divided into head node features, middle node features, and tail node features.
[0110] The electronic device extracts corresponding node information from the node features of each category according to the classification result. For example, one or more node information of the head is extracted from the head node features, one or more node information of the middle is extracted from the middle node features, and one or more node information of the tail is extracted from the tail node features. A relationship tree is constructed based on the node information of each category mentioned above. For example, a relationship tree can be constructed according to the mutual relationship between the extracted node information. For instance, the mutual relationship among one or more node information of the tail is obtained, the node information of the root node is judged according to the mutual relationship, the root of the relationship tree is constructed based on the node information of the root node, the nodes in the trunk above the root node in the relationship tree are found in turn according to the obtained information of the root node, and the nodes on the branch corresponding to the nodes on the trunk are searched from the remaining node information according to the node information on the main path of the trunk, and these nodes are connected to each other to form the branch and the trunk of the relationship tree. Then, according to the node information on the branch and the trunk, the information of the leaf nodes on the branch is obtained to complete the construction of the relationship tree.
[0111] The electronic device fills or fuses the feature attributes on the nodes into the relationship tree according to the constructed relationship tree, so that each node in the relationship tree includes one or more data. The data in the features can also be mapped to each connection line in the relationship tree to form the edges of the graph data, and finally the graph data of the dataset to be processed is generated. The generated graph data structures and visualizes the data in the dataset to be processed. The mutual relationships such as the positional relationship and the structural relationship between the data in the dataset can be intuitively reflected from the graph data, and the attribute information of the data itself can also be included.
[0112] 204. The electronic device filters the data clusters corresponding to the nodes in the graph data.
[0113] For example, the electronic device randomly selects a node in the graph data, takes the selected node as the target node, and finds the K adjacent nodes of the target node through cosine similarity. K can be any value. The adjacent nodes can include the nodes directly adjacent to the target node, and can also include the nodes whose distance from the target node in the graph data is within a preset distance threshold. For instance, in the graph data, the target node and the remaining other nodes are converted into vectors in space, and the cosine value of the angle between the vectors converted from the remaining other nodes and the target vector converted from the target node is used as a measure or judgment of the cosine distance between the remaining other nodes and the target node. By judging the cosine distance between the remaining other nodes and the target node and the preset distance threshold, the remaining nodes with the cosine distance within the preset distance threshold are used as the adjacent nodes of the target node.
[0114] The electronic device performs hierarchical clustering on the target node and its corresponding adjacent nodes in the graph data to obtain a clustering graph of the target node. For example, in the graph data, each node can be regarded as one or more pieces of data, and the obtained target node and its corresponding adjacent nodes can include multiple or multiple types of data. By calculating the similarity between multiple pieces of data of the target node and its corresponding adjacent nodes, the two most similar pieces of data among all the data are combined, and this process is iterated repeatedly. Finally, the data in the target node and its corresponding adjacent nodes are divided into two or more categories, and a clustering graph is generated.
[0115] The electronic device filters the data cluster corresponding to the target node in the clustering graph. For example, in the clustering graph, according to the target node, a clustering sub-graph of the target node is generated, and the clustering sub-graph is used as the data cluster corresponding to the target node. Among them, the clustering sub-graph can be regarded as a sub-topological graph composed of one or more pieces of data that are closest or most similar to the target node in the clustering graph. It can be seen that the data cluster contains multiple pieces of data, and the multiple pieces of data can be one or more categories.
[0116] 205. The electronic device calculates the data purity of the data cluster to obtain the intra-cluster purity of the data cluster.
[0117] For example, the electronic device uses the trained graph recognition model to extract features from the data cluster to obtain the data information of the data cluster. For example, the data cluster can be subjected to feature extraction through a Graph Convolutional Network (GCN) to obtain the data information of the data cluster. Specifically, it can be as follows:
[0118] Each node in the data cluster sends its own feature information to its respective adjacent nodes after transformation. At this time, the adjacent nodes include the adjacent nodes directly connected by the side lines. Each node aggregates the feature information sent by its respective adjacent nodes to perform feature information fusion to obtain the data information in the data cluster. Among them, the data information includes the total number of all data in the data cluster and also includes the attribute information of all the data.
[0119] Among them, the selected target node v j is used as the central vertex input to the GCN model. The GCN model uses the clustering sub-graph (data cluster) composed of the vertex set associated with the visual features of the target node v j as the input, and performs feature extraction on the clustering sub-graph (data cluster) to obtain the data information of the clustering sub-graph (data cluster). The calculation formula is as follows:
[0120]
[0121] Among them, A(P i ) is the target node v jA clustering subgraph composed of vertex sets that are visually feature - related and are respectively the data information of the target node and any associated node in the clustering subgraph is a diagonal matrix, I is the identity matrix, F l (P i ) is the feature representation of the l - th layer, W l is the feature mapping learned by the GCN model of the l - th layer, and σ is the activation function. Here, it should be noted that the activation function can be selected as ReLU (an activation function).
[0122] After the GCN model in the electronic device obtains the data information within the data cluster through feature extraction, it classifies the data within the data cluster according to the data information. For example, it can classify according to the attribute information of the data, or it can also classify according to the structure of the data. For instance, classifying according to the structure of the data can include classifying data with the same data structure into one category. According to the classification results, the quantity of each category of data and the total quantity of data in the data cluster are obtained in the data information. For example, according to the classification results, it can be divided into category A data, category B data, and category C data. Among them, category A data includes data 1 and data 2, category B data includes data 3 and data 4, category C data includes data 5 and data 6. The quantities of data 1 to data 6 are obtained in the data information. Based on the data of data 1 to data 6 obtained, the quantity of category A data is the sum of the quantities of data 1 and data 2. Similarly, the quantities of category B data and category C data can be obtained, and at the same time, the total quantity of all data within the data cluster can also be obtained.
[0123] Among the quantities of each category of data, the data with the largest quantity is selected as the target data. For example, if the quantity of category A data is 100, the quantity of category B data is 20, and the quantity of category C data is 10, then among the three categories of data A, B, and C, category A data is selected as the data with the largest quantity and used as the target data. It can also be regarded as using category A data as the class represented by this data cluster, and the within - cluster purity of this data cluster is the data purity of category A data. Calculate the ratio of the target data to the total quantity of data in the data cluster to obtain the within - cluster purity of the data cluster. The calculation formula is as follows
[0124]
[0125] where, purity(P i ,C gt ) is the within - cluster purity of the target data, w k is the total quantity of data within the data cluster, C gt =c1,c2,…,c M is the result of the original classification, and c j is the quantity of each data within the data cluster.
[0126] For example, still taking the number of Class A data as 100, the number of Class B data as 20, and the number of Class C data as 10 as an example, if the target data is Class A data, then the intra-cluster purity of this data cluster is the ratio of the number of 100 Class A data to the total number of 130 data within the data cluster, which is approximately equal to 0.769.
[0127] 206. When the intra-cluster purity is lower than the preset purity threshold, the electronic device obtains the data corresponding to the node in the data set to be processed, and obtains the mined data.
[0128] For example, when the intra-cluster purity is lower than the preset purity threshold, the electronic device determines the target node corresponding to the data. For example, it calculates that the intra-cluster purity of the data cluster corresponding to Node A is 0.769, and the preset intra-cluster purity is 0.8. Then the intra-cluster purity of the data cluster corresponding to Node A is lower than the preset intra-cluster purity threshold, indicating that Node A is the target node corresponding to the data to be mined.
[0129] The electronic device filters the graph data corresponding to the target node in the graph data of the data set to be processed. For example, after determining that the target node is Node A, it obtains the position of Node A in the clustering graph of the graph data according to the clustering sub-graph of Node A. Based on the position of Node A in the clustering graph of the graph data, it can continue to obtain the target area of Node A in the graph data, and obtain the data corresponding to Node A in the data set according to the target area of the graph data. This data is used as the data to be mined in the data set to be processed, that is, this data is a Bad case to be mined in the data set. After the mining is completed, it continues to calculate the intra-cluster purity of the data cluster corresponding to the next node until the intra-cluster purity of all the data clusters corresponding to the nodes in the graph data is calculated.
[0130] 207. When the intra-cluster purity is not lower than the preset intra-cluster purity threshold, continue to calculate the intra-cluster purity of the data cluster corresponding to the next node.
[0131] For example, when the intra-cluster purity is not lower than the preset intra-cluster purity threshold, the electronic device obtains the data cluster corresponding to the next target node. For example, the intra-cluster purity of the data cluster corresponding to Node A is 0.9, and the preset intra-cluster purity threshold is 0.8. Then the data cluster corresponding to Node A does not contain Bad cases. Obtain the data cluster corresponding to Node B in the graph data, and calculate the intra-cluster purity of the data cluster corresponding to Node B. Process the remaining nodes in the graph data corresponding to the data set to be processed in turn until the intra-cluster purity of all the data clusters corresponding to the nodes is calculated.
[0132] As can be seen from the above, in this embodiment, the electronic device extracts features from the dataset to be processed to construct a feature space, extracts node features in the feature space to generate graph data of the dataset to be processed, and the graph data includes at least one node. Then, the data clusters corresponding to the nodes are screened out from the graph data, and the data purity of the data clusters is calculated to obtain the intra-cluster purity of the data clusters. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the nodes in the dataset to be processed is obtained to get the mined data. Since this solution not only examines all the feature information within the data clusters, but also evaluates bad cases through the intra-cluster purity within the data clusters, and then conducts bad case mining, reducing the excessive dependence on feature representation, it can more quickly, efficiently, and accurately mine bad cases (Bad case) in the data, thereby improving the hit rate of bad cases in the data.
[0133] To better implement the above method, an embodiment of the present invention further provides a data mining device. The data mining device can be integrated in an electronic device, such as a server or a terminal, etc. The terminal can include a tablet computer, a laptop computer, and / or a personal computer, etc.
[0134] For example, as Figure 6 shown, the data mining device may include an extraction unit 301, a generation unit 302, a screening unit 303, a calculation unit 304, and an acquisition unit 305, as follows:
[0135] (1) Extraction unit 301;
[0136] The extraction unit 301 is configured to extract features from the dataset to be processed to construct a feature space.
[0137] The extraction unit 301 may include an acquisition subunit 3011 and an extraction subunit 3012, as Figure 7 shown, specifically as follows:
[0138] The acquisition subunit 3011 is configured to acquire the dataset to be processed;
[0139] The first extraction subunit 3012 is configured to extract features from the dataset to be processed to construct a feature space.
[0140] For example, the acquisition subunit 3011 acquires the dataset to be processed, and the extraction subunit 3012 extracts features from the dataset to be processed to construct a feature space.
[0141] (2) Generation unit 302;
[0142] The generation unit 302 is configured to extract node features in the feature space to generate graph data of the dataset to be processed, and the graph data includes at least one node.
[0143] Among them, the generation unit 302 may include a second extraction subunit 3021, a first classification subunit 3022, and a generation subunit 3023, as Figure 8 shown;
[0144] The second extraction subunit 3021 is configured to extract node features in the feature space;
[0145] The first classification subunit 3022 is configured to classify the node features;
[0146] The generation subunit 3023 is configured to generate graph data of the dataset to be processed according to the classification result.
[0147] For example, the second extraction subunit 3021 extracts node features in the feature space, the first classification subunit 3022 classifies the node features, and the generation subunit 3023 generates graph data of the dataset to be processed according to the classification result.
[0148] (3) The screening unit 303;
[0149] The screening unit 303 is configured to screen out the data clusters corresponding to the nodes in the graph data;
[0150] Among them, the screening unit 303 may include a search subunit 3031, a clustering subunit 3032, and a screening subunit 3033, as Figure 9 shown,
[0151] The search subunit 3031 is configured to search for the adjacent nodes corresponding to the nodes in the graph data;
[0152] The clustering subunit 3032 is configured to cluster the nodes and the corresponding adjacent nodes in the graph data to obtain a clustering graph of the nodes;
[0153] The screening subunit 3033 is configured to screen out the data clusters corresponding to the nodes in the clustering graph.
[0154] For example, the search subunit 3031 searches for the adjacent nodes corresponding to the nodes in the graph data, the clustering subunit 3032 clusters the nodes and the corresponding adjacent nodes in the graph data to obtain a clustering graph of the nodes, and the screening subunit 3033 screens out the data clusters corresponding to the nodes in the clustering graph.
[0155] (4) The calculation unit 304;
[0156] The calculation unit 304 is configured to calculate the data purity of the data clusters to obtain the intra-cluster purity of the data clusters.
[0157] Among them, the calculation unit 304 may include a third extraction unit 3041, a second classification unit 3042, and a calculation subunit 3043, as Figure 10 shown, specifically as follows:
[0158] The third extraction unit 3041 is configured to extract features from the data clusters by using the trained graph recognition model to obtain the data information of the data clusters;
[0159] The second classification subunit 3042 is configured to classify the data within the data clusters according to the data information;
[0160] The calculation subunit 3043 is configured to calculate the data purity of the data clusters according to the classification results to obtain the intra-cluster purity of the data clusters.
[0161] For example, the third extraction unit 3041 extracts features from the data clusters by using the trained graph recognition model to obtain the data information of the data clusters. The second classification subunit 3042 classifies the data within the data clusters according to the data information. The calculation subunit 3043 calculates the data purity of the data clusters according to the classification results to obtain the intra-cluster purity of the data clusters.
[0162] (5) The acquisition unit 305;
[0163] The acquisition unit 305 is configured to, when the intra-cluster purity is lower than the preset intra-cluster purity threshold, acquire the data corresponding to the node in the to-be-processed data set and use the data as the data to be mined.
[0164] For example, when the cluster purity is lower than the preset intra-cluster purity threshold, determine the target node corresponding to the data cluster, screen the graph data corresponding to the target node in the graph data of the to-be-processed data set, acquire the data corresponding to the node in the to-be-processed data set according to the graph data corresponding to the node, and use the data as the data to be mined in the to-be-processed data set; when the intra-cluster purity is not lower than the preset intra-cluster purity threshold, continue to calculate the intra-cluster purity of the data clusters corresponding to the next node.
[0165] Optionally, the trained recognition model can be obtained not only by being pre-set by the operation and maintenance personnel, but also by self-training of the graph recognition model. That is, as Figure 11 shown, the recognition model may further include an acquisition unit 306 and a training unit 307, as follows:
[0166] The acquisition unit 306 is configured to acquire multiple data set samples, and the data cluster samples include data clusters with labeled intra-cluster purity.
[0167] For example, the acquisition unit 306 can download data of known data types and quantities from the Internet to form data clusters, form data set samples from these data clusters, calculate the intra-cluster purity of the data clusters according to the calculation formula and perform labeling, and can also upload the data set samples of known data types and data and the corresponding intra-cluster purity of the data cluster samples of the data set to the data mining device.
[0168] A training unit 307 is configured to predict the intra-cluster purity of a dataset sample using a preset graph recognition model, obtain the predicted intra-cluster purity, and converge the preset graph recognition model based on the predicted intra-cluster purity and the labeled cluster purity to obtain a trained graph recognition model.
[0169] For example, the training unit 307 can specifically extract features from the dataset samples to construct a feature space, extract node features in the feature space to generate graph data of the dataset samples, the graph data includes at least one node, filter out the data clusters corresponding to the nodes in the graph data, calculate the data purity of the data clusters to obtain the predicted intra-cluster purity of the data clusters in the dataset. Thereafter, the preset graph recognition model can be converged based on the predicted intra-cluster purity and the labeled cluster purity to obtain a trained graph recognition model.
[0170] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the method embodiments described above, which will not be elaborated here.
[0171] As can be seen from the above, in this embodiment, the extraction unit 301 extracts features from the dataset to be processed to construct a feature space, the generation unit 302 extracts node features in the feature space to generate graph data of the dataset to be processed, the graph data includes at least one node, the screening unit 303 filters out the data clusters corresponding to the nodes in the graph data, the calculation unit 304 calculates the data purity of the data clusters to obtain the intra-cluster purity of the data clusters, and the acquisition unit 305 obtains the data corresponding to the nodes in the dataset to be processed when the intra-cluster purity is lower than a preset purity threshold to obtain the mined data; since this solution not only examines all the feature information within the data clusters, but also evaluates bad cases through the intra-cluster purity within the data clusters, and then conducts bad case mining, reducing the excessive dependence on feature representation, it can more quickly, efficiently, and accurately mine the bad cases (Badcase) in the data, thereby improving the hit rate of bad cases in the data.
[0172] An embodiment of the present invention further provides an electronic device, such as Figure 12 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:
[0173] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 12 the structural diagram of the electronic device shown in
[0174] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and calling the data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby performing an overall detection of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.
[0175] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.
[0176] The electronic device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0177] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0178] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:
[0179] Feature extraction is performed on the dataset to be processed to construct a feature space. Node features are extracted in the feature space to generate graph data of the dataset to be processed. The graph data includes at least one node. In the graph data, the data cluster corresponding to the node is screened, and the data purity of the data cluster is calculated to obtain the intra-cluster purity of the data cluster. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the node in the dataset to be processed is obtained to get the mined data.
[0180] For example, data can be specifically obtained from the Internet, such as downloading or collecting, and the data is formed into a dataset. It can also include that the user uploads data to the server, and the data mining device obtains the data uploaded by the user from the server to form a dataset. A deep residual network is used to perform feature extraction on the dataset to be processed, and the feature information of the data in the dataset to be processed is extracted. Node features are extracted in the feature space, and the node features are classified by hierarchical clustering. The K-nearest neighbor algorithm can also be used to classify the node features to obtain different types of node features. According to the classification results, the corresponding node information is extracted from each type of node feature, and a relationship tree is constructed based on the extracted node information of each type. According to the constructed relationship tree, the graph data of the dataset to be processed is generated. A node is randomly selected in the graph data, and the selected node is used as the target node. The adjacent nodes corresponding to the target node are searched using the selected target node. Hierarchical clustering is performed on the target node and its corresponding adjacent nodes in the graph data to obtain the clustering graph of the target node. The data cluster corresponding to the target node is screened in the clustering graph, and the trained graph recognition model is used to perform feature extraction on the data cluster to obtain the data information of the data cluster. After the GCN model obtains the data information in the data cluster through feature extraction, the data in the data cluster is classified according to the data information. For example, it can be classified according to the attribute information of the data, or it can also be classified according to the structure of the data. According to the classification results, the quantity of each category of data and the total quantity of the data in the data cluster are obtained from the data information. The data with the largest quantity is screened from the quantity of each category of data as the target data, and the ratio of the target data to the total quantity of the data in the data cluster is calculated to obtain the intra-cluster purity of the data cluster. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the node in the dataset to be processed is obtained to get the mined data. When the intra-cluster purity is not lower than the preset intra-cluster purity threshold, the calculation of the intra-cluster purity of the data cluster corresponding to the next node is continued.
[0181] Optionally, the trained graph recognition model can be obtained not only by being pre-set by the operation and maintenance personnel, but also by being self-trained by the data mining device. That is, the instruction can also perform the following steps:
[0182] Collect multiple dataset samples, where the dataset samples include data clusters with labeled intra-cluster purity. Use a preset graph recognition model to predict the intra-cluster purity of the dataset samples to obtain the predicted intra-cluster purity. Converge the preset graph recognition model based on the predicted intra-cluster purity and the labeled cluster purity to obtain the trained graph recognition model.
[0183] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.
[0184] As can be seen from the above, in the embodiments of the present invention, feature extraction is performed on the dataset to be processed to construct a feature space, node features are extracted in the feature space to generate graph data of the dataset to be processed, the graph data includes at least one node, data clusters corresponding to the nodes are screened out in the graph data, the data purity of the data clusters is calculated to obtain the intra-cluster purity of the data clusters. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the nodes in the dataset to be processed is obtained to get the mined data. Since this solution not only examines all the feature information within the data clusters, but also evaluates bad cases through the intra-cluster purity within the data clusters, and then conducts bad case mining, reducing the excessive dependence on feature representation, it can more quickly, efficiently, and accurately mine bad cases (Bad case) in the data, thereby improving the hit rate of bad cases in the data.
[0185] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0186] Therefore, an embodiment of the present invention provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any of the data mining methods provided by the embodiments of the present invention. For example, the instructions can execute the following steps:
[0187] Perform feature extraction on the dataset to be processed to construct a feature space, extract node features in the feature space to generate graph data of the dataset to be processed, the graph data includes at least one node, screen out the data clusters corresponding to the nodes in the graph data, calculate the data purity of the data clusters to obtain the intra-cluster purity of the data clusters. When the intra-cluster purity is lower than the preset purity threshold, obtain the data corresponding to the nodes in the dataset to be processed to get the mined data.
[0188] For example, data can be specifically obtained from the Internet, such as downloading or collecting, and the data is formed into a data set. It can also include that the user uploads data to the server, and the data mining device obtains the data uploaded by the user from the server to form a data set. The deep residual network is used to extract features from the data set to be processed, and the feature information of the data in the data set to be processed is extracted. Node features are extracted in the feature space, and the node features are classified by hierarchical clustering. The K-nearest neighbor algorithm can also be used to classify the node features to obtain different types of node features. According to the classification results, the corresponding node information is extracted from each type of node feature, and a relationship tree is constructed based on the extracted node information of each type. According to the constructed relationship tree, the graph data of the data set to be processed is generated. A node is randomly selected from the graph data, and the selected node is used as the target node to search for the adjacent nodes corresponding to the target node. Hierarchical clustering is performed on the target node and its corresponding adjacent nodes in the graph data to obtain the clustering graph of the target node. The data cluster corresponding to the target node is screened in the clustering graph, and the trained graph recognition model is used to extract the features of the data cluster to obtain the data information of the data cluster. After the GCN model obtains the data information in the data cluster through feature extraction, it classifies the data in the data cluster according to the data information. For example, it is classified according to the attribute information of the data, or it can also be classified according to the structure of the data. According to the classification results, the number of data in each category and the total number of data in the data cluster are obtained from the data information. The data with the largest number is screened from the number of data in each category as the target data, and the ratio of the target data to the total number of data in the data cluster is calculated to obtain the intra-cluster purity of the data cluster. When the intra-cluster purity is lower than the preset purity threshold, the data corresponding to the node in the data set to be processed is obtained to get the mined data. When the intra-cluster purity is not lower than the preset intra-cluster purity threshold, the calculation of the intra-cluster purity of the data cluster corresponding to the next node is continued.
[0189] Optionally, the trained graph recognition model can be obtained not only by being pre-set by the operation and maintenance personnel, but also by being self-trained by the data mining device. That is, the instruction can also perform the following steps:
[0190] Collect multiple data set samples, where the data cluster samples include data clusters with labeled intra-cluster purity. Use the preset graph recognition model to predict the intra-cluster purity of the data set samples to obtain the predicted intra-cluster purity. Converge the preset graph recognition model according to the predicted intra-cluster purity and the labeled cluster purity to obtain the trained graph recognition model.
[0191] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.
[0192] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0193] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the data mining methods provided by the embodiments of the present invention, the beneficial effects achievable by any of the data mining methods provided by the embodiments of the present invention can be realized. For details, please refer to the previous embodiments and will not be elaborated here.
[0194] The above has introduced in detail a data mining method, device, and computer-readable storage medium provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A data mining method, characterized in that, Applied to an electronic device, including: Performing feature extraction on a dataset to be processed to construct a feature space, where the types of data in the dataset to be processed include image type or text type; Extracting node features in the feature space to generate graph data of the dataset to be processed, where the graph data includes at least one node; Filtering out the data clusters corresponding to the nodes in the graph data, including: searching for adjacent nodes corresponding to the nodes in the graph data; clustering the nodes and the corresponding adjacent nodes in the graph data to obtain a clustering graph of the nodes, where the adjacent nodes include nodes directly adjacent to the nodes and nodes whose distance from the nodes in the graph data is within a preset distance threshold; generating a clustering subgraph of the nodes in the clustering graph, and using the clustering subgraph as the data cluster corresponding to the nodes; Calculating the data purity of the data cluster to obtain the intra-cluster purity of the data cluster, including: each node in the data cluster aggregates the feature information sent by the adjacent nodes directly connected to it to perform feature information fusion, obtaining the data information in the data cluster; classifying the data in the data cluster according to the data information; calculating the data purity of the data cluster according to the classification result to obtain the intra-cluster purity of the data cluster; When the intra-cluster purity of the data cluster corresponding to the node is lower than a preset intra-cluster purity threshold, obtaining the data corresponding to the node in the dataset to be processed according to the position of the clustering subgraph of the node in the clustering graph, obtaining the mined data, where the mined data is a bad file, and the data clusters corresponding to the nodes with intra-cluster purity not lower than the preset intra-cluster purity threshold do not contain the bad file, and the bad file means that there are multiple different types of data in the corresponding data cluster.
2. The data mining method according to claim 1, wherein Calculating the data purity of the data cluster to obtain the intra-cluster purity of the data cluster, including: Obtaining the quantity of data of each category and the total quantity of data in the data cluster from the data information according to the classification result; Filtering out the data with the largest quantity among the quantities of data of each category as the target data; Calculating the ratio of the target data to the total quantity of data in the data cluster to obtain the intra-cluster purity of the data cluster.
3. The data mining method according to claim 1, characterized in that Before performing feature extraction on the data cluster using the trained graph recognition model, it further includes: Collecting multiple dataset samples, where the dataset samples include data clusters with labeled cluster purities; Predicting the cluster purity of the dataset samples using a preset graph recognition model to obtain a predicted cluster purity; Converging the preset graph recognition model according to the predicted cluster purity and the labeled cluster purity to obtain a trained graph recognition model.
4. The data mining method according to any one of claims 1 to 3, characterized in that When the intra-cluster purity of the data cluster corresponding to the node is lower than a preset intra-cluster purity threshold, obtaining the data corresponding to the node in the dataset to be processed and using the data as the data to be mined, including: When the cluster purity of the data cluster corresponding to the node is lower than a preset intra-cluster purity threshold, determining the target node corresponding to the data cluster; Filter the graph data corresponding to the target node from the graph data of the dataset to be processed; According to the graph data corresponding to the node, obtain the data corresponding to the node from the dataset to be processed, and use the data as the data to be mined in the dataset to be processed.
5. The data mining method according to any one of claims 1 to 3, characterized in that Extract node features in the feature space to generate the graph data of the dataset to be processed, where the graph data includes at least one node, including: Extract node features in the feature space; Classify the node features; Generate the graph data of the dataset to be processed according to the classification result.
6. The data mining method according to claim 5, wherein Generate the graph data of the dataset to be processed according to the classification result, including: Extract the node information in the node features of each category according to the classification result; Construct a relationship tree according to the node information; Generate the graph data of the dataset to be processed based on the constructed relationship tree.
7. A data mining device, characterized in that, Applied to an electronic device, including: An extraction unit for extracting features from a dataset to be processed to construct a feature space, where the data types in the dataset to be processed include image types or text types; A generation unit for extracting node features in the feature space to generate the graph data of the dataset to be processed, where the graph data includes at least one node; A screening unit for screening out the data cluster corresponding to the node in the graph data, including: searching for the adjacent nodes corresponding to the node in the graph data; clustering the node and the corresponding adjacent nodes in the graph data to obtain the clustering graph of the node, where the adjacent nodes include the nodes directly adjacent to the node and the nodes whose distance from the node in the graph data is within a preset distance threshold; generating a clustering subgraph of the node in the clustering graph, and using the clustering subgraph as the data cluster corresponding to the node; A calculation unit for calculating the data purity of the data cluster to obtain the intra-cluster purity of the data cluster, including: each node in the data cluster aggregates the feature information sent by the adjacent nodes directly connected to it for feature information fusion to obtain the data information in the data cluster; classifying the data in the data cluster according to the data information; calculating the data purity of the data cluster according to the classification result to obtain the intra-cluster purity of the data cluster; An acquisition unit for, when the intra-cluster purity of the data cluster corresponding to the node is lower than a preset purity threshold, obtaining the data corresponding to the node in the dataset to be processed according to the position of the clustering subgraph of the node in the clustering graph, to obtain the mined data, where the mined data is a bad file, and the data clusters corresponding to the nodes with intra-cluster purity not lower than the preset intra-cluster purity threshold do not contain the bad file, and the bad file means that there are multiple different types of data in the corresponding data cluster.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the data mining method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Sample processing method, device, apparatus, and storage medium
CN109242106A