Data label automatic classification method based on semi-supervised learning
By constructing an initial classifier to predict the label-free sample dataset and perform accuracy tests, the classification overfitting problem caused by too few labels in semi-supervised learning is solved, and the credibility and generalization performance of label prediction is improved.
Patent Information
- Application Number
- CN202411797253.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-06
AI Technical Summary
In semi-supervised learning, the classification overfitting results in too few labels, and the generated labels are not reliable, lack specific standard verification processes, which reduces the credibility of label prediction.
By obtaining labeled sample data sets, building an initial classifier, label prediction is performed on labelless sample data sets, and accuracy test is performed based on preset evaluation indicators, improving the credibility and generalization performance of label prediction.
The classification accuracy and performance of the classifier are improved, and the credibility of the label prediction results is judged through accuracy test, which enhances the generalization performance of label prediction.
Smart Images

Figure CN119939299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of label data classification, and in particular to a method for automatic classification of data labels based on semi-supervised learning. Background Art
[0002] At present, with the popularization of computer equipment, data information in computer data processing has shown an exponential growth. These data tend to be unstructured or semi-structured, and unstructured data cannot be used efficiently. In real task scenarios, only a few samples are labeled, and the vast majority are unlabeled sample data. Traditional text feature extraction cannot obtain high-quality category labels and cannot solve the problem of multi-label text classification. Therefore, the classification effect of unlabeled data is not good. In traditional label classification tasks, text categories are mutually exclusive, and any text has only one category label. In reality, a text has multiple category labels. Therefore, when there are fewer labeled sample data and the number of unlabeled samples is small, the number of unlabeled samples is small. When there is too much data, it is easy to cause overfitting of the classification results. For example, the invention patent CN201910883908 - A semi-supervised image classification algorithm based on ACGAN solves the problem of overfitting caused by too few labels in semi-supervised learning. The generator in the ACGAN network generates data and its corresponding labels, and puts the data into the classifier for classification, increasing the amount of labeled data, thereby improving the generalization ability of the classification model. However, this technical solution will make the generated labels unreliable due to the classification performance of the classifier, and there is no specific standard verification process for the credibility of the label prediction results, which reduces the credibility of the label prediction. Summary of the invention
[0003] The present invention provides a data label automatic classification method based on semi-supervised learning, which is used to solve the situation that the label prediction result has low accuracy and lacks reliability.
[0004] A method for automatic classification of data labels based on semi-supervised learning, comprising:
[0005] Obtain a labeled sample data set, and construct a classifier according to the label information of the labeled sample data set to determine an initial classifier;
[0006] Obtain an unlabeled sample data set, perform label prediction on the unlabeled sample data set based on the initial classifier, and obtain a label prediction result;
[0007] The data label prediction result is subjected to an accuracy test based on a preset evaluation index to obtain an accuracy test result.
[0008] As an embodiment of the present invention, the step of obtaining a labeled sample data set, and constructing a classifier according to label information of the labeled sample data set to determine an initial classifier includes:
[0009] Obtain a sample data set, and classify the sample data set according to whether it has a label or not, and determine the classification result of the sample data set; wherein,
[0010] The sample data set classification results include: a labeled sample data set and an unlabeled sample data set;
[0011] Obtaining text and corresponding labels in the labeled sample data set, and establishing a mapping relationship between the text and the corresponding labels;
[0012] Based on the mapping relationship between text and labels in the labeled sample data set, an initial classifier is constructed.
[0013] As an embodiment of the present invention, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result includes:
[0014] Based on the initial classifier, a mapping relationship between text and labels in the labeled sample data set is obtained, and a mapping label is generated according to the mapping relationship between the text and the label;
[0015] Preprocessing the unlabeled sample data to determine the preprocessing results; wherein the preprocessing content includes: sentence decomposition, part of speech decomposition, stop word removal, and noise cleaning;
[0016] According to the preprocessing result of the unlabeled sample data and the mapping label, label prediction is performed on the preprocessed unlabeled sample data set to determine the label prediction result.
[0017] As an embodiment of the present invention: the data label prediction result is subjected to an accuracy test based on preset evaluation indicators to obtain an accuracy test result, wherein the evaluation indicators include: Hamming loss, sorting loss, coverage efficiency, precision, and error rate.
[0018] As an embodiment of the present invention, the step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier further includes:
[0019] Performing label correlation calculation on the labeled sample data to obtain label correlation calculation results;
[0020] Constructing a label co-occurrence matrix according to the label correlation calculation result and the labeled sample data set to obtain category label distribution information in the labeled sample data set; wherein the label co-occurrence matrix is used to calculate the distribution between labels in the labeled sample data set;
[0021] Performing a correlation analysis on the category label distribution information of the labeled sample data set to obtain the correlation coefficient between the category labels;
[0022] An initial classifier is constructed based on the correlation coefficients between the class labels.
[0023] As an embodiment of the present invention, the step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier further includes:
[0024] Obtain a sample data set, determine whether the sample data set is a linear data set, and when the determination result shows that the sample data set is a linear data set, determine whether the sample data set is in a linearly separable state.
[0025] When the judgment result shows that the sample data set is in a linearly separable state, a hyperplane of the sample data set is obtained; wherein the hyperplane is used to divide the linear space into two non-intersecting parts; according to the hyperplane, a corresponding normal vector, an offset, and a support vector are obtained, and a hyperplane segmentation distance is calculated based on the support vector, and an optimal hyperplane is determined according to the hyperplane segmentation distance; wherein the optimal hyperplane is a hyperplane with the largest corresponding hyperplane segmentation distance;
[0026] When the judgment result shows that the sample data set is in a linearly inseparable state, the sample data set is converted into a high-dimensional space through a kernel function, and the hyperplane of the sample data set in the high-dimensional space is obtained. The corresponding normal vector, offset, and support vector are obtained according to the hyperplane in the high-dimensional space, and the hyperplane segmentation distance is calculated based on the support vector, and the optimal hyperplane is determined according to the hyperplane segmentation distance; wherein the kernel function is used to convert the linearly inseparable state of the data set into a linearly separable state in the high-dimensional space.
[0027] As an embodiment of the present invention, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes:
[0028] Performing an integrated analysis on the category labels according to the labeled sample data set to obtain the co-occurrence frequency corresponding to the category labels; wherein the co-occurrence frequency is used to calculate the correlation coefficient between the labels;
[0029] According to the co-occurrence frequency of the category labels in the labeled sample data, the label information of the neighboring text samples is counted to determine the label information statistics of the neighboring text samples; wherein the neighboring text samples are text samples whose label co-occurrence frequency in the sample data set is within a preset threshold range;
[0030] Based on the co-occurrence frequency of the category labels and the statistical results of the label information of the neighboring text samples, a label prediction model is established, the unlabeled sample data set is used as the input item of the label prediction model, the label prediction result is output, and the label prediction sequence of the unlabeled sample data is obtained; wherein the label prediction model is used to output the corresponding label category information according to the input sample data set.
[0031] As an embodiment of the present invention, the step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier further includes:
[0032] Randomly sampling the labeled sample data set to divide it into a first labeled sample data set and a second labeled sample data set;
[0033] Constructing a first initial classifier based on the first labeled sample data set, and constructing a second initial classifier based on the second labeled sample data set; wherein the first initial classifier is trained and obtained using the IMML-KNN algorithm, and the second initial classifier is trained and obtained using the MLPLSA-KNN algorithm;
[0034] Traversing the first initial classifier and the second initial classifier respectively on the unlabeled sample data set to obtain a first traversal result and a second traversal result;
[0035] Calculate the confidence of the first traversal result and the second traversal result, select the ones with the highest confidence as pseudo labels, save the pseudo labels corresponding to the first traversal result in a first set, and save the pseudo labels corresponding to the second traversal result in a second set;
[0036] Adding the first set to a first labeled sample data set as an updated first labeled sample data set, and adding the second set to a second labeled sample data set as an updated second labeled sample data set;
[0037] The updated first labeled sample data set and the updated second labeled sample data set are respectively updated for the first initial classifier and the second initial classifier to obtain the updated first initial classifier and the updated second initial classifier, accuracy calculation is performed for the updated first initial classifier and the updated second initial classifier to obtain the accuracy calculation result, and according to the accuracy calculation result, the one with a higher corresponding calculation result is used as the target classifier, and the unlabeled sample data set is classified based on the target classifier to obtain the label classification result.
[0038] As an embodiment of the present invention: the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: performing entity type prediction on a knowledge graph based on the label prediction result, wherein the entity type prediction step of the knowledge graph includes:
[0039] Step 1: Obtain the knowledge graph input by the user, parse the knowledge graph, obtain the knowledge graph information, generate the corresponding information description according to the knowledge graph information, and determine the knowledge graph information description result; wherein the knowledge graph information includes: structured information and unstructured information;
[0040] Step 2: According to the description result of the knowledge graph information, determine each entity in the knowledge graph, and extract features for the text information and link information of each entity to generate a feature data set; wherein the feature data set includes: a labeled sample data set and an unlabeled sample data set, the labeled sample data set is used as training data, and the unlabeled sample data set is used as prediction data;
[0041] Step 3: Based on the results of the knowledge graph information description, perform model training on the training data, determine the training model, input the predicted data as input to the training model, obtain the predicted type of the entity, generate type knowledge based on each entity in the knowledge graph and the corresponding predicted type, and determine the predicted result of the entity type in the knowledge graph; wherein the training model makes predictions for the entity types in the knowledge graph.
[0042] As an embodiment of the present invention, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: detecting data faults in the sensor network according to the label prediction result, and determining a fault detection result; wherein the execution step of the fault detection includes:
[0043] Acquire a data stream in the sensor network, segment the data stream to obtain at least two segments of data stream, and perform data stream category label prediction based on the label prediction result for each segment of data stream;
[0044] The data stream label prediction result is compared with a preset threshold range. When the data stream label prediction result is not within the threshold range, it is determined that a fault has occurred in the process of building the initial classifier, and the fault type is determined; wherein the fault types include: outlier fault, spike fault, and high noise fault.
[0045] The beneficial effects of the above technical solution are:
[0046] In the present invention, a classifier is constructed according to the mapping relationship between text and labels in a labeled sample data set, as well as the connection between labels, which is beneficial to improving the classification accuracy and performance of the classifier. Label prediction is performed on an unlabeled sample data set by the constructed classifier, and the label prediction results of the unlabeled sample data set are verified, which is beneficial to judging whether the result of label judgment is credible and improving the generalization performance of label prediction.
[0047] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0048] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0050] Figure 1 A schematic diagram of a flow chart of a method for automatic classification of data labels based on semi-supervised learning in an embodiment of the present invention;
[0051] Figure 2 A schematic diagram of a label prediction process in a method for automatic classification of data labels based on semi-supervised learning in an embodiment of the present invention;
[0052] Figure 3 The present invention is a schematic diagram of the accuracy test process for predicting labels in a method for automatic classification of data labels based on semi-supervised learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0054] It should be noted that when a component is referred to as being "fixed to" or "disposed on" another component, it can be directly on the other component or indirectly on the other component. When a component is referred to as being "connected to" another component, it can be directly or indirectly connected to the other component.
[0055] It should be understood that the orientation or position relationship indicated by terms such as "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside" and "outside" are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0056] In addition, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations, and "multiple" means two or more, unless otherwise clearly and specifically limited. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0057] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
[0058] In many scenarios, such as image recognition, speech recognition, and natural language processing, a large amount of unlabeled data needs to be annotated for further use. However, manually assigning labels to each picture or each text is a time-consuming and costly task. Therefore, how to automatically generate labels from a large amount of unlabeled data has become a problem that needs to be solved. The prior art can build a classifier for classification. When the classifier automatically classifies the labels, there may be problems with the performance of the classifier, and there may be insufficient training data, resulting in the result of label classification not necessarily being trustworthy. The present invention has made improvements to the technical problems in this regard, such as the following embodiments:
[0059] Embodiment 1:
[0060] The embodiment of the present invention provides a method for automatic classification of data labels based on semi-supervised learning. Figure 1 As shown, including:
[0061] Obtain a labeled sample data set, and construct a classifier according to the label information of the labeled sample data set to determine an initial classifier;
[0062] Obtain an unlabeled sample data set, perform label prediction on the unlabeled sample data set based on the initial classifier, and obtain a label prediction result;
[0063] The data label prediction result is subjected to an accuracy test based on a preset evaluation index to obtain an accuracy test result.
[0064] The principle of the above technical solution is:
[0065] In the actual implementation process of the present invention, in the prior art, when classifying samples, any sample only contains one category label. In this case, it is only necessary to find the closest category label. However, in the actual sample classification task, a sample generally contains multiple category label information, and the multiple category labels are independent of each other. If the method of sampling to find the closest category label is still used, it is easy to cause the text information to be single, which is poor in effect when performing information retrieval and recommendation, and does not meet the actual usage scenarios and needs.
[0066] Therefore, for the sample data set, it is considered that the sample data set includes a labeled sample data set and an unlabeled sample data set.
[0067] The present invention constructs a corresponding classifier for a labeled sample data set, and the classifier can be composed of label information as the classification basis, thereby realizing the classification function in the prior art. Therefore, an initial classifier is constructed based on the mapping relationship between the text and labels in the labeled sample data set and the connection between labels. The initial classifier can be any classifier trained with a labeled sample data set, such as a naive Bayes classifier, a support vector machine classifier, a random forest classifier, etc. The choice of classifier should be based on the specific situation, and the type of classifier that is most suitable for the current task can be inferred based on known information. If there are multiple classifiers that can be used to solve the current task, one of the classifiers can be selected as the initial classifier.
[0068] For the unlabeled sample data set, please use it as the data for verification of classification by the initial classifier, perform label prediction, and determine the label prediction results of the unlabeled sample data set. In this process, the prediction is combined with a semi-supervised learning algorithm, which may include rule-based algorithms, graph-theory-based algorithms, deep learning-based algorithms, etc. Therefore, the present invention uses a constructed classifier to perform label prediction on the sample data in the unlabeled sample data set, and performs label prediction based on the contextual relationship, semantic relationship, and logical structure in the labeled sample data set, taking full account of the intrinsic connection between the labels and the relationship between terms. Finally, the accuracy of the label prediction results of the unlabeled sample data set is checked to improve the credibility of the label prediction. In terms of credibility prediction, the evaluation indicators include accuracy, precision, recall, and F1 value.
[0069] The beneficial effects of the above technical solution are:
[0070] In the present invention, a classifier is constructed according to the mapping relationship between text and labels in a labeled sample data set, as well as the connection between labels, which is beneficial to improving the classification accuracy and performance of the classifier. Label prediction is performed on an unlabeled sample data set by the constructed classifier, and the label prediction results of the unlabeled sample data set are verified, which is beneficial to judging whether the result of label judgment is credible and improving the generalization performance of label prediction.
[0071] Embodiment 2:
[0072] In Example 1, as shown in the attached Figure 2 As shown, the step of obtaining a labeled sample data set, constructing a classifier according to the label information of the labeled sample data set, and determining an initial classifier includes:
[0073] Obtain a sample data set, and classify the sample data set according to whether it has a label or not, and determine the classification result of the sample data set; wherein,
[0074] The sample data set classification results include: a labeled sample data set and an unlabeled sample data set;
[0075] Obtaining text and corresponding labels in the labeled sample data set, and establishing a mapping relationship between the text and the corresponding labels;
[0076] Based on the mapping relationship between text and labels in the labeled sample data set, an initial classifier is constructed.
[0077] The principle of the above technical solution is:
[0078] In the actual implementation process, because the sample data set integrates different types of data, including labeled sample data and unlabeled sample data, and the formats of the data are also different, the existing technology classifies the labels of the sample data in a supervised manner, and directly classifies the labels according to the order in the data set. In this way, the system memory is occupied and the classification efficiency is low. Only the labeled sample data can be classified. When classifying the unlabeled sample data, overfitting is prone to occur, resulting in low credibility of the label classification.
[0079] When the present invention is implemented, the sample data set is first classified according to whether it has a label or not, and is divided into a labeled sample data set and an unlabeled sample data set. A mapping relationship is established between the text and the label in the labeled sample data set, and a classification is established based on the mapping relationship between the text and the label.
[0080] More specifically, for labeled sample datasets, the text and corresponding labels are read and stored in a directed acyclic graph structure. In this structure, each node represents a word, and the color of the node represents its label. For example, if a node is labeled "red", it can point to other nodes with the same label, forming a directed connected structure. For unlabeled sample datasets, all possible word combinations are traversed and each combination is labeled. In this process, existing word embedding models can be used to assist in labeling, such as Word2Vec or GloVe. Once all word combinations are labeled, these combinations can be input into the classifier and classified according to the output of the classifier. Finally, the output of the classifier is compared with the original label to determine the accuracy test result of the calculation result.
[0081] The beneficial effects of the above technical solution are:
[0082] The present invention establishes a mapping relationship between text and labels in a labeled sample data set, and constructs a classifier through the intrinsic mapping relationship between text and labels, which is beneficial to improving the performance of the classifier, and the credibility of using this classifier in an unlabeled sample data set is higher without overfitting.
[0083] Embodiment 3:
[0084] In Example 2, as shown in the attached Figure 3 As shown, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result includes:
[0085] Based on the initial classifier, a mapping relationship between text and labels in the labeled sample data set is obtained, and a mapping label is generated according to the mapping relationship between the text and the label;
[0086] Preprocessing the unlabeled sample data to determine the preprocessing results; wherein the preprocessing content includes: sentence decomposition, part of speech decomposition, stop word removal, and noise cleaning;
[0087] According to the preprocessing result of the unlabeled sample data and the mapping label, label prediction is performed on the preprocessed unlabeled sample data set to determine the label prediction result.
[0088] The principle of the above technical solution is:
[0089] When the present invention is implemented, the data in the unlabeled sample data set is preprocessed, that is, the data in the data set is sentence decomposition, part of speech decomposition, stop word removal, and noise cleaning are performed. The words remaining after the preprocessing will form a dictionary, and the expression of text features is based on this dictionary. For Chinese text, the preprocessing process is more important because Chinese text is generally based on words, and there are no spaces between words. By applying the constructed classifier to the preprocessed unlabeled sample data set, a better classification effect can be obtained.
[0090] In specific implementation, the mapping labels of the present invention can map labels in an existing labeled data set to words in an unlabeled data set. Mapping labels can be generated using the most frequently occurring labels and words in an existing data set. The role of preprocessing is to effectively remove noise and irrelevant information in the text, and help the model better understand the meaning and context of the text. And improve the accuracy of subsequent classification, and prepare for the generation of mapping labels. The semi-supervised learning method can use existing labeled data sets and mapping labels to train the model, and then predict new data. The fully supervised learning method can directly use unlabeled data sets and mapping labels for model training without the need for labeled data sets. When the present invention is implemented on a large scale, these two functions can be combined.
[0091] The beneficial effects of the above technical solution are:
[0092] The present invention preprocesses the data in the unlabeled sample data set, which is beneficial to obtaining the text features of the unlabeled sample data, thereby improving the efficiency and accuracy of label classification.
[0093] Embodiment 4:
[0094] Under Example 1, the embodiment of the present invention provides the method of performing accuracy test on the data label prediction result based on preset evaluation indicators to obtain the accuracy test result, wherein the evaluation indicators include: Hamming loss, sorting loss, coverage efficiency, precision, and error rate.
[0095] The principle of the above technical solution is:
[0096] In a real scenario: generally, evaluation indicators such as precision, accuracy, and recall are used to test the results of text classification. However, in multi-label text classification tasks, since each sample may have multiple irrelevant category labels, the applicability of the evaluation indicators used for text classification is low and the performance of the classifier cannot be fully tested.
[0097] When the present invention is implemented, for multi-label text classification based on "text and labels" in the multi-label classification process, the hierarchical association between category labels is tested using evaluation indicators such as Hamming loss and sorting loss, coverage efficiency, accuracy, and error rate to test the results of label prediction from multiple angles.
[0098] The beneficial effects of the above technical solution are:
[0099] The present invention constructs a classifier based on the hierarchical association between category labels in a labeled sample data set, which can effectively improve the classification level of the classifier. By using Hamming loss, ranking loss, coverage efficiency, accuracy, and error rate as the evaluation criteria of the classifier, the results of label prediction and the construction performance of the classifier can be more comprehensively tested from different angles.
[0100] Embodiment 5:
[0101] In Embodiment 1, the step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier further includes:
[0102] Performing label correlation calculation on the labeled sample data to obtain label correlation calculation results;
[0103] Constructing a label co-occurrence matrix according to the label correlation calculation result and the labeled sample data set to obtain category label distribution information in the labeled sample data set; wherein the label co-occurrence matrix is used to calculate the distribution between labels in the labeled sample data set;
[0104] Performing a correlation analysis on the category label distribution information of the labeled sample data set to obtain the correlation coefficient between the category labels;
[0105] An initial classifier is constructed based on the correlation coefficients between the class labels.
[0106] The principle of the above technical solution is:
[0107] In a practical scenario: when classifying the data labels in the unlabeled sample data set, the correlation between the labels in the labeled sample data set is not considered. Therefore, the unlabeled data cannot be used correctly to improve the performance of the classifier, and overfitting is prone to occur.
[0108] When the present invention is implemented, when constructing a classifier, by constructing a label co-occurrence matrix in a labeled sample data set, the degree of association of the category information in the labeled sample is calculated, the corresponding category label is obtained, and the distribution of different labels in the sample data set is viewed more intuitively. Finally, by normalizing the distribution, the probability of occurrence of labels with high similarity, that is, the co-occurrence frequency, can be obtained. The co-occurrence frequency can be used to obtain the intrinsic connection between the labels in a more detailed manner. The role of the co-occurrence matrix is to represent the frequency of co-occurrence of labels or other label elements in the data.
[0109] In a specific embodiment, constructing a label co-occurrence matrix for the labeled sample data set mainly includes the following steps:
[0110] Step 1: Assume that there is a labeled sample data set represented by T = {Y, Z, M, i}, and the corresponding text sample space is Y = {y1, y2, …y j …, y n}, the label space is Z = {z1,z2,…,z j , …z n}, the set of all possible labels is M = {M1, ..., M m …, M q}, the training set can be represented as {(y1,z1),(y2,z2),…,(y n ,z n )}(y j ∈ i represents the label classification function: i:y j →z j , text sample y j The label set z j is predictable, so there will be label information in the labeled sample data set to construct a co-occurrence matrix N q*n ;
[0111]
[0112] Among them, N q*n represents the co-occurrence matrix constructed based on the labeled sample data set, q represents the number of rows in the co-occurrence matrix, n represents the number of columns corresponding to the co-occurrence matrix, j = 1, 2, ..., n, m = 1, 2, ..., q, fqn If it contains a category label, it takes 0, if it does not contain a category label, it takes 1; M q Represents a label set and a row data set;
[0113] Step 2: The co-occurrence matrix can be used to obtain the distribution of category labels in the labeled sample data. In order to visualize the distribution, it is normalized:
[0114]
[0115] Among them, w j ,w v Respectively represent the matrix N q*n All label information in the jth and vth rows, Represents the matrix N qvn The normalized processing result of all label information in the jth and vth rows, N lj Represents the matrix N q*n The label information of the lth row and jth column in N vj Represents the matrix N q*n The label information of the vth row and jth column in
[0116] Step 3: Calculate w based on the normalized results of all label information in rows j and v j ,w v The co-occurrence frequency of the label M j and M v The co-occurrence frequency of:
[0117]
[0118] Among them, the T j,v Indicates label M j and M v The co-occurrence frequency calculation results are: Represents the co-occurrence matrix N q*n The normalized processing result of all label information in the jth row and vth column is: Represents the co-occurrence matrix N q*n The normalization result of all the data in the j-th row and j-th column is: Represents the co-occurrence matrix N q*n The normalized processing result of all the data in the vth row and vth column in is obtained. By counting the co-occurrence frequencies of all labels, a co-occurrence frequency matrix can be established, and the correlation degree between any two labels can be obtained through the co-occurrence frequency matrix.
[0119] The beneficial effects of the above technical solution are:
[0120] In the present invention, by constructing a label co-occurrence matrix in a labeled sample data set, it is beneficial to obtain the correlation coefficient between labels. The performance of the classifier constructed in this way is higher, and the label prediction accuracy is higher when it is used in an unlabeled sample data set. In addition, by calculating the co-occurrence frequency between labels and generating a co-occurrence frequency matrix, it is beneficial to obtain the degree of association between any two labels in the sample data set, and when performing label prediction on unlabeled sample data, the efficiency and reliability are higher.
[0121] Embodiment 6:
[0122] In Embodiment 1, the step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier further includes:
[0123] Obtain a sample data set, determine whether the sample data set is a linear data set, and when the determination result shows that the sample data set is a linear data set, determine whether the sample data set is in a linearly separable state.
[0124] When the judgment result shows that the sample data set is in a linearly separable state, a hyperplane of the sample data set is obtained; wherein the hyperplane is used to divide the linear space into two non-intersecting parts; according to the hyperplane, a corresponding normal vector, an offset, and a support vector are obtained, and a hyperplane segmentation distance is calculated based on the support vector, and an optimal hyperplane is determined according to the hyperplane segmentation distance; wherein the optimal hyperplane is a hyperplane with the largest corresponding hyperplane segmentation distance;
[0125] When the judgment result shows that the sample data set is in a linearly inseparable state, the sample data set is converted into a high-dimensional space through a kernel function, and the hyperplane of the sample data set in the high-dimensional space is obtained. The corresponding normal vector, offset, and support vector are obtained according to the hyperplane in the high-dimensional space, and the hyperplane segmentation distance is calculated based on the support vector, and the optimal hyperplane is determined according to the hyperplane segmentation distance; wherein the kernel function is used to convert the linearly inseparable state of the data set into a linearly separable state in the high-dimensional space.
[0126] The principle of the above technical solution is:
[0127] When the present invention is implemented, a support vector machine is used to analyze linear problems. When linearly separable, there is a hyperplane that allows the training samples to be completely separated. The normal vector and offset of the hyperplane determine the distance between the hyperplane and the origin. By calculating the distance between the samples in the labeled sample data set and the segmented hyperplane, the optimal hyperplane can be obtained when the corresponding distance sum is maximized. When the sample data set is in a linearly inseparable state, the linearly inseparable state of the input feature space is converted into a linearly separable state of a high-dimensional space by introducing a kernel function. Although it is converted into a high-dimensional space, the complexity of the algorithm does not increase.
[0128] The present invention can process not only linear data sets, but also data sets in a linear inseparable state. This is achieved by introducing a kernel function, which can transform the linear inseparable state in the data set into a linear separable state in a high-dimensional space, thereby enabling better processing of nonlinear data sets. In addition, by calculating the hyperplane segmentation distance and determining the optimal hyperplane, the data set can be divided into two parts in the high-dimensional space, and the segmentation between the two parts is optimal.
[0129] The beneficial effects of the above technical solution are:
[0130] The present invention introduces a support vector machine when training sample data, which is beneficial for solving label classification problems in the fields of small samples, high dimensions, nonlinearity and local minima, and has a wide range of applications and high reliability.
[0131] Embodiment 7:
[0132] In Embodiment 1, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes:
[0133] Performing an integrated analysis on the category labels according to the labeled sample data set to obtain the co-occurrence frequency corresponding to the category labels; wherein the co-occurrence frequency is used to calculate the correlation coefficient between the labels;
[0134] According to the co-occurrence frequency of the category labels in the labeled sample data, the label information of the neighboring text samples is counted to determine the label information statistics of the neighboring text samples; wherein the neighboring text samples are text samples whose label co-occurrence frequency in the sample data set is within a preset threshold range;
[0135] Based on the co-occurrence frequency of the category labels and the statistical results of the label information of the neighboring text samples, a label prediction model is established, the unlabeled sample data set is used as the input item of the label prediction model, the label prediction result is output, and the label prediction sequence of the unlabeled sample data is obtained; wherein the label prediction model is used to output the corresponding label category information according to the input sample data set.
[0136] The principle of the above technical solution is:
[0137] When the present invention is implemented, a K-nearest neighbor algorithm is used to perform label prediction on a labeled sample data set, with the purpose of defining a nearest neighbor number K and an unlabeled sample data set to ultimately obtain a label prediction result of the unlabeled sample data set, establish a label prediction model through the co-occurrence frequency of category labels and the label information statistics of neighboring text samples, use the unlabeled sample data set as the input of the label prediction model, output the label prediction result, and obtain the label prediction sequence of the unlabeled sample data.
[0138] In the specific implementation, the initial classifier can be any classifier suitable for semi-supervised learning, such as support vector machine, random forest, etc. Through this classifier, a set of label prediction results can be obtained. The co-occurrence frequency can be used to calculate the correlation coefficient between labels, because the correlation between two labels can be measured by their co-occurrence frequency in the text. Finally, the established label prediction model can realize the output of label prediction results, thereby determining the label prediction sequence.
[0139] The beneficial effects of the above technical solution are:
[0140] The present invention adopts a multi-label K-nearest neighbor algorithm to construct a label prediction model, and performs iterative prediction on an unlabeled sample data set, which is beneficial to improving the credibility of the label prediction model and increasing the stability of the parameters in the model.
[0141] Embodiment 8:
[0142] In Embodiment 1, the step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier further includes:
[0143] Randomly sampling the labeled sample data set to divide it into a first labeled sample data set and a second labeled sample data set;
[0144] Constructing a first initial classifier based on the first labeled sample data set, and constructing a second initial classifier based on the second labeled sample data set; wherein the first initial classifier is trained and obtained using the IMML-KNN algorithm, and the second initial classifier is trained and obtained using the MLPLSA-KNN algorithm;
[0145] Traversing the first initial classifier and the second initial classifier respectively on the unlabeled sample data set to obtain a first traversal result and a second traversal result;
[0146] Calculate the confidence of the first traversal result and the second traversal result, select the ones with the highest confidence as pseudo labels, save the pseudo labels corresponding to the first traversal result in a first set, and save the pseudo labels corresponding to the second traversal result in a second set;
[0147] Adding the first set to a first labeled sample data set as an updated first labeled sample data set, and adding the second set to a second labeled sample data set as an updated second labeled sample data set;
[0148] The updated first labeled sample data set and the updated second labeled sample data set are respectively updated for the first initial classifier and the second initial classifier to obtain the updated first initial classifier and the updated second initial classifier, accuracy calculation is performed for the updated first initial classifier and the updated second initial classifier to obtain the accuracy calculation result, and according to the accuracy calculation result, the one with a higher corresponding calculation result is used as the target classifier, and the unlabeled sample data set is classified based on the target classifier to obtain the label classification result.
[0149] The principle of the above technical solution is:
[0150] When the sample data set contains only a small amount of labeled data and a large amount of unlabeled sample data, the construction of the trainer can only rely on a small amount of labeled data, resulting in low classification performance of the classifier. In addition, in real tasks, the data sets often cannot meet the requirements of being completely independent of each other, so the constructed classifier is prone to redundancy.
[0151] When the present invention is implemented, two classifiers are trained on the same attribute set, and the sample space is divided into several equal classes in each classifier. In the collaborative training process, statistical techniques are used to estimate the confidence of the classifier in marking unlabeled samples, and then sample labels with high confidence are selected to mark the unlabeled samples. The samples and the marking results of the samples are put into another classifier training set to train the classifier. This process is cross-cut in the two classifiers until a certain stop condition is met.
[0152] The beneficial effects of the above technical solution are:
[0153] In the present invention, different algorithms are used to construct classifiers, and the classifier with higher accuracy is used as the target classifier to perform label prediction on the unlabeled sample data set. This method is beneficial to improving the performance of the classifier and also beneficial to improving the efficiency and accuracy of label prediction of unlabeled sample data information.
[0154] Embodiment 9:
[0155] In Embodiment 1, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: performing entity type prediction on a knowledge graph based on the label prediction result, wherein the entity type prediction step of the knowledge graph includes:
[0156] Step 1: Obtain the knowledge graph input by the user, parse the knowledge graph, obtain the knowledge graph information, generate the corresponding information description according to the knowledge graph information, and determine the knowledge graph information description result; wherein the knowledge graph information includes: structured information and unstructured information;
[0157] Step 2: According to the description result of the knowledge graph information, determine each entity in the knowledge graph, and extract features for the text information and link information of each entity to generate a feature data set; wherein the feature data set includes: a labeled sample data set and an unlabeled sample data set, the labeled sample data set is used as training data, and the unlabeled sample data set is used as prediction data;
[0158] Step 3: Based on the description result of the knowledge graph information, perform model training on the training data, determine the training model, input the predicted data as an input item into the training model, obtain the predicted type of the entity, generate type knowledge based on each entity in the knowledge graph and the corresponding predicted type, and determine the prediction result of the entity type in the knowledge graph; wherein the training model predicts the entity type in the knowledge graph;
[0159] When the present invention is implemented, the multi-label classification method of semi-supervised learning is applied to the knowledge graph to predict the type of entity, and the information in the knowledge graph is described. Then, for each entity in the knowledge graph, the text information and link information of the entity are extracted. Finally, features are extracted from the text information and link information respectively to generate a data set for the entity set. In the data set, the labeled sample data set is used as training data, and the unlabeled sample data set is used as data to be predicted. By training the training model, the entity type in the knowledge graph is finally predicted.
[0160] The principle of the above technical solution is:
[0161] The process of performing model training on the training data and determining the training model includes:
[0162] Step 1: Assume that the sample data set is l = [l1,l2,…,l i ,…l m ], where i = 1, 2, ..., m, where m represents the number of samples in the data set, and each l i Each represents a sample. By segmenting each word in the text, obtaining the semantic features of the sub-text, and calculating the deep features through the fully connected layer, multi-label text classification is finally achieved. Assuming that the sample data set passes through the first layer of the network, the parameters of the neuron are X, b, and the result of the model training is expressed as:
[0163] R=Φ(X·l i +b)
[0164] Among them, R represents the result of model training, X and b are the variable parameters and constant parameters of the neural network respectively, and l i Represents any sample in the sample data set, Φ() represents the neuron activation function. When setting the activation function, the linear rectifier function ReLU is selected, ReLU = max(0, x). The reason for selecting ReLU as the activation function is that it can alleviate the problem of gradient diffusion in the training process of the neural network, and can also improve the calculation speed and convergence speed;
[0165] Step 2: After the first layer of training, when the entity type in the knowledge graph is input as an input item, the output result is the prediction result of the knowledge type and entity type:
[0166]
[0167] Among them, X, b are the variable parameters and constant parameters of the neural network respectively, σ represents the prediction result of the knowledge type, ω 2 Represents the prediction result of entity type, n represents the number of samples;
[0168] The beneficial effects of the above technical solution are as follows: in the present invention, by applying the multi-label classification method to the prediction of entity types in the knowledge graph, the use scenarios of semi-supervised learning are increased, and the field of multi-label classification is made more diversified. The use of ReLU as the activation function is not only conducive to alleviating the problem of gradient diffusion in the training process of the neural network, but also conducive to improving the calculation speed and convergence speed.
[0169] Embodiment 10:
[0170] In Embodiment 1, the step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: detecting data faults in the sensor network according to the label prediction result, and determining a fault detection result; wherein the execution step of the fault detection includes:
[0171] Acquire a data stream in the sensor network, segment the data stream to obtain at least two segments of data stream, and perform data stream category label prediction based on the label prediction result for each segment of data stream;
[0172] Compare the data stream label prediction result with a preset threshold range. When the data stream label prediction result is not within the threshold range, determine that a fault occurs during the construction of the initial classifier, and determine the type of fault; wherein the fault type includes: outlier fault, spike fault, and high noise fault;
[0173] In a real scenario: In a sensor network, there are many data failures. Generally, the data is tracked and analyzed, and the data is compared with the preset threshold to determine the abnormal data, so as to locate and detect the failure. In this way, when the amount of data involved is large, the efficiency of fault detection and location is low;
[0174] When the present invention is implemented, a multi-label classification method is applied to fault detection in a sensor network. Since there are many types of faults and there are many internal connections between different faults, a multi-label classification model is used in the present invention to detect sensor network data faults. Label prediction is performed by analyzing the internal connections between different faults, and the fault type can be determined based on the results of label prediction.
[0175] The beneficial effects of the above technical solution are:
[0176] The present invention applies a multi-label classification method to fault detection in a sensor network, which is beneficial for detecting various types of faults in the sensor network and can improve the performance index of the classifier.
[0177] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0178] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0179] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0180] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0181] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for automatic classification of data labels based on semi-supervised learning, characterized in that: include: Obtain a labeled sample data set, and construct a classifier according to the label information of the labeled sample data set to determine an initial classifier; Obtain an unlabeled sample data set, perform label prediction on the unlabeled sample data set based on the initial classifier, and obtain a label prediction result; The data label prediction result is subjected to an accuracy test based on a preset evaluation index to obtain an accuracy test result.
2. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining a labeled sample data set, constructing a classifier according to label information of the labeled sample data set, and determining an initial classifier includes: Obtain a sample data set, and classify the sample data set according to whether it has a label or not, and determine the classification result of the sample data set; wherein, The sample data set classification results include: a labeled sample data set and an unlabeled sample data set; Obtaining text and corresponding labels in the labeled sample data set, and establishing a mapping relationship between the text and the corresponding labels; An initial classifier is constructed based on the mapping relationship between text and labels in the labeled sample data set.
3. The method for automatic classification of data labels based on semi-supervised learning according to claim 2, characterized in that: The step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result includes: Based on the initial classifier, a mapping relationship between text and labels in the labeled sample data set is obtained, and a mapping label is generated according to the mapping relationship between the text and the label; Preprocessing the unlabeled sample data to determine the preprocessing results; wherein the preprocessing content includes: sentence decomposition, part of speech decomposition, stop word removal, and noise cleaning; According to the preprocessing result of the unlabeled sample data and the mapping label, label prediction is performed on the preprocessed unlabeled sample data set to determine the label prediction result.
4. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The data label prediction result is subjected to an accuracy test based on preset evaluation indicators to obtain an accuracy test result, wherein the evaluation indicators include: Hamming loss, sorting loss, coverage efficiency, precision, and error rate.
5. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining a labeled sample data set, constructing a classifier according to the label information of the labeled sample data set, and determining an initial classifier further includes: Performing label correlation calculation on the labeled sample data to obtain label correlation calculation results; Constructing a label co-occurrence matrix according to the label correlation calculation result and the labeled sample data set to obtain category label distribution information in the labeled sample data set; wherein the label co-occurrence matrix is used to calculate the distribution between labels in the labeled sample data set; Performing a correlation analysis on the category label distribution information of the labeled sample data set to obtain the correlation coefficient between the category labels; An initial classifier is constructed based on the correlation coefficients between the class labels.
6. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining a labeled sample data set, constructing a classifier according to the label information of the labeled sample data set, and determining an initial classifier further includes: Obtain a sample data set, determine whether the sample data set is a linear data set, and when the determination result shows that the sample data set is a linear data set, determine whether the sample data set is in a linearly separable state. When the judgment result shows that the sample data set is in a linearly separable state, a hyperplane of the sample data set is obtained; wherein the hyperplane is used to divide the linear space into two non-intersecting parts; according to the hyperplane, a corresponding normal vector, an offset, and a support vector are obtained, and a hyperplane segmentation distance is calculated based on the support vector, and an optimal hyperplane is determined according to the hyperplane segmentation distance; wherein the optimal hyperplane is a hyperplane with the largest corresponding hyperplane segmentation distance; When the judgment result shows that the sample data set is in a linearly inseparable state, the sample data set is converted into a high-dimensional space through a kernel function, and the hyperplane of the sample data set in the high-dimensional space is obtained. The corresponding normal vector, offset, and support vector are obtained according to the hyperplane in the high-dimensional space, and the hyperplane segmentation distance is calculated based on the support vector, and the optimal hyperplane is determined according to the hyperplane segmentation distance; wherein the kernel function is used to convert the linearly inseparable state of the data set into a linearly separable state in the high-dimensional space.
7. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: Performing an integrated analysis on the category labels according to the labeled sample data set to obtain the co-occurrence frequency corresponding to the category labels; wherein the co-occurrence frequency is used to calculate the correlation coefficient between the labels; According to the co-occurrence frequency of the category labels in the labeled sample data, the label information of the neighboring text samples is counted to determine the label information statistics of the neighboring text samples; wherein the neighboring text samples are text samples whose label co-occurrence frequency in the sample data set is within a preset threshold range; Based on the co-occurrence frequency of the category labels and the statistical results of the label information of the neighboring text samples, a label prediction model is established, the unlabeled sample data set is used as the input item of the label prediction model, the label prediction result is output, and the label prediction sequence of the unlabeled sample data is obtained; wherein the label prediction model is used to output the corresponding label category information according to the input sample data set.
8. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining a labeled sample data set, constructing a classifier according to the label information of the labeled sample data set, and determining an initial classifier further includes: Randomly sampling the labeled sample data set to divide it into a first labeled sample data set and a second labeled sample data set; Constructing a first initial classifier based on the first labeled sample data set, and constructing a second initial classifier based on the second labeled sample data set; wherein the first initial classifier is trained and obtained using the IMML-KNN algorithm, and the second initial classifier is trained and obtained using the MLPLSA-KNN algorithm; Traversing the first initial classifier and the second initial classifier respectively on the unlabeled sample data set to obtain a first traversal result and a second traversal result; Calculate the confidence of the first traversal result and the second traversal result, select the ones with the highest confidence as pseudo labels, save the pseudo labels corresponding to the first traversal result in a first set, and save the pseudo labels corresponding to the second traversal result in a second set; Adding the first set to a first labeled sample data set as an updated first labeled sample data set, and adding the second set to a second labeled sample data set as an updated second labeled sample data set; The updated first labeled sample data set and the updated second labeled sample data set are respectively updated for the first initial classifier and the second initial classifier to obtain the updated first initial classifier and the updated second initial classifier, accuracy calculation is performed for the updated first initial classifier and the updated second initial classifier to obtain the accuracy calculation result, and according to the accuracy calculation result, the one with a higher corresponding calculation result is used as the target classifier, and the unlabeled sample data set is classified based on the target classifier to obtain the label classification result.
9. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: performing entity type prediction on a knowledge graph based on the label prediction result, wherein the entity type prediction step of the knowledge graph includes: Step 1: Obtain the knowledge graph input by the user, parse the knowledge graph, obtain the knowledge graph information, generate the corresponding information description according to the knowledge graph information, and determine the knowledge graph information description result; wherein the knowledge graph information includes: structured information and unstructured information; Step 2: According to the description result of the knowledge graph information, determine each entity in the knowledge graph, and extract features for the text information and link information of each entity to generate a feature data set; wherein the feature data set includes: a labeled sample data set and an unlabeled sample data set, the labeled sample data set is used as training data, and the unlabeled sample data set is used as prediction data; Step 3: Based on the results of the knowledge graph information description, perform model training on the training data, determine the training model, input the predicted data as input to the training model, obtain the predicted type of the entity, generate type knowledge based on each entity in the knowledge graph and the corresponding predicted type, and determine the predicted result of the entity type in the knowledge graph; wherein the training model makes predictions for the entity types in the knowledge graph.
10. The method for automatic classification of data labels based on semi-supervised learning according to claim 1, characterized in that: The step of obtaining an unlabeled sample data set, performing label prediction on the unlabeled sample data set based on the initial classifier, and obtaining a label prediction result further includes: detecting data faults in the sensor network according to the label prediction result, and determining a fault detection result; wherein the execution step of the fault detection includes: Acquire a data stream in the sensor network, segment the data stream to obtain at least two segments of data stream, and perform data stream category label prediction based on the label prediction result for each segment of data stream; The data stream label prediction result is compared with a preset threshold range. When the data stream label prediction result is not within the threshold range, it is determined that a fault has occurred in the process of building the initial classifier, and the fault type is determined; wherein the fault types include: outlier fault, spike fault, and high noise fault.
Citation Information
Patent Citations
Semi-supervised classification algorithm based on ACGAN image
CN110647927A
Text label extraction method and device, equipment and storage medium
CN112699232A
Text classification data processing method and device, storage medium and program product
CN113722493A
Small sample text classification method and system based on semi-supervised learning
CN114036947A
Text classification method and device, computer equipment and storage medium
CN114117048A