Blockchain Abnormal Transaction Detection Method Based on Semi-Supervised Learning

Through a semi-supervised learning method, combined with integrated learning and improved Tri-Training, the feature selection and data imbalance in blockchain abnormal transaction detection are solved, the detection accuracy and model generalization capabilities are improved, and an efficient Multi-Training model is built.

CN116644283BActive Publication Date: 2025-07-25NORTHEASTERN UNIV CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310605595.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-26
Publication Date
2025-07-25
Estimated Expiration
2043-05-26

AI Technical Summary

Technical Problem

In the detection of blockchain abnormal transactions, the existing technology has problems such as feature selection that does not take into account the lack of fitting capabilities of subsequent learners, underfitting data classification and noise impacts in reverse optimization, especially in the case of unbalanced data sets.

Method used

Using a semi-supervised learning method, combining feature selection of integrated learning and a classified data preprocessing, high confidence pseudo-label data is obtained through the improved Tri-Training method, and model training is used for model training to build a Multi-Training model.

Benefits of technology

It improves the accuracy and generalization ability of blockchain abnormal transaction detection, reduces redundant feature interference, improves the training effect of the model, and reduces the error classification rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116644283B_ABST
    Figure CN116644283B_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting abnormal blockchain transactions based on semi-supervised learning, which relates to the technical field of blockchain. In view of the diversity of characteristics of blockchain transaction nodes, a feature selection method based on ensemble learning is proposed; in view of the phenomenon of data set imbalance in blockchain transactions, a data preprocessing method based on one-class classification is proposed; in view of the problem of reverse optimization caused by prediction deviation in the traditional Tri-Training method, the present invention proposes a Multi-Training method to obtain pseudo-label data with high confidence for training the blockchain abnormal transaction detection model. The present invention makes full use of labeled data and unlabeled data, enhances the detection accuracy, reduces the interference of redundant features, effectively solves the problem of underfitting of negative example samples caused by data imbalance, and improves the generalization ability of the training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of blockchain, and in particular to a method for detecting abnormal transactions in a blockchain based on semi-supervised learning. Background Art

[0002] In recent years, cryptocurrency based on blockchain has received wide attention, and a large number of illegal and criminal incidents have occurred. However, the global blockchain abnormal transaction detection field is still in its infancy, and it is not perfect enough in terms of integrity, systematicness, practicality, etc. Blockchain abnormal transaction detection can be regarded as a classification task for nodes, finding abnormal nodes that are different from normal transaction nodes, which can effectively detect the attack behavior of lawbreakers, and then achieve multi-level security access control to effectively ensure the security of the blockchain system.

[0003] The product research and development of blockchain technology has made remarkable progress, but at the same time, it has also exposed risks such as abnormal transactions, fraud, false information induction, and malicious price manipulation. Therefore, it is necessary to accurately identify abnormal transactions on the blockchain, extract the original features from the intricate blockchain transactions, and accurately and efficiently classify a small number of illegal transactions from the continuously growing massive transactions. Studying the detection of blockchain abnormal transactions can effectively detect the attack behavior of lawbreakers, and give early warnings of possible attacks or illegal transactions, effectively ensuring the security of the blockchain system, which has important theoretical significance and practical value. Strengthen the theoretical analysis of blockchain abnormal events and enhance the research of abnormal detection technology, break through the bottleneck of blockchain abnormal transaction detection, identify user identities through abnormal transaction detection, achieve multi-level security access control, and then provide security support for the development and full implementation of blockchain applications.

[0004] Chinese Patent CN202210183812.9[P] provides a method for detecting blockchain anomalies based on PCA and RF, including calling the PCA model to reduce the dimension of the original transaction data to obtain the data to be detected; calling the Bayesian optimization model to optimize and train the data to be detected to obtain the optimal hyperparameters of the random forest model; training the random forest model based on the obtained hyperparameters to obtain a blockchain anomaly detection model; calling the pre-trained random forest model to calculate the data to be detected to obtain the anomaly detection result corresponding to the original transaction data. By reducing the dimension of the original transaction data through the PCA model, the interference of redundant features can be reduced and the anomaly detection performance can be improved; by implementing intelligent optimization of the random forest hyperparameters through the Bayesian optimization model, the classification performance can be improved and the influence of the extremely unbalanced positive and negative samples of blockchain transaction data can be eliminated.

[0005] Traditional research on blockchain abnormal transaction detection mainly focuses on supervised learning and unsupervised learning. Supervised abnormal detection methods require a labeled training set to train a classifier to identify blockchain abnormal transactions. Unsupervised abnormal transaction detection methods do not require a pre-labeled training set and have received more attention and expectations. In existing research, it can be seen that supervised learning-based abnormal detection methods need to label all sample data sets, and a large number of unlabeled samples lead to poor generalization performance of the learner. In blockchain, abnormal transactions often hide among a large number of normal transactions, resulting in an unclear boundary between legal and illegal transactions and a small difference in their eigenvalue, so there is a high misclassification rate in unsupervised method detection.

[0006] The current methods for blockchain abnormal transaction detection mainly have the following defects:

[0007] (1) In the feature selection stage, the learner to be used later is not considered, weakening the fitting ability of the learner;

[0008] (2) There is a certain risk of underfitting when classifying data;

[0009] (3) The impact of noise and misclassification on the reverse optimization of the model is ignored. Summary of the Invention

[0010] The technical problem to be solved by the present invention is to provide a blockchain abnormal transaction detection method based on semi-supervised learning in view of the above-mentioned deficiencies of the prior art. In view of the diversity of blockchain transaction node features, a feature selection method based on ensemble learning is proposed; in view of the imbalance phenomenon in the blockchain transaction data set, a data preprocessing method based on one-class classification is proposed; in view of the problem of reverse optimization caused by prediction deviation in the traditional Tri-Training method, the present invention proposes a Multi-Training method to obtain pseudo-label data with high confidence for training the blockchain abnormal transaction detection model.

[0011] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0012] A blockchain abnormal transaction detection method based on semi-supervised learning, comprising the following steps:

[0013] Step 1: Feature engineering and feature selection; obtain the initial blockchain transaction data set and extract the effective features of the data set;

[0014] Step 2: Data preprocessing; preprocess the initial data set with unbalanced numbers of positive and negative examples;

[0015] Step 3: Model training; classify a large amount of unlabeled data using the improved Tri-Training method, obtain pseudo-labeled data with high confidence based on DBSCAN, and iteratively optimize the base classifier using the training set data and the pseudo-labeled data;

[0016] Step 4: Input the target monitoring data into the abnormal transaction detection model to obtain the detection result.

[0017] The beneficial effects of adopting the above technical solutions are as follows: The blockchain abnormal transaction detection method based on semi-supervised learning provided by the present invention proposes an innovative semi-supervised model, which makes full use of labeled data and unlabeled data to enhance the detection accuracy; combines the advantages of the filtering method and the wrapper method to judge whether a variable should be selected from the situation of a single variable itself and the relationship between multiple variables, and proposes a feature selection method based on ensemble learning to extract effective features of the data set and reduce the interference of redundant features; proposes a one-class classification method to process the imbalanced data set, effectively solving the problem of underfitting of negative example samples caused by data imbalance; proposes a Multi-Training method to obtain pseudo-labeled data with high confidence and improve the generalization ability of the training model. Brief Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of the blockchain abnormal transaction detection method based on semi-supervised learning provided by the embodiment of the present invention;

[0019] Figure 2 It is a schematic flowchart of the feature selection method based on ensemble learning provided by the embodiment of the present invention;

[0020] Figure 3 It is a schematic diagram of the modeling process of the blockchain abnormal transaction detection based on semi-supervised learning provided by the embodiment of the present invention;

[0021] Figure 4 It is a schematic diagram of the system structure of the blockchain abnormal transaction detection based on semi-supervised learning provided by the embodiment of the present invention. Detailed Embodiments

[0022] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0023] As Figure 1 shown, the method of this embodiment is described as follows.

[0024] Step 1: Feature engineering and feature selection. Obtain the initial blockchain transaction data set, extract the effective features of the data set, and combine the filtering method ANOVA and the wrapper method RFECV to extract the features for effectively detecting blockchain abnormal transactions through the situation of a single variable itself and the relationship between multiple variables.

[0025] The characteristics of the initial data set of blockchain transactions include local characteristics and aggregated characteristics. The local information of nodes, denoted as local characteristics, such as 94 characteristics like time step, transaction fee, number of inputs / outputs, etc.; 72 characteristics such as maximum value, minimum value, standard deviation, and correlation coefficient statistically aggregated from one-hop forward / backward transaction information from the central node, marked as aggregated characteristics. Specifically, an initial data set constructed from 166 characteristics such as the waiting time of the maximum output cost of a transaction, the variance of all output amounts of a transaction, the mean of the time interval of input transactions, the median of the time interval of output transactions, the number of output transactions of a transaction, and the number of input transactions of a transaction is used, and the data is labeled according to the types of legal entities (such as exchanges, wallet providers, legal services, etc.) and illegal entities (such as fraud, malware, ransomware, etc.) through a heuristic reasoning process.

[0026] Weights are assigned to each dimension of the characteristics, and the characteristics are sorted according to the weights. A base classifier is used for multiple rounds of training. After each round of training, several low-weight characteristics are eliminated. The next round of training is based on the new feature set, and the Ranking of each characteristic is obtained accordingly. Based on the Ranking, feature subsets are sequentially selected for model training and cross-validation, and finally the feature subset with the highest average score is selected. As Figure 2 shown, the flowchart of the feature selection method based on ensemble learning provided in this embodiment is as follows:

[0027] For the initial feature set F = {F1, F2, F3,..., F N-1 , F N}, ANOVA is used to generate a preliminary screening feature set F' = {F'1, F'2, F'3,..., F' m-1 , F' m}, and the purpose is to test whether there are significant differences in the means under different groups. First, calculate the means, the sample mean of the i-th population and the total mean, n i is the number of sample observations of the i-th level population. Then calculate the sum of squared errors. The sum of squared errors between the group means and the total mean reflects the degree of difference between the sample means of each level population, representing the influence brought by the differences in the theoretical means of each level of factor A, denoted as the sum of squares between groups SSA, and the formula for SSA is as follows:

[0028]

[0029] where k is the number of levels of the control variable, k is 2 in this embodiment, including normal transactions and abnormal transactions; n i is the number of sample observations of the i-th population, n1 is the number of sample observations of normal transactions, and n2 is the number of sample observations of abnormal transactions in this embodiment; is the average value of each group of samples; is the average value of the overall samples.

[0030] The sum of the squares of the errors between each sample data in each group and its group average value reflects the dispersion of each observed value of each sample, represents the influence of random errors, and is denoted as the within-group sum of squares SSE. The formula for SSE is as follows:

[0031]

[0032] where, x ij is the j-th sample value at the i-th level of the control variable.

[0033] The sum of the squares of the errors between all observed values and the total average value reflects the dispersion degree of all observed values, and is denoted as the total sum of squares SST. The formula for SST is as follows:

[0034]

[0035] The mean square between groups MSA and the mean square within groups MSE are obtained by dividing the sum of squares of errors by their respective degrees of freedom. The ratio of MSA / MSE forms an F-distribution:

[0036]

[0037] where, n is the number of samples. As a statistical test statistic, F also has a corresponding probability distribution, that is, it satisfies the probability distribution of F(k - 1, n - k). The larger the F value, the larger SSA is compared to SSE, that is, the greater the difference between groups, the more likely the data in different groups are sampled from different populations, and the lower the possibility (p-value) of the null hypothesis holding. Conversely, the higher the possibility of the null hypothesis holding. Therefore, select the initial screening feature set F' = {F'1, F'2, F'3,..., F' m-1 , F' m}.

[0038] The recursive feature elimination method RFE uses a base model for multiple rounds of training. After each round of training, several features with low weights are eliminated, and then the next round of training is carried out based on the new feature set. RFECV is RFE + CV (cross-validation). Its operation mechanism is: first use REF to obtain the ranking of each feature, and then based on the ranking, sequentially select feature subsets with the number of features of [min_features_to_select, len(feature)] for model training and cross-validation, and finally select the feature subset with the highest average score.

[0039] Step 2: Data preprocessing. The number of abnormal transactions and normal transactions is extremely unbalanced. Preprocess the initial data set with unbalanced positive and negative example numbers so that it is divided into data subsets with uniform positive and negative sample numbers.

[0040] The data is divided into multiple subsets and added to the base classifier for training. For the positive class samples N, a subset N is randomly sampled from them i , and it is added to the negative class samples P. That is, each classified subset contains all abnormal transactions and some normal transactions.

[0041] Step 3: Model training. Use the improved Tri-Training method to classify a large amount of unlabeled data, obtain pseudo-labeled data with high confidence based on DBSCAN, and use the training set data and pseudo-labeled data to iteratively optimize the base classifier.

[0042] Train i base classifiers respectively based on the i labeled data training sets preprocessed in Step 2, and predict the unlabeled data based on this respectively.

[0043] The Tri-training algorithm first performs repeatable sampling on the labeled dataset to obtain three different labeled data training sets. Train three base classifiers respectively based on these three labeled data training sets, and predict the unlabeled data based on this respectively. Obtain the prediction results of the unlabeled data based on the principle of majority vote, and add the unlabeled data values and the predicted results to the labeled data training set, and train the base classifier again. This has significantly improved the effects of model training and model optimization. However, the Tri-training algorithm also has defects. Generally, the algorithm needs to calculate and estimate the data confidence. The Tri-training algorithm implicitly analyzes the training data based on the majority vote mechanism. Although the calculation amount is reduced, it is not accurate enough and may get wrong classifications, which has a reverse optimization effect on the model. To increase the detection accuracy, this embodiment improves the classic Tri-Training algorithm. As Figure 3 shown, this embodiment provides a schematic diagram of the blockchain abnormal transaction detection modeling process based on semi-supervised learning, and the specific implementation process is as follows:

[0044] Step 3.1: Obtain pseudo-labeled data based on KNN. Input the labeled dataset L and the unlabeled dataset U, and use the base classifier H i to iteratively train the data subset X i , predict U, and obtain the prediction result i_pred;

[0045] Base classifier H1, base classifier H2,..., base classifier H i-1 、base classifier H i 、base classifier H i+1, … make predictions on U to obtain prediction results 1_pred, 2_pred, …, i-1_pred, i_pred, i+1_pred, ...

[0046] Step 3.2: Select the datasets with the same predicted labels in U, and set the neighborhood radius Eps and the threshold MinPts for the number of data objects in the neighborhood.

[0047] Step 3.3: Arbitrarily select a data object point p from the dataset. If, for the parameters Eps and MinPts, the selected data object point p is a core point, then find all the data object points that are density-reachable from the data object point p to form a cluster; if the selected data object point p is a border point, then select another data object point.

[0048] Step 3.4: Repeat Step 3.2 and Step 3.3 until all data object points are processed, and finally output the density-connected clusters.

[0049] Step 3.5: Use the training set data and the pseudo-label data to iteratively optimize the base classifier. If the samples are distributed within the same cluster, then add the unlabeled data values and the predicted results to the training set to iteratively optimize the base classifier.

[0050] In the embodiment of the present invention, the base classifier for the blockchain abnormal transaction detection modeling process based on semi-supervised learning adopts the KNN algorithm, and the KNN algorithm is a method that can better balance the accuracy and the memory occupation during the training process. The specific implementation process is as follows:

[0051] Input the training dataset L = {(x1, y1), (x2, y2),...(x N , y N )}, where, is the feature vector of the n-dimensional instance; y i ∈ Y = {c1, c2,..., c k} is the category of the instance; i = 1, 2,..., N.

[0052] According to the given distance metric, find the k nearest points to x in the training set L, and denote the neighborhood of x covering these k points as N k (x).

[0053] The specific distance metric uses the calculation formula of the Manhattan distance:

[0054] In N k (x), determine the category y of x according to the classification decision rule (majority voting):

[0055]

[0056] In the above formula, I is an indicator function:

[0057] In this embodiment, in the blockchain abnormal transaction detection modeling process based on semi-supervised learning, steps 3.2 to 3.4 are to obtain pseudo-label data with high confidence based on DBSCAN. The specific implementation process is as follows:

[0058] Select a data set with the same predicted label in U, set the neighborhood radius Eps and the threshold MinPts for the number of data objects in the neighborhood. Arbitrarily select a data object point p from the data set. If for the parameters Eps and MinPts, the selected data object point p is a core point, then find all the data object points that are density-reachable from p to form a cluster. If the selected data object point p is a border point, then select another data object point. Repeat this process until all data object points are processed, and finally output the density-connected clusters. If the samples are distributed in the same cluster, add the unlabeled data values and the predicted results to the labeled data training set to train the base classifier again, and perform iterative optimization on the base classifier. Finally, integrate it into a Multi-Training classifier.

[0059] Step 4: Input the target monitoring data into the abnormal transaction detection model to obtain the detection result.

[0060] Through the blockchain abnormal transaction detection method based on semi-supervised learning in this embodiment, all the features of the data set are input into the feature selection model, and the model is modeled through ensemble learning to extract features that can effectively detect blockchain abnormal transactions. The number of abnormal transactions and normal transactions is extremely unbalanced. The initial data set with unbalanced positive and negative examples is preprocessed to divide it into data subsets with uniform numbers of positive and negative samples. The classified data subsets are added to the base classifier for training, and the samples predicted by the base classifier with the same label are trained again to reduce the misclassification rate. The base classifier is continuously iteratively optimized, and the above models are integrated and adjusted to construct a Multi-Training model that can effectively detect blockchain abnormal transactions, such as Figure 4 the system shown.

[0061] A blockchain abnormal transaction detection method provided by the present invention includes: obtaining an initial blockchain transaction data set; performing feature selection based on ensemble learning; preprocessing an unbalanced data set based on one-class classification; and proposing a Multi-Training method to obtain pseudo-label data with high confidence for training a blockchain abnormal transaction detection model. It is consistent with the current goals of detecting blockchain abnormal transactions in different ways, but the research of the present invention is very different from the existing research results:

[0062] (1) Most of the current research methods for detecting abnormal blockchain transactions are based on supervised or unsupervised methods. The present invention researches an innovative semi-supervised model, which makes full use of labeled data and unlabeled data to enhance the detection accuracy.

[0063] (2) Considering the diversity of the characteristics of blockchain transaction data, the correlation between features also plays a crucial role in the training of the model. The current research on feature selection is mainly based on the selection of single features. This project combines the advantages of the filtering method and the wrapper method to judge whether a variable should be selected from the situation of a single variable itself and the relationship between multiple variables.

[0064] (3) For the preprocessing of the dataset, the current research mainly uses oversampling or undersampling to process the imbalanced dataset. Based on the one-class classification method, the data is divided into multiple subsets and added to the base classifier for training, effectively improving the training accuracy of the model.

[0065] (4) The present invention researches the comprehensive utilization of labeled data and unlabeled data, effectively combines the two, iteratively optimizes the base classifier, and finally integrates it into the Multi-Training model, which is significantly superior to the current state-of-the-art model in terms of detection accuracy.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A blockchain abnormal transaction detection method based on semi-supervised learning, characterized in that: It includes the following steps: Step 1: Feature engineering and feature selection; Obtain the initial blockchain transaction dataset, and extract the effective features of the dataset; Combine the filtering method ANOVA and the wrapper method RFECV, and extract the features for effectively detecting blockchain abnormal transactions through the situation of single variables themselves and the relationships between multiple variables. The specific method is as follows: For the initial feature set F = {F1, F2, F3,..., F N-1 , F N}, the filtering method ANOVA is used to generate the preliminary screening feature set F' = {F'1, F'2, F'3,..., F' m-1 , F' m}; then the wrapper method RFECV is used to finally select the feature subset with the highest average score; The specific method for generating the initial screening feature set using the filtering method ANOVA is as follows: First, calculate the means, the sample mean of the \(i\)-th population and the overall mean, \(n\) i is the number of sample observations of the \(i\)-th level population, and then calculate the sum of squared errors; The sum of squared errors of the average values of each group and the total average value reflects the degree of difference between the sample means of each level of the population, representing the influence brought by the differences in the theoretical average values of each level of factor A, denoted as the sum of squares between groups SSA. The formula for SSA is as follows: where k is the number of levels of the control variable; is the mean of each group of samples; is the mean of the overall samples; The sum of squared errors of each sample data in each group and its group average value reflects the discrete situation of each observed value of each sample, representing the influence of random errors, denoted as the sum of squares within groups SSE. The formula for SSE is as follows: where x ij is the j sample values at the i-th level of the control variable; The sum of squared errors of all observed values and the total average value reflects the degree of dispersion of all observed values, denoted as the total sum of squares of errors SST. The formula for SST is as follows: The mean square between groups MSA and the mean square within groups MSE are obtained by dividing the sum of squared errors by their respective degrees of freedom df A and df E respectively, and the MSA / MSE ratio forms an F-distribution: Among them, n is the number of samples; F is used as a statistical test statistic and also has a corresponding probability distribution, that is, it satisfies the probability distribution of F(k - 1, n - k). The larger the F value, the larger SSA is compared to SSE, which means the greater the between-group differences, the more likely the data in different groups are sampled from different populations, and the lower the p-value of the possibility that the null hypothesis holds. On the contrary, the higher the possibility that the null hypothesis holds. Therefore, select the initial screening feature set F’ = {F’1, F’2, F’3,..., F’ m-1 , F’ m}; The wrapper method RFECV first uses the recursive feature elimination method REF to obtain the ranking of each feature, and then based on the ranking, sequentially selects feature subsets with the number of features [min_features_to_select, len(feature)] for model training and cross-validation, and finally selects the feature subset with the highest average score; Step 2: Data preprocessing; Preprocess the initial dataset with an imbalanced number of positive and negative examples; Step 3: Model training; Use the improved Tri-Training method to classify a large amount of unlabeled data, obtain pseudo-labeled data with high confidence based on DBSCAN, and iteratively optimize the base classifier using the training set data and the pseudo-labeled data; Step 4: Input the target monitoring data into the abnormal transaction detection model to obtain the detection result.

2. The method for detecting abnormal transactions in a blockchain based on semi-supervised learning according to claim 1, characterized in that: The features of the initial blockchain transaction dataset include local features and aggregated features; The local information of the node, denoted as local features; The features statistically obtained from aggregating transaction information one hop forward / backward from the central node are marked as aggregated features.

3. The method for detecting abnormal transactions in a blockchain based on semi-supervised learning according to claim 1, wherein: The specific method for the data preprocessing in Step 2 is as follows: Divide the data into multiple subsets and add them to the base classifier for training; for the positive class samples N, randomly sample a subset N from them i and add it to the negative class samples P; that is, each classified subset contains all abnormal transactions and some normal transactions.

4. The method for detecting abnormal transactions in a blockchain based on semi-supervised learning according to claim 1, characterized in that: The specific method for Step 3 is as follows: Step 3.1: Obtain pseudo-labeled data based on the KNN algorithm. Input the labeled dataset L and the unlabeled dataset U, and use the base classifier H i for the data subset X i Perform iterative training, make predictions on U, and obtain the prediction result i_pred; Base classifier H1, base classifier H2, …, base classifier H i-1 , base classifier H i , base classifier H i+1 , … make predictions on U to obtain prediction results 1_pred, 2_pred, …, i-1_pred, i_pred, i+1_pred, ... respectively; Step 3.2: Select the dataset with the same predicted labels in U, and set the neighborhood radius Eps and the threshold MinPts for the number of data objects in the neighborhood; Step 3.3: Arbitrarily select a data object point p from the dataset. If, for the parameters Eps and MinPts, the selected data object point p is a core point, then find all the data object points that are density-reachable from the data object point p to form a cluster; If the selected data object point p is a border point, then select another data object point; Step 3.4: Repeat Step 3.2 and Step 3.3 until all data object points are processed, and finally output the density-connected clusters; Step 3.5: Use the training set data and the pseudo-labeled data to iteratively optimize the base classifier. If the samples are distributed within the same cluster, then add the unlabeled data values and the predicted results to the training set to iteratively optimize the base classifier.

Citation Information

Patent Citations

  • Block chain anomaly detection method and related device based on PCA and RF

    CN114331731A

  • Artificial intelligence supervised learning based laboratory chest pain data check assisted recognition method

    CN109907751A

  • KR20200101504A