Semi-supervised classification learning method and system based on double reconstruction similarity

By constructing a dual reconstruction similarity map, using paired typicality and propagation probability methods, the time-consuming and adaptive problems in semi-supervised classification learning are solved, and high-precision and fast label-free sample classification are achieved.

CN120339716APending Publication Date: 2025-07-18JIANGSU COLLEGE OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510499648.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing semi-supervised classification learning methods have problems such as time-consuming, inability to operate adaptively and similarity learning is not paid enough attention to in the process of label dissemination, resulting in classification complexity and inefficiency.

Method used

By introducing paired typicality and propagation probability, a double reconstruction similarity graph is constructed, and the feature and label information is represented by weighted undirected graphs and one-hot encoding, the non-affinity and propagation probability are calculated, and the non-affinity and propagation probability are converted into a standardized propagation matrix and chunked to achieve accurate classification of label-free samples.

Benefits of technology

It improves classification accuracy, reduces class overlap problem, reduces computational complexity, realizes fast adaptive classification on medium or large-scale data sets, and avoids the dependence of matrix inversion operations and KNN graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339716A_ABST
    Figure CN120339716A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning and classification, in particular to a semi-supervised classification learning method and system based on double reconstruction similarity, and the method comprises the steps: representing feature information through a weighted undirected graph and label information through one-hot coding, and obtaining an initial similarity graph based on the weighted undirected graph; through pairwise typicality, in combination with the initial similarity graph, determining a first reconstruction similar graph; calculating the non-affinity of any node pair to obtain a propagation probability; determining a second reconstruction similarity graph, and converting the second reconstruction similarity graph into a standardized propagation matrix; partitioning the standardized propagation matrix, calculating label predicted values of all label-free samples in the label-free sample set by combining one-hot coding, and determining corresponding label categories; by introducing the pairwise typicality and the propagation probability, the classification precision and the operation speed of the semi-supervised classification task are improved, no extra physical law or statistical law is needed for supporting, and self-adaptive operation can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning and classification, and particularly to a semi-supervised classification learning method and system based on dual reconstruction similarity. Background Art

[0002] In the real world, such as object detection and social media information classification, it is costly to obtain labeled samples, requiring a large amount of manpower and resources. Machine learning relies on training samples and summarizes potential rules to predict new samples. Therefore, classification is extremely challenging in the case of few labeled samples. In order to more reliably classify unlabeled samples based on a very small number of labeled samples, semi-supervised classification learning is usually adopted in recent years.

[0003] Common assumptions in semi-supervised classification learning include the clustering assumption and the manifold assumption. Among them, the clustering assumption believes that the labels of samples in the same cluster are the same, and the decision boundary between different classes is in the low-density region, and some specific requirements and standards need to be formulated. Its learning objective is complex and difficult to optimize. The manifold assumption believes that the labels of samples are the same under the same manifold. The dataset is usually represented by a graph, the similarity between nodes is represented by edges and weights, and a propagation matrix is constructed using the similarity graph and the labels are propagated iteratively, which is expressed as the Label Propagation (LP) algorithm. However, the existing LP algorithm has some limitations: First, it involves explicit or implicit matrix inversion operations, which cannot guarantee the non-negativity of each element in the propagation matrix during the entire label propagation process and is time-consuming. Second, before the LP algorithm runs, the propagation matrix needs to be redefined by calling the K-Nearest Neighbor (KNN) graph each time, and for medium or large-scale datasets, the value of the parameter K is sensitive and difficult to determine in advance, and it cannot run adaptively, resulting in an increase in the algorithm running time. Third, similarity learning has not been taken seriously. In practical applications, it still relies on the initial similarity matrix and cannot truly reconstruct the similarity in an adaptive manner, but only fine-tunes the similarity matrix, that is, partial learning, resulting in a complex label propagation rule. Summary of the Invention

[0004] In order to solve the technical problems that the existing semi-supervised classification learning objective is complex and difficult to optimize, time-consuming, unable to run adaptively, and the learning of similarity has not been paid enough attention, the purpose of the present invention is to provide a semi-supervised classification learning method based on dual reconstruction similarity, and the specific technical solution adopted is as follows:

[0005] Obtain the original dataset, divide the original dataset into a labeled sample set and an unlabeled sample set, and the original dataset has label information and feature information. The feature information is respectively represented by a weighted undirected graph and the label information is represented by one-hot encoding, and an initial similarity graph is obtained based on the weighted undirected graph;

[0006] Define the feature information of any labeled sample or unlabeled sample as a node, and determine the first reconstructed similarity graph by pairwise typicality and combining with the initial similarity graph.

[0007] Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph to obtain the propagation probability.

[0008] Determine the second reconstructed similarity graph according to the dissimilarity and the propagation probability, and convert the second reconstructed similarity graph into a normalized propagation matrix.

[0009] Partition the normalized propagation matrix, combine with one-hot encoding, calculate the label prediction values of all unlabeled samples in the unlabeled sample set and determine the corresponding label categories.

[0010] Preferably, divide the original data set into a labeled sample set and an unlabeled sample set, represent the feature information by a weighted undirected graph and the label information by one-hot encoding respectively, and obtain the initial similarity graph based on the weighted undirected graph, including:

[0011] Divide the original data set into a labeled sample set according to whether it is labeled or not, denoted as and an unlabeled sample set, denoted as where x represents the feature information; y represents the label information; both i and j represent any sample in the sample set, and i = 1, 2, …, l, j = 1, 2, …, u, N = l + u, l represents the total number of samples in the labeled sample set; u represents the total number of samples in the unlabeled sample set; N represents the total number of samples in the original data set; k represents the total number of categories of the label information;

[0012] Represent the feature information by a weighted undirected graph, denoted as G=(V, W), where, and both contain d-dimensional feature information; W represents an N×N adjacency matrix, denoted as the initial similarity graph;

[0013] Convert the label information of the labeled sample set into an l×k-dimensional one-hot encoding y l , and convert the label information of the unlabeled sample set into a u×k-dimensional zero matrix 0 for storage.

[0014] Preferably, W represents an N×N adjacency matrix, denoted as the initial similarity graph, including:

[0015] Calculate any element in the initial similarity graph based on the Gaussian kernel function, and the corresponding calculation formula is:

[0016] w ij =exp(-‖x i -x j‖ 2 / 2σ 2 )

[0017] Among them, w ij represents the element formed by the feature information of the i-th labeled sample and the j-th unlabeled sample in the N×N adjacency matrix; x i represents the feature information of the i-th labeled sample; x j represents the feature information of the j-th unlabeled sample; σ represents the bandwidth of the Gaussian kernel function.

[0018] Preferably, by pairwise typicality, combined with the initial similarity graph, determine the first reconstructed similarity graph, including:

[0019] Calculate pairwise typicality according to the adjacency matrix, and the corresponding calculation formula is:

[0020] t ij = m i m j / ‖W‖

[0021]

[0022] Among them, t ij represents the pairwise typicality of node i and node j; m i and m j respectively represent the strengths of node i and node j; W represents the N×N adjacency matrix; N represents the total number of samples in the original dataset; w ij represents the element formed by node i and node j in the adjacency matrix; n represents the total number of nodes;

[0023] Calculate any element in the first reconstructed similarity graph in sequence to determine the first reconstructed similarity graph, and the corresponding calculation formula is:

[0024]

[0025] Among them, represents the element based on node i and node j in the first reconstructed similarity graph W r ; γ represents a hyperparameter.

[0026] Preferably, calculate the dissimilarity of any node pair based on the first reconstructed similarity graph to obtain the propagation probability, including:

[0027] Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph, and the corresponding calculation formula is:

[0028]

[0029] Among them, d ij represents the dissimilarity of node i and node j; Denote the first - order reconstructed similarity graph as \(W\). r The element formed by nodes \(i\) and \(j\) in it; \(w\). r Denote the first - order reconstructed similarity graph as \(W\). r Any element in it;

[0030] Calculate the propagation probability through the Gaussian exponential distribution, combined with the dissimilarity of node pairs. The corresponding calculation formula is:

[0031]

[0032] Among them, \(p\). ij Represents the propagation probability between nodes \(i\) and \(j\); \(\lambda\) represents the hyperparameter; \(d\). ii Represents the dissimilarity between nodes \(i\) and \(i\); \(d\). jj Represents the dissimilarity between nodes \(j\) and \(j\);

[0033] Based on Judge whether it is greater than 0. If not, adjust the hyperparameter \(\gamma\) until is greater than 0, and recalculate the propagation probability; if so, determine the propagation probability.

[0034] Preferably, determine the second - order reconstructed similarity graph according to the dissimilarity and the propagation probability, and convert the second - order reconstructed similarity graph into a normalized propagation matrix, including:

[0035] Determine the objective function of the second - order reconstructed similarity graph according to the dissimilarity and the propagation probability. The corresponding calculation formula is:

[0036]

[0037] Among them, \(J\) represents the objective function; \(A\) represents the second - order reconstructed similarity graph; \(i\), \(j\), \(q\) all represent nodes; \(a\). ij Represents the element formed by nodes \(i\) and \(j\) in the second - order reconstructed similarity graph; \(\tau\) and \(\lambda\) both represent hyperparameters; \(d\). ii Represents the dissimilarity between nodes \(i\) and \(i\); \(d\). jj Represents the dissimilarity between nodes \(j\) and \(j\); \(d\). ij Represents the dissimilarity between nodes \(i\) and \(j\); \(p\). ij Represents the propagation probability between nodes \(i\) and \(j\);

[0038] Optimize the objective function by solving the stationary point, obtain the optimal solution of the second - order reconstructed similarity graph, and convert it into a normalized propagation matrix based on the optimal solution.

[0039] Preferably, block the normalized propagation matrix, combined with one - hot encoding, calculate the label prediction values of all unlabeled samples in the unlabeled sample set and determine the corresponding label categories, including:

[0040] Perform block partitioning on the normalized propagation matrix using the one-hot encoding that combines the label information of the labeled sample set and the label information of the unlabeled sample set, and calculate the label prediction values of all unlabeled samples in the unlabeled sample set. The corresponding calculation formula is:

[0041]

[0042] Wherein, and respectively represent the one-hot encoding of the label prediction values of the predicted labeled samples and unlabeled samples; A ll , A lu , A ul and A uu represent the matrix elements after block partitioning of the normalized propagation matrix; y l represents the label information of the labeled sample set transformed into a unique encoding;

[0043] Based on the nodes j = 1, 2,..., u in the unlabeled sample set for indexing, determine the label category of any unlabeled sample according to the obtained one-hot encoding of the label prediction value. The corresponding calculation formula is:

[0044]

[0045] Wherein, k j represents the label category of node j; represents the one-hot encoding of the label prediction value of the unlabeled sample the j-th row of.

[0046] To solve the above problems, the present application also provides: a semi-supervised classification learning system based on double reconstruction similarity, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a semi-supervised classification learning method based on double reconstruction similarity as described in any one of the foregoing.

[0047] The present invention has the following beneficial effects:

[0048] 1. Based on the introduced pairwise typicality and propagation probability, pairwise typicality can quantitatively characterize the importance of each node pair, making the connections between similar samples in the reconstructed similarity graph closer, weakening the similarity between dissimilar samples accordingly, and effectively reducing the class overlap problem in the semi-supervised classification task; the propagation concept can represent the difference between the dissimilarity of any node pair and the average dissimilarity of a single sample itself, enabling the doubly reconstructed similarity graph to have a powerful inter-class or intra-class segmentation ability, with higher accuracy compared to the existing LP algorithm. Furthermore, it does not require a complex propagation process involving matrix inversion, that is, it does not require additional physical laws or statistical rules for support. Only by partitioning the normalized propagation matrix and converting the label information into one-hot encoding can accurate classification of the unlabeled sample set be achieved, with faster running speed; and since this application does not need to call the KNN graph to ensure the sparsity of the similarity graph, it does not require manual optimization to set the hyperparameter K, enabling it to run adaptively, and being more convenient in practical applications, especially when facing medium or large-scale data sets.

[0049] 2. The present invention also provides a semi-supervised classification learning system based on double reconstruction similarity for implementing the foregoing semi-supervised classification learning method based on double reconstruction similarity. This system has the same beneficial effects as the foregoing semi-supervised classification learning method based on double reconstruction similarity and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a flowchart of the implementation of a semi-supervised classification learning method based on double reconstruction similarity provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, will detail the specific implementation manners, structures, features, and effects of a semi-supervised classification learning method and system based on double reconstruction similarity proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs.

[0054] The following specifically describes the specific solutions of a semi-supervised classification learning method and system provided by the present invention based on double reconstruction similarity in conjunction with the accompanying drawings.

[0055] The existing semi-supervised classification learning uses complex clustering hypothesis learning objectives that are difficult to optimize; the manifold hypothesis involves explicit or implicit matrix inversion operations, which cannot guarantee the non-negativity of each element in the propagation matrix during the entire label propagation process and is time-consuming; it requires additional physical laws or statistical laws for support and cannot run adaptively, with a slow running speed; similarity learning is not emphasized, resulting in complex label propagation rules; the first embodiment of the present invention provides a semi-supervised classification learning method based on double reconstruction similarity. By introducing pairwise typicality and propagation probability, the similarity between samples of the same class is enhanced, the possibility of class overlap is reduced, the segmentation ability of the similarity graph after double reconstruction is stronger, and block division is performed to achieve accurate classification of the unlabeled sample set; to solve a semi-supervised classification learning method based on double reconstruction similarity, the second embodiment of the present invention provides a semi-supervised classification learning system based on double reconstruction similarity. This system is essentially a software system composed of units that implement corresponding functions. Now, the specific steps in this method will be introduced in detail.

[0056] Please refer to Figure 1 , which shows the implementation flowchart of a semi-supervised classification learning method based on double reconstruction similarity provided by an embodiment of the present invention. The method includes:

[0057] Step S1: Obtain the original data set, divide the original data set into a labeled sample set and an unlabeled sample set, and for the label information and feature information of the original data set, represent the feature information using a weighted undirected graph and the label information using one-hot encoding, and obtain an initial similarity graph based on the weighted undirected graph;

[0058] Step S2: Define the feature information of any labeled sample or unlabeled sample as a node, and determine the first reconstructed similarity graph by combining the initial similarity graph through pairwise typicality;

[0059] Step S3: Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph to obtain the propagation probability;

[0060] Step S4: Determine the second reconstructed similarity graph according to the dissimilarity and the propagation probability, and convert the second reconstructed similarity graph into a normalized propagation matrix;

[0061] Step S5: Partition the standardized propagation matrix, combine with one-hot encoding, calculate the label prediction values of all unlabeled samples in the unlabeled sample set, and determine the corresponding label categories.

[0062] For better illustration, for the constructed initial similarity graph, the feature information of any labeled sample or unlabeled sample therein is used as a node to quantify the similarity relationship between the feature information.

[0063] Further, in step S1, it includes:

[0064] Step S11: Divide the original dataset into a labeled sample set denoted as and an unlabeled sample set denoted as where x represents the feature information; y represents the label information; both i and j represent any sample in the sample set, and i = 1, 2, …, l, j = 1, 2, …, u, N = l + u, l represents the total number of samples in the labeled sample set; u represents the total number of samples in the unlabeled sample set; N represents the total number of samples in the original dataset; k represents the total number of categories of the label information;

[0065] Step S12: Represent the feature information using a weighted undirected graph denoted as G=(V, W), where, and both contain d-dimensional feature information; W represents an N×N adjacency matrix, denoted as the initial similarity graph;

[0066] Step S13: Convert the label information of the labeled sample set into an l×k-dimensional one-hot encoding y l , and convert the label information of the unlabeled sample set into a u×k-dimensional zero matrix 0 for storage.

[0067] Explanation is made that a weighted undirected graph means connecting two nodes to form an edge and assigning a weight to represent the importance between the two connected nodes, used to describe the node relationship; one-hot encoding means converting the label information into binary values, with each category of label information corresponding to a specific binary bit, enabling the label categories to be accepted by machine learning algorithms while preserving the independence between the label categories, without implicitly implying the order or magnitude relationship between the label categories; where, x i1 represents the first feature information component of the i-th sample in the labeled sample set; x i2 represents the second feature information component of the i-th sample in the labeled sample set, and so on; Similarly represents the feature information in the unlabeled sample set.

[0068] Further, in step S12, W represents an N×N adjacency matrix, denoted as the initial similarity graph, including:

[0069] Calculating any element in the initial similarity graph based on the Gaussian kernel function, and the corresponding calculation formula is:

[0070] w ij = exp(-‖x i - x j ‖ 2 / 2σ 2 )

[0071] where w ij represents the element formed by the feature information of the i-th labeled sample and the j-th unlabeled sample in the N×N adjacency matrix; x i represents the feature information of the i-th labeled sample; x j represents the feature information of the j-th unlabeled sample; σ represents the bandwidth of the Gaussian kernel function.

[0072] It should be noted that using the Gaussian kernel function to measure the similarity of nodes in the sample to construct the adjacency matrix can effectively capture the local structural features in the feature information; among them, σ represents the bandwidth of the Gaussian kernel function, so that the similarity measure of the Gaussian kernel function can adapt to different scales in the entire set of node pairs. In this embodiment, it is simply set as the median of the Euclidean distances of all node pairs in the original dataset to reduce the influence of outliers on the similarity measure, so as to ensure that the Gaussian kernel function can cover the scale changes in the original dataset and then accurately reflect the similarity between nodes in the sample.

[0073] Further, in step S2, it includes:

[0074] Step S21: Calculating the pairwise typicality according to the adjacency matrix, and the corresponding calculation formula is:

[0075] t ij = m i m j / ‖W‖

[0076]

[0077] where t ij represents the pairwise typicality of node i and node j; m i and m j respectively represent the intensities of node i and node j; W represents the N×N adjacency matrix; N represents the total number of samples in the original dataset; w ij represents the element formed by node i and node j in the adjacency matrix; n represents the total number of nodes.

[0078] It is explained that the adjacency matrix represents the connection relationship between nodes in the initial similarity graph; pairwise typicality is used to represent the importance of each node pair, that is, the degree of similarity between two nodes. Among them, the connection between samples of the same category is close, while the similarity between samples of different categories is correspondingly weakened.

[0079] Step S22: Calculate any element in the first reconstructed similarity graph in sequence to determine the first reconstructed similarity graph. The corresponding calculation formula is:

[0080]

[0081] where, represents the element formed based on node i and node j in the first reconstructed similarity graph W r ; γ represents a hyperparameter.

[0082] It can be explained that γ represents a hyperparameter and γ ∈ (0, 1), which is used to control the proportion of pairwise typicality.

[0083] Furthermore, in step S3, it includes:

[0084] Step S31: Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph. The corresponding calculation formula is:

[0085]

[0086] where, d ij represents the dissimilarity between node i and node j; represents the element formed based on node i and node j in the first reconstructed similarity graph W r ; w r represents the first reconstructed similarity graph W r ; any element in it.

[0087] It is explained that the dissimilarity of any node pair in the initial similarity graph is calculated for the adjacency matrix by the normalization method. Among them, the dissimilarity refers to the reverse measure of the relationship degree between two nodes of any node pair in the first reconstructed similarity graph, that is, the closer the relationship degree between two nodes, the lower the dissimilarity.

[0088] Step S32: Calculate the propagation probability by combining the Gaussian exponential distribution with the dissimilarity of the node pair. The corresponding calculation formula is:

[0089]

[0090] where, p ij represents the propagation probability between node i and node j; λ represents a hyperparameter; d ii represents the dissimilarity between node i and node i; d jjDenote the dissimilarity between node j and node j.

[0091] It should be noted that the Gaussian exponential distribution is used to ensure that the propagation probability decreases as the dissimilarity between nodes increases. The propagation probability refers to the likelihood of feature information propagating from one node to another, and can measure the propagation efficiency and range of feature information.

[0092] Step S33: Based on Judge whether it is greater than 0. If not, adjust the hyperparameter γ until is greater than 0, and recalculate the propagation probability; if so, determine the propagation probability.

[0093] It should be explained that making a judgment based on the data value is to ensure that the superscript of the exponential function e in step S32 is negative, avoiding the propagation probability being greater than 1, so as to ensure that the obtained propagation probability is meaningful; and the adjustment of the hyperparameter γ is to find an optimal value between 0 and 1 according to different datasets to be classified.

[0094] Furthermore, in step S4, it includes:

[0095] Step S41: Determine the objective function for the second - stage reconstruction similarity graph according to the dissimilarity and the propagation probability. The corresponding calculation formula is:

[0096]

[0097] where J represents the objective function; A represents the second - stage reconstruction similarity graph; i, j, q all represent nodes; a ij represents the element formed by nodes i and j in the second - stage reconstruction similarity graph; τ and λ are both hyperparameters; d ii represents the dissimilarity between node i and node i; d jj represents the dissimilarity between node j and node j; d ij represents the dissimilarity between node i and node j; p ij represents the propagation probability between node i and node j;

[0098] Step S42: Optimize the objective function by solving the stationary point, obtain the optimal solution of the second - stage reconstruction similarity graph, and convert it into a normalized propagation matrix based on the optimal solution.

[0099] It should be noted that a stationary point refers to a point where the gradient of the objective function is zero during the optimization process, representing an extreme point or a stationary point of the objective function. That is, by solving the stationary point, the optimal solution or the region near the optimal solution of the objective function can be found. Then, the optimal solution is converted into a standardized propagation matrix according to the shape of the Hessian matrix, which is used to describe the propagation and update of feature information, ensuring that the changes of the propagation matrix are uniform in all directions, making the propagation process more stable and uniform, and avoiding the influence of too large or too small matrix elements. Among them, the Hessian matrix is a second-order derivative matrix used to describe the curvature of the objective function at a certain point; τ represents a hyperparameter used to control the second-order reconstruction similarity graph A to satisfy the matrix properties of the Hessian matrix.

[0100] Furthermore, in step S5, it includes:

[0101] Step S51: Block the standardized propagation matrix by combining the one-hot encoding of the label information of the labeled sample set and the label information of the unlabeled sample set, and calculate the label prediction values of all unlabeled samples in the unlabeled sample set. The corresponding calculation formula is:

[0102]

[0103] Among them, and respectively represent the one-hot encoding of the label prediction values of the predicted labeled samples and unlabeled samples; A ll 、A lu 、A ul and A uu represent the matrix elements after the standardized propagation matrix is blocked; y l represents the label information of the labeled sample set converted into a unique encoding;

[0104] Step S52: Index based on the nodes j = 1, 2,..., u in the unlabeled sample set, and determine the label category of any unlabeled sample according to the obtained one-hot encoding of the label prediction value. The corresponding calculation formula is:

[0105]

[0106] Among them, k j represents the label category of node j; represents the one-hot encoding of the label prediction value of the unlabeled sample of the j-th row.

[0107] It should be noted that the standardized propagation matrix refers to a symmetric and positive definite matrix, which is used to represent the similarity or adjacency relationship between matrix elements, and then it is blocked. Among them, A ll represents the propagation relationship between samples in the labeled sample set; Alu Represents the propagation relationship between the labeled sample set and the unlabeled sample set; A ul Represents the propagation relationship between the unlabeled sample set and the labeled sample set; A uu Represents the propagation relationship between samples in the unlabeled sample set; combined with the one-hot encoding of the label information of the previously obtained labeled sample set and unlabeled sample set to obtain the corresponding one-hot encoding of the label prediction value; and then determine the class label of the unlabeled sample according to the prediction value, so as to realize the semi-supervised classification learning of the original data set.

[0108] For better illustration, this application uses a proposed semi-supervised classification learning method based on double reconstruction similarity and the Mincut algorithm (Minimum Cut algorithm), HF algorithm (Harmonic Functions), FLAP algorithm (Fick’s Law Assisted Propagation), AELP-WL algorithm (Adaptive Embedded Label Propagation with Weight Learning), and ALP-TMR algorithm (Triple Matrix Recovery-Based Robust Auto-Weighted Label Propagation) to conduct experimental verification on semi-supervised classification tasks on 18 groups of UCI (University of California Irvine) benchmark data sets (i.e., the machine learning library of the University of California, Irvine). That is, each comparison method runs 20 times on each data set under the same experimental conditions to obtain the experimental results, as shown in Table 1, the accuracy rates of different semi-supervised classification methods running on the UCI data sets, and Table 2, the running times (seconds) of different semi-supervised classification methods on the UCI data sets.

[0109] Table 1 Accuracy rates of different semi-supervised classification methods running on the UCI data sets

[0110] data set the method of the present invention Mincut HF FLAP AELP-WL ALP-TMR ala 0.8692±0.0012 0.8247±0.0079 0.7965±0.0063 0.8452±0.0020 0.8603±0.0180 0.8552±0.0162 Amazon 0.7379±0.0047 0.6654±0.0082 0.6370±0.0105 0.7462±0.0058 0.7271±0.0154 0.7389±0.0133 Amphibians 0.7135±0.0196 0.6245±0.0133 0.6457±0.0143 0.7058±0.0188 0.7011±0.0385 0.7102±0.0364 audit 0.9083±0.0025 0.8563±0.0046 0.8541±0.0051 0.8677±0.0030 0.8754±0.0094 0.8837±0.0103 Avila 0.9304±0.0037 0.8872±0.0060 0.9004±0.0077 0.9145±0.0042 0.9156±0.0142 0.9110±0.0168 bank 0.9718±0.0125 - - 0.9550±0.0166 0.9738±0.0235 0.9794±0.0216 epilepsy 0.9401±0.0015 0.8591±0.0011 0.8978±0.0027 0.9010±0.0016 0.9231±0.0067 0.9170±0.0085 fertility 0.7759±0.0234 0.7269±0.0355 0.7038±0.0396 0.7652±0.0227 0.7459±0.0338 0.7362±0.0397 Gait 0.7872±0.0220 0.7149±0.0280 0.7244±0.0305 0.7321±0.0263 0.7543±0.0384 0.7219±0.0292 German 0.7836±0.0088 0.7311±0.0146 0.7257±0.0167 0.7747±0.0090 0.7671±0.0174 0.7725±0.0166 HCV 0.9135±0.0192 0.8537±0.0344 0.8935±0.0321 0.8835±0.0199 0.8976±0.0257 0.9025±0.0299 HTRU2 0.9212±0.0018 0.8661±0.0064 0.8695±0.0088 0.9026±0.0077 0.8865±0.0121 0.9056±0.0169 madelon 0.6072±0.0055 0.5927±0.0085 0.5856±0.0099 0.5943±0.0067 0.5933±0.0168 0.6102±0.0235 Parkinson 0.8609±0.0035 0.7971±0.0102 0.8254±0.0126 0.8475±0.0094 0.8436±0.0235 0.8316±0.0212 Roman 0.8899±0.0060 0.7963±0.0085 0.8406±0.0093 0.8235±0.0109 0.8578±0.0160 0.8644±0.0112 seeds 0.9774±0.0019 0.9215±0.0077 0.9053±0.0054 0.9629±0.0043 0.9614±0.0072 0.9759±0.0064 Synchronous 0.8170±0.0159 0.7250±0.0159 0.7380±0.0259 0.7492±0.0184 0.7570±0.0254 0.7319±0.0271 wine 0.9648±0.0010 0.9269±0.0017 0.9445±0.0024 0.9578±0.0011 0.9632±0.0076 0.9562±0.0073

[0111] Table 2 Running times (seconds) of different semi-supervised classification methods on the UCI data sets

[0112]

[0113]

[0114] It should be noted that, according to the numerical display of the experimental results, a semi-supervised classification learning method based on double reconstruction similarity provided by this application has achieved the highest accuracy and the fastest running speed on most UCI datasets, indicating that the method proposed in this application has higher accuracy and faster running speed compared with the existing LP algorithm.

[0115] An embodiment of the present invention also proposes a semi-supervised classification learning system based on double reconstruction similarity, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a semi-supervised classification learning method according to any one of the foregoing embodiments.

[0116] It is explained that a semi-supervised classification learning system based on double reconstruction similarity has the same beneficial effects as the semi-supervised classification learning method provided above, and will not be elaborated here.

[0117] It should be noted that: the above-mentioned order of the embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0118] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.

Claims

1. A semi-supervised classification learning method based on double reconstruction similarity, characterized in that The method includes: Obtain the original data set, divide the original data set into a labeled sample set and an unlabeled sample set, and the original data set has label information and feature information. Represent the feature information using a weighted undirected graph and the label information using one-hot encoding respectively, and obtain an initial similarity graph based on the weighted undirected graph; Define the feature information of any labeled sample or unlabeled sample as a node, and determine the first reconstructed similarity graph by pairwise typicality in combination with the initial similarity graph; Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph to obtain the propagation probability; Determine the second reconstructed similarity graph according to the dissimilarity and the propagation probability, and convert the second reconstructed similarity graph into a normalized propagation matrix; Partition the normalized propagation matrix, and in combination with one-hot encoding, calculate the label prediction values of all unlabeled samples in the unlabeled sample set and determine the corresponding label categories.

2. A semi-supervised classification learning method based on double reconstruction similarity according to claim 1, characterized in that Dividing the original data set into a labeled sample set and an unlabeled sample set, representing the feature information using a weighted undirected graph and the label information using one-hot encoding respectively, and obtaining an initial similarity graph based on the weighted undirected graph, includes: The original dataset is divided into a labeled sample set, denoted as and an unlabeled sample set, denoted as where x represents feature information; y represents label information; both i and j represent any sample in the sample set, and i = 1, 2, …, l, j = 1, 2, …, u, N = l + u, l represents the total number of samples in the labeled sample set; u represents the total number of samples in the unlabeled sample set; N represents the total number of samples in the original dataset; k represents the total number of categories of label information; The feature information is represented by a weighted undirected graph, denoted as G=(V, W), where, and both contain d-dimensional feature information; W represents an N×N adjacency matrix, denoted as the initial similarity graph; Convert the label information of the labeled sample set into a one-hot encoding y of l×k dimensions l , and convert the label information of the unlabeled sample set into a zero matrix 0 of u×k dimensions for storage.

3. A semi-supervised classification learning method based on double reconstruction similarity according to claim 2, characterized in that, Let \(W\) denote the \(N\times N\) adjacency matrix, denoted as the initial similarity graph, including: Calculate any element in the initial similarity graph based on the Gaussian kernel function, and the corresponding calculation formula is: w ij = exp(-||x i - x j || 2 / 2σ 2 ) where, w ij represents the element formed by the feature information of the \(i\)-th labeled sample and the \(j\)-th unlabeled sample in the \(N\times N\) adjacency matrix; \(x\) i represents the feature information of the \(i\)-th labeled sample; \(x\) j represents the feature information of the \(j\)-th unlabeled sample; \(\sigma\) represents the bandwidth of the Gaussian kernel function.

4. A semi-supervised classification learning method based on double reconstruction similarity according to claim 2, characterized in that Determine the first reconstructed similarity graph by pairwise typicality in combination with the initial similarity graph, including: Calculate pairwise typicality according to the adjacency matrix, and the corresponding calculation formula is: t ij = m i m j / ||W|| where, t ij represents the pairwise typicality of nodes i and j; m i and m j respectively represent the strengths of nodes t and j; W represents the N×N adjacency matrix; N represents the total number of samples in the original dataset; w ij represents the element formed by nodes i and j in the adjacency matrix; n represents the total number of nodes; Calculate any element in the first reconstructed similarity graph in sequence to determine the first reconstructed similarity graph, and the corresponding calculation formula is: Among them, represents the element formed by nodes i and j in the first reconstructed similarity graph W; r γ represents a hyperparameter.

5. A semi-supervised classification learning method based on double reconstruction similarity according to claim 4, characterized in that, Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph to obtain the propagation probability, including: Calculate the dissimilarity of any node pair based on the first reconstructed similarity graph, and the corresponding calculation formula is: Among them, d ij represents the dissimilarity between node i and node j; represents the element formed by node i and node j in the first - order reconstructed similarity graph W r ; w r represents the first - order reconstructed similarity graph W r and any element in it; Calculate the propagation probability by the Gaussian exponential distribution in combination with the dissimilarity of the node pair, and the corresponding calculation formula is: Among them, p ij represents the propagation probability between node i and node j; λ represents a hyperparameter; d ii represents the inaffinity between node i and node i; d jj represents the inaffinity between node j and node j; Based on Judge whether it is greater than 0. If not, adjust the hyperparameter γ until it is greater than 0 and recalculate the propagation probability; if so, determine the propagation probability.

6. A semi-supervised classification learning method based on double reconstruction similarity according to claim 1, characterized in that Determine the second reconstructed similarity graph according to the dissimilarity and the propagation probability, and convert the second reconstructed similarity graph into a normalized propagation matrix, including: Determine the objective function of the second reconstructed similarity graph according to the dissimilarity and the propagation probability, and the corresponding calculation formula is: Among them, J represents the objective function; A represents the second - order reconstruction similarity graph; i, j, and q all represent nodes; a ij represents the element formed by nodes i and j in the second - order reconstruction similarity graph; τ and λ both represent hyperparameters; d ii represents the dissimilarity between node i and node i; d jj represents the dissimilarity between node j and node j; d ij represents the dissimilarity between node i and node j; p ij represents the propagation probability between node i and node j; Optimize the objective function by solving the stationary point to obtain the optimal solution of the second reconstructed similarity graph, and convert it into a normalized propagation matrix based on the optimal solution.

7. A semi-supervised classification learning method based on double reconstruction similarity according to claim 2, characterized in that, Partition the normalized propagation matrix, and in combination with one-hot encoding, calculate the label prediction values of all unlabeled samples in the unlabeled sample set and determine the corresponding label categories, including: Partition the normalized propagation matrix in combination with the one-hot encoding of the label information of the labeled sample set and the label information of the unlabeled sample set, and calculate the label prediction values of all unlabeled samples in the unlabeled sample set, and the corresponding calculation formula is: Among them, and respectively represent the one-hot encoded label prediction values of the predicted labeled samples and unlabeled samples; A ll 、A lu 、A ul and A uu represent the matrix elements after the normalized propagation matrix is partitioned; y l represents the label information of the labeled sample set after transformation; Index based on the nodes \(j = 1, 2, \ldots, u\) in the unlabeled sample set, and determine the label category of any unlabeled sample according to the obtained one-hot encoded label prediction value, and the corresponding calculation formula is: Among them, k j represents the label category of node j; represents the j-th row of the one-hot encoding of the label prediction value of the unlabeled sample of.

8. A semi-supervised classification learning system based on double reconstruction similarity, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a semi-supervised classification learning method based on double reconstructed similarity as described in any one of claims 1 to 7.