Tensor depth semi-supervised learning method for high-dimensional small sample data classification
By constructing a second-order similarity matrix and a third-order similarity tensor combined with label propagation, the problem that traditional methods are difficult to capture complex structures in high-dimensional small sample data is solved, and higher classification accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510427048.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional deep semi-supervised learning methods are difficult to capture complex data structures and multi-order correlation information in high-dimensional small sample data, resulting in susceptibility to noise interference during information transmission and pseudo-label generation, affecting model performance.
By constructing a second-order similarity matrix and a third-order similarity tensor, combining label propagation, the network's modeling ability of complex data structures is improved, and deep neural networks are used for iterative optimization to generate high-quality pseudo-labels.
The classification accuracy and robustness of high-dimensional small sample data are improved, and the classification performance on high-dimensional small sample data sets is significantly improved.
Smart Images

Figure CN120296562A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a tensor deep semi-supervised learning method for high-dimensional small-sample data classification. Background Art
[0002] Semi-supervised learning combines a small amount of labeled data with a large amount of unlabeled data, and uses specific learning strategies to utilize the potential information in the unlabeled data. According to the different ways of using unlabeled data, semi-supervised learning methods can be divided into generative methods, graph-based methods, and deep learning-based methods, etc. Generative methods assume that the data distribution conforms to a certain generative model, and estimate the model parameters to utilize the unlabeled data. Graph-based methods construct a graph structure between data points and use the Laplacian matrix of the graph to capture the intrinsic structure information of the data. Deep learning-based methods design specific network structures and loss functions to utilize unlabeled data for feature learning and model optimization.
[0003] Traditional deep semi-supervised learning methods have achieved certain results, but the existing methods mainly rely on low-order similarity to construct the correlation relationship between samples. However, low-order similarity often only reflects the direct relationship between samples, and it is difficult to capture the complex geometric structure and multi-order correlation information inside the data, resulting in being vulnerable to noise interference during the information transmission and pseudo-label generation processes, which affects the overall performance of the model. Especially in high-dimensional small-sample data or complex scenarios, this limitation makes traditional methods have obvious deficiencies in revealing potential data structures and improving prediction accuracy. Therefore, through richer similarity metrics, more abundant interaction information and local structure features between samples can be captured, so as to more accurately describe the intrinsic geometry and distribution pattern of the data. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose a tensor deep semi-supervised learning method for high-dimensional small-sample data classification. By fusing the second-order similarity matrix and the third-order similarity tensor, and combining label propagation, the modeling ability of the network for complex data structures is improved, and higher classification accuracy and robustness are achieved.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: a tensor deep semi-supervised learning method for high-dimensional small-sample data classification, including the following steps:
[0006] S1: Preprocess the original high-dimensional small-sample data, including cleaning, missing value filling, normalization, and noise suppression;
[0007] S2: Construct a deep neural network containing a feature extraction module and a classifier module. The feature extraction module obtains low-dimensional embedded features through a 13-layer fully connected neural network, and the classifier module outputs class probabilities, i.e., predicted labels, through the Softmax function. Based on the labeled samples in the preprocessed dataset, use the cross-entropy loss function to perform supervised pre-training on the feature extraction module to complete the initialization of network parameters;
[0008] S3: Based on the low-dimensional embedded features obtained from pre-training, construct a second-order similarity matrix and a third-order similarity tensor. The second-order similarity matrix is obtained by calculating the Euclidean distance between low-dimensional embedded features and mapping it to the similarity space using the Gaussian kernel function. The third-order similarity tensor is obtained by comprehensively considering the relative positions and directions among three data points in the low-dimensional embedded features;
[0009] S4: Combine the second-order similarity matrix and the third-order similarity tensor to construct an objective function containing multi-order smoothing constraints. Subsequently, use the low-dimensional embedded features obtained from pre-training as input, and through a label propagation network composed of fully connected layers, adopt the gradient descent algorithm to iteratively optimize the objective function to generate pseudo-labels;
[0010] S5: Input the original high-dimensional small sample data, low-dimensional embedded features, and pseudo-labels into the deep neural network together. By alternately optimizing the parameters of the feature extraction module and the classifier module, iteratively update the network in a semi-supervised manner until convergence, and output the final prediction result.
[0011] Further, in step S2, the feature extraction module is a 13-layer fully connected neural network, and each layer is sequentially connected to a batch normalization layer and a Dropout layer. Among them, the first to 12th layers use the ReLU activation function, and the output of the last layer is normalized by the l2 norm to generate unit norm low-dimensional embedded features; the classifier module is composed of a Softmax layer, which maps the low-dimensional embedded features to the class probability space.
[0012] Further, in step S3, the second-order similarity matrix is obtained by calculating the Euclidean distance between low-dimensional embedded features and mapping it to the similarity space using the Gaussian kernel function, and the third-order similarity tensor captures the high-order correlation information among multiple samples;
[0013] a. The second-order similarity matrix is defined as:
[0014]
[0015] where d ij is the Euclidean distance between samples i and j, is the second-order similarity between samples i and j, and σ is the bandwidth parameter of the Gaussian kernel function;
[0016] b. Third - order similarity tensor: Construct complex associations between samples through a third - order tensor, defined as:
[0017]
[0018] where d ij and d jk are the Euclidean distances between samples i and j, and j and k respectively, and v i , v j , v k are the low - dimensional embedded features of samples i, j, and k respectively. <·,·> represents the vector inner product, is the third - order similarity between samples i, j, and k.
[0019] Furthermore, the specific operation steps of step S4 are as follows:
[0020] S41: Construct a label propagation network composed of dynamically constructed fully - connected layers, which is used to fuse the second - order similarity matrix and the third - order similarity tensor, and optimize the label propagation process through spectral clustering;
[0021] Capture the local geometric structure between samples through the third - order tensor mode product, avoiding the noise sensitivity of traditional separable tensors; introduce tensor spectral analysis technology to optimize the smoothness constraint of label propagation; design the objective function L tensor (F) of label propagation. Balance the contributions of the third - order similarity tensor and the second - order similarity matrix in the objective function. The specific form is:
[0022]
[0023] where is the third - order similarity tensor, is the second - order similarity matrix, C is the number of categories, F is the pseudo - label matrix, represents the transpose of the c - th column of matrix . Here, represents the real number field, m is the total number of samples, Y is the true - label matrix, α and β are balance coefficients, ‖·‖ Frob represents the Frobenius norm, and × k is the k - th dimensional mode product of the tensor;
[0024] In the objective function of the label propagation, the first term is the smoothness constraint term constructed based on the second - order similarity matrix and the third - order similarity tensor:
[0025]
[0026] S42: The optimization of the objective function of the label propagation adopts a dynamic learning rate decay strategy, including: using Stochastic Gradient Descent (SGD) combined with a dynamic learning rate decay strategy to optimize the objective function, and adjusting the initial learning rate by cosine annealing;
[0027] For solving the objective function of the above-mentioned label propagation, a numerical iteration method based on gradient descent is adopted for optimization. After derivation, it is proved that the gradient of the objective function of label propagation with respect to F is as follows:
[0028]
[0029] In the formula, M is the auxiliary matrix after tensor unfolding, which is used for gradient calculation;
[0030] The initial learning rate of stochastic gradient descent is set to 0.05, and the learning rate is periodically adjusted during the training process through the cosine annealing algorithm; the learning rate is adaptively adjusted according to the convergence of the objective function after each iteration. The learning rate η used in the t-th iteration (t) The formula is:
[0031]
[0032] In the formula, η max and η min are the upper and lower bounds of the learning rate, and T′ is the total number of iterations.
[0033] Furthermore, the specific operation steps of step S5 are as follows:
[0034] a. Define the joint loss function:
[0035]
[0036] In the formula, L net is the joint loss function, l is the number of labeled samples, m is the total number of samples, l s is the cross-entropy loss, f θ (x i ) is the predicted output of the deep neural network for the input sample xi, y i is the true label, is the pseudo label, is the class balance weight, and γ i is the pseudo label confidence weight, which is calculated through the entropy function H(·):
[0037]
[0038] The calculation of the pseudo label confidence weight γ i further includes row normalization of the predicted probability matrix F, represents the probability distribution after normalizing the i-th row of the pseudo label matrix F, and we get Then, the prediction uncertainty is measured through the entropy function. The lower the entropy value, the higher the confidence. The formula is:
[0039]
[0040] In the formula, H(·) is the entropy function, is the predicted probability matrix after normalization, is the element in the i-th row and c-th column of
[0041] Combined with the class balance weight δ c , to prevent the samples of the minority class from being ignored, it is defined as:
[0042]
[0043] In the formula, |L c | and |U c | are the numbers of samples of class c in the labeled and unlabeled data respectively;
[0044] b. By alternately updating the network parameters and the pseudo-labels until convergence, the final classification result is output.
[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0046] 1. The present invention introduces a third-order similarity tensor, which effectively captures the complex internal correlations between samples, thus making up for the deficiencies of traditional pairwise similarities in high-dimensional small-sample scenarios;
[0047] 2. The present invention designs label propagation that fuses second-order and third-order information, uses a deep neural network to guide the generation of pseudo-labels, and realizes more accurate and robust label propagation.
[0048] 3. The present invention designs a tensor deep semi-supervised learning framework for high-dimensional small-sample data classification, uses limited labeled data and rich unlabeled data to improve the quality of feature representation, and thus improves the overall classification performance.
[0049] In summary, the present invention can fully explore the internal high-order correlations of data by introducing high-order similarities, realize the complementarity of multi-order correlation information and accurate label propagation, and improve the semi-supervised classification accuracy on high-dimensional small-sample data through the joint optimization framework of the deep neural network and label propagation. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is the framework diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0052] As Figure 1As shown in the figure, this embodiment discloses a tensor deep semi-supervised learning method for high-dimensional small-sample data classification, which uses a deep neural network as the base model for feature extraction, and includes the following steps:
[0053] 1) Perform multi-step preprocessing on the original high-dimensional small-sample data to ensure the stability and accuracy of subsequent model training. The preprocessing steps include:
[0054] Feature cleaning: Use the variance threshold method to filter low-variance features, and set the threshold to 0.01. This method can eliminate redundant dimensions and uninformative features to ensure the effectiveness of input features;
[0055] Missing value filling: Adopt the K-nearest neighbor interpolation method, where the value of K is taken as 5, and at the same time, the mean value of the nearest neighbor samples is selected based on the Euclidean distance to fill the missing values to ensure data integrity;
[0056] Normalization: Perform Z-score normalization on each feature dimension to obtain the normalized value of each feature after Z-score normalization.
[0057] Noise suppression: Apply the Local Outlier Factor (LOF) algorithm to detect noise samples (set the number of neighbors to 20). For the samples detected as noise, replace them with the mean value of the same-class samples to reduce the interference of noise on subsequent model training.
[0058] 2) Deep feature extraction and pre-training. The main purpose of the pre-training stage is to obtain stable and discriminative deep feature representations using limited labeled data. This stage is trained with 180 epochs, and the specific steps are as follows:
[0059] a. Network initialization, randomly initialize the parameters θ of the deep neural network, where the network includes two parts:
[0060] Feature extraction module: Feature extractor φ θ Used to map the original data to a low-dimensional space to obtain deep features, and adopt a 13-layer fully connected neural network; after the output of φ θ Add an l2 normalization layer to generate a feature descriptor with unit norm;
[0061] Classifier module: Connect a fully connected layer after the feature extractor and use the softmax activation function to output the class probability distribution.
[0062] b. Full-supervised training, use the labeled sample set for training, and the loss function adopts the cross-entropy form:
[0063]
[0064] In the formula, L sis the supervised loss function, l is the number of labeled samples, and f θ (x i ) is the predicted output of the deep neural network for the input sample xi, and l s is the cross-entropy loss, and y i is the true label. By using optimization methods such as Stochastic Gradient Descent (SGD) to continuously update the parameter θ until satisfactory performance is achieved on the validation set, a stable and discriminative feature representation is obtained. To further improve the robustness of the features, after the pre-training is completed, an l2 normalization operation is introduced at the output layer of the feature extractor to unitize the feature vectors of each sample.
[0065] 3) To fully characterize the relationships between samples, similarity metrics based on low-order and high-order information are constructed respectively;
[0066] 3.1) Construct a second-order similarity matrix According to the Euclidean distance d ij between samples, a similarity matrix is constructed, and its formula is:
[0067]
[0068] where d ij is the Euclidean distance between samples i and j, and σ is a parameter that controls the standard deviation of the similarity distribution. This matrix can characterize the pairwise similarity between samples, but there may be a "distance concentration" problem in the high-dimensional small-sample scenario.
[0069] 3.2) Construct a third-order similarity tensor To solve the problem that low-order similarity cannot fully reflect the complex internal relationships between samples, a third-order tensor is constructed based on the joint relationship between sample features:
[0070]
[0071] In the formula, d ij and d jk are the Euclidean distances between samples i and j and between j and k respectively, v i 、v j 、v k are the low-dimensional feature representations of samples i, j, and k respectively, <·,·> represents the inner product of vectors, is the third-order similarity between samples i, j, and k;
[0072] This formula measures the similarity between sample x i and x j from the perspective of sample x k and can capture high-order structural information, making up for the deficiencies of second-order similarity.
[0073] 4) After constructing the similarity metric, use the tensor label propagation network h ω , with its parameter ω, to generate pseudo-labels. The process is as follows:
[0074] 4.1) Using the feature matrix V = [v1, v2,..., v m as the input, generate a preliminary pseudo-label matrix F through the network h ω , that is:
[0075] F = hω(V)
[0076] Design the pseudo-label objective function L tensor (F). This objective function combines the smoothness constraint of tensor spectral clustering and the pseudo-label fitting term, as follows:
[0077]
[0078] where C is the number of classes, is the label matrix, here represents the real number field, m is the total number of samples, the labeled sample rows are represented by one-hot encoding, and the unlabeled sample rows are all zeros, represents the transpose of the c-th column of the matrix , α is the balance coefficient, β is the hyperparameter for adjusting the contribution of the second-order tensor, and × k represents the k-th dimensional mode product of the tensor.
[0079] 4.2) Gradient descent optimization: Since the objective function is difficult to obtain a closed-form solution, a numerical iterative method based on gradient descent is used to update the pseudo-label matrix F and the propagation network parameter ω. After derivation, the gradient of the objective function with respect to F is:
[0080]
[0081] where,
[0082] and represents the mode-1 unfolding of the tensor , represents the Kronecker product.
[0083] Moreover, the hyperparameters α and β are respectively selected from the sets {0.1, 0.3, 0.5, 0.7, 0.9} and {0.01, 0.1, 1, 10, 100} through grid search to obtain the optimal combination; the initial learning rate l0 is set to 0.05, and the cosine annealing strategy is used for decay.
[0084] 4.3) Parameter iterative update: Use the stochastic gradient descent (SGD) method to update the label propagation network parameter ω. The update formula is:
[0085]
[0086] where η is the learning rate, t is the number of training rounds, and ω (t) is the network parameter at the t-th round.
[0087] An iterative method based on gradient descent is used to update the pseudo-label matrix F and the label propagation network parameter ω. In each epoch of the semi-supervised stage, a maximum of 20 optimization iterations are set. In each iteration, after normalizing the pseudo-label vector, the maximum probability index is taken to generate the pseudo-labels of unlabeled samples Meanwhile, the confidence γ of each pseudo-label is calculated through the entropy function i , and this confidence is used to adjust the weight of the pseudo-label in the joint training stage, so that the pseudo-labels with high confidence contribute more to the network training.
[0088] 5) After obtaining the pseudo-label matrix F, the true labels and pseudo-labels are combined to construct an overall joint loss function, and the deep neural network f θ is jointly trained. The specific process is as follows:
[0089] 5.1) To alleviate the problem of uneven class distribution, a class balance weight δ is defined for each class c c
[0090]
[0091] where |L j | and |U j | represent the number of samples belonging to class c in the labeled samples and unlabeled samples respectively. Let and represent the balance weights of the corresponding samples, that is, take the balance weight value of the class to which they belong.
[0092] For the pseudo-labels of unlabeled samples, the confidence γ i is also introduced to adjust the contribution of the pseudo-labels. The specific definition is as follows:
[0093]
[0094] where H is the entropy function, represents the probability distribution after normalizing the i-th row of the pseudo-label matrix F, and its normalization formula is:
[0095]
[0096] And the entropy function is defined as:
[0097]
[0098] where, is the predicted probability matrix after normalization, is the element in the \(i\)-th row and \(c\)-th column of
[0099] Based on the above definitions, the overall joint loss function \(L\) net is defined as:
[0100]
[0101] where \(L\) net is the joint loss function, \(l\) is the number of labeled samples, and \(m\) is the total number of samples. \(l\) s represents the cross-entropy loss function, \(f\) θ (·) is the output of the entire deep neural network, \(y\) i is the true label of the labeled sample, and is determined by the index of the maximum value in each row of the pseudo-label matrix \(F\), and are the class balance weights for the labeled samples and pseudo-labeled samples respectively, which are used to alleviate the problem of uneven class distribution.
[0102] 5.2) The iterative process of the overall joint training is as follows:
[0103] 1. Feature extraction and pseudo-label update:
[0104] In each epoch, first fix the network parameters \(\theta\) after pre-training, and use the current model to extract the deep features of all samples. Based on the extracted features, update the second-order similarity matrix and the third-order similarity tensor and use the label propagation network \(h\) ω to generate a new pseudo-label matrix \(F\).
[0105] In each iteration, after normalizing the pseudo-label vector, obtain the pseudo-labels of the unlabeled samples by taking the index of the maximum probability, and calculate the confidence \(\gamma\) of each pseudo-label using the entropy function i .
[0106] 2. Deep neural network parameter update:
[0107] Subsequently, jointly construct the joint loss function using the updated pseudo-labels and the true labels of the labeled samples, and update the deep neural network through backpropagation using stochastic gradient descent or other optimizers.
[0108] 3. Iteration and learning rate decay:
[0109] The entire training process is an alternating iteration among three stages: pre-training, pseudo-label generation, and joint training. Within each epoch, first fix θ to extract features and update the pseudo-labels, and then update the parameters of f using the true labels and the newly generated pseudo-labels. θ The parameters of.
[0110] During the iteration process, continuously optimize the accuracy of the pseudo-labels, improve the utilization rate of unlabeled data, until the joint loss L net converges or reaches the preset maximum number of iterations.
[0111] Through the above iteration process, the model can continuously improve the classification accuracy and robustness under high-dimensional small-sample data.
[0112] To verify the effectiveness of the TLPDSL model under high-dimensional small-sample (HDLSS) data, this paper selects six public datasets, including five HDLSS datasets and one conventional dataset. The specific information is as follows:
[0113] Table 1 Dataset Details
[0114] Serial number Data set Number of samples Number of features Number of classes 1 Colon 62 2000 2 2 Leukemia 72 7070 2 3 ALLAML 72 7129 2 4 Prostate GE 102 5966 2 5 Lung 203 3312 5 6 USPS 1000 256 10
[0115] There are significant differences in the number of samples, feature dimensions, and number of classes among the datasets, providing a sufficient test basis for evaluating the robustness of the model in the HDLSS scenario. In the preprocessing stage, the same training, validation, and test division method is adopted for each dataset. Among them, 80% of the data is used as the training set, 20% of the data is used as the test set, and only 20% of the samples in the training set are labeled with the true class, and the remaining samples are trained using the pseudo-labels generated by the model.
[0116] To comprehensively evaluate the model performance, this paper adopts the following evaluation metrics:
[0117] Accuracy: Measure the overall proportion of correct predictions;
[0118] Recall: Measure the ability to identify positive samples;
[0119] F1-Score: Consider both precision and recall comprehensively;
[0120] ROC curve and AUC value: Calculate the AUC of non-binary classification tasks through the One-vs-One strategy, comprehensively reflecting the ability of the model to distinguish positive and negative samples.
[0121] This paper compares with a variety of current mainstream semi-supervised methods. The baseline methods include: MLP, A2LP, lapoleaf, ILP, GCN, GAT, LPDSL.
[0122] The experimental results are shown in Table 2:
[0123] Table 2 Analysis Table of Experimental Results
[0124]
[0125]
[0126] Compared with the above methods, in addition to comprehensively using second-order similarity, TLPDSL also introduces third-order similarity information, which can more accurately capture the complex internal relationships between samples, thus generating higher-quality pseudo-labels and significantly improving the classification performance on the HDLSS dataset.
[0127] The experimental results fully demonstrate the effectiveness of the proposed deep semi-supervised learning method based on high-order similarity. By fusing third-order similarity information, this method can fully explore the complex and diverse associations between samples, accurately capture the subtle structures in the intersection area, and significantly improve the classifier's ability to identify complex data structures. Using high-order similarity for label propagation not only extends the limitations of traditional pairwise similarity in local neighborhood capture but also reduces noise interference and error accumulation through information interaction among multiple samples, thereby improving classification accuracy. In addition, the joint optimization of deep feature extraction and high-order similarity further enhances the model's ability to model the internal relationships of data, achieving a significant improvement in classification performance on high-dimensional small-sample datasets and conventional datasets. Through the above steps, the complex structure information of the data can be fully utilized to further improve the overall classification effect.
[0128] Experimental Conclusion: Aiming at the problem that existing algorithms are difficult to comprehensively capture the complex relationships in high-dimensional small-sample data, the present invention proposes a tensor deep semi-supervised learning method for high-dimensional small-sample data classification. Experimental evaluations on multiple publicly available datasets show that the method of the present invention can make full use of multi-order similarity information to accurately reveal the internal associations between samples, and thus is significantly superior to existing methods in terms of classification accuracy, recall rate, and F-Score. In the next research, applications in larger-scale datasets and other application fields such as medical image analysis and genomic data mining will be explored, and advanced technologies such as transfer learning and adversarial training will be combined to further improve the robustness and adaptability of the model. This method has broad application prospects and is worthy of promotion in a wider range of scenarios.
[0129] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. Tensor deep semi-supervised learning method for high-dimensional small-sample data classification, characterized in that It includes the following steps: S1: Preprocess the original high-dimensional small-sample data, including cleaning, missing value filling, normalization, and noise suppression; S2: Construct a deep neural network including a feature extraction module and a classifier module. Among them, the feature extraction module obtains low-dimensional embedded features through a 13-layer fully connected neural network, and the classifier module outputs class probabilities, that is, predicted labels, through the Softmax function; based on the labeled samples in the preprocessed dataset, the cross-entropy loss function is used to perform supervised pre-training on the feature extraction module to complete the initialization of network parameters; S3: Based on the low-dimensional embedded features obtained by pre-training, construct a second-order similarity matrix and a third-order similarity tensor. The second-order similarity matrix is obtained by calculating the Euclidean distance between low-dimensional embedded features and mapping it to the similarity space by combining the Gaussian kernel function. The third-order similarity tensor is obtained by comprehensively considering the relative positions and directions of three data points in the low-dimensional embedded features; S4: Combine the second-order similarity matrix and the third-order similarity tensor to construct an objective function containing multi-order smoothing constraints; then, using the low-dimensional embedded features obtained by pre-training as input, through the label propagation network composed of fully connected layers, use the gradient descent algorithm to iteratively optimize the objective function to generate pseudo-labels; S5: Input the original high-dimensional small-sample data, low-dimensional embedded features, and pseudo-labels into the deep neural network together. By alternately optimizing the parameters of the feature extraction module and the classifier module, iteratively update the network in a semi-supervised manner until convergence, and output the final prediction result.
2. The tensor deep semi-supervised learning method for high-dimensional small-sample data classification according to claim 1, wherein In step S2, the feature extraction module is a 13-layer fully connected neural network, and each layer is sequentially connected to a batch normalization layer and a Dropout layer. Among them, the ReLU activation function is used in the first to twelfth layers, and the output of the last layer is normalized by the norm to generate a unit norm low-dimensional embedded feature; The classifier module is composed of a Softmax layer, which maps low-dimensional embedded features to the class probability space.
3. The tensor deep semi-supervised learning method for high-dimensional small-sample data classification according to claim 1, wherein In step S3, the second-order similarity matrix is obtained by calculating the Euclidean distance between low-dimensional embedded features and mapping it to the similarity space using the Gaussian kernel function. The third-order similarity tensor captures the high-order correlation information among multiple samples; a. The second-order similarity matrix is defined as: where d ij is the Euclidean distance between samples i and j, is the second-order similarity between samples i and j, and σ is the bandwidth parameter of the Gaussian kernel function; b. Third-order similarity tensor: Construct complex associations between samples through a third-order tensor, defined as: where d ij and d jk are the Euclidean distances between samples i and j, and j and k respectively, and v i , v j , and v k are the low-dimensional embedded features of samples i, j, and k respectively. <·,·> represents the vector inner product, is the third-order similarity between samples i, j, and k.
4. The tensor deep semi-supervised learning method for high-dimensional small-sample data classification according to claim 1, characterized in that The specific operation steps of step S4 are: S41: A label propagation network is constructed by dynamically constructed fully connected layers, which is used to fuse the second-order similarity matrix and the third-order similarity tensor, and optimize the label propagation process through spectral clustering; Capture the local geometric structure between samples through the third-order tensor mode product, avoiding the noise sensitivity of traditional separable tensors; introduce tensor spectral analysis technology to optimize the smoothness constraint of label propagation; design the objective function L tensor (F), balance the contributions of the third-order similarity tensor and the second-order similarity matrix in the objective function, and the specific form is: In the formula, is the third-order similarity tensor, is the second-order similarity matrix, C is the number of categories, F is the pseudo-label matrix, represents the transpose of the c-th column of the matrix , where represents the real number field, m is the total number of samples, Y is the true label matrix, α and β are balance coefficients, ‖·‖ Frob represents the Frobenius norm, and × k is the k-th dimensional mode product of the tensor; In the objective function of the label propagation, the first term is a smoothing constraint term constructed based on the second-order similarity matrix and the third-order similarity tensor: S42: The optimization of the objective function of the label propagation adopts a dynamic learning rate decay strategy, including: using Stochastic Gradient Descent (SGD) combined with a dynamic learning rate decay strategy to optimize the objective function, and adjusting the initial learning rate by cosine annealing; For solving the above objective function of label propagation, a numerical iterative method based on gradient descent is used for optimization. After derivation, it is proved that the gradient of the objective function of label propagation with respect to F is as follows: In the formula, M is an auxiliary matrix after tensor unfolding, which is used for gradient calculation; The initial learning rate of stochastic gradient descent is set to 0.05, and the learning rate is adjusted periodically during training by the cosine annealing algorithm; the learning rate is adaptively adjusted according to the convergence of the objective function after each iteration. The learning rate η used at the t-th iteration (t) The formula is: where η max and η min are the upper and lower bounds of the learning rate, and T' is the total number of iterations.
5. The tensor deep semi-supervised learning method for high-dimensional small-sample data classification according to claim 1, characterized in that The specific operation steps of step S5 are: a. Define a joint loss function: where, L net is the combined loss function, l is the number of labeled samples, m is the total number of samples, is the cross-entropy loss, f θ (x i ) is the predicted output of the deep neural network for the input sample x i , y i is the true label, is the pseudo label, is the class balance weight, γ i is the pseudo label confidence weight, calculated by the entropy function H(·): The pseudo-label confidence weight γ i The calculation further includes row-normalizing the prediction probability matrix F, representing the probability distribution after normalizing the i-th row of the pseudo-label matrix F, and obtaining then measuring the prediction uncertainty through the entropy function. The lower the entropy value, the higher the confidence. The formula is: where \(H(\cdot)\) is the entropy function, is the normalized predicted probability matrix, is the element at the \(i\)-th row and \(c\)-th column of Combined with the class balance weight δ c , to prevent minority class samples from being ignored, defined as: where, |L c | and |U c | are the numbers of samples of class c in the labeled and unlabeled data, respectively; b. By alternately updating network parameters and pseudo-labels until convergence, output the final classification result.
Citation Information
Cited By
Wi-Fi-oriented intelligent performance sensing and predicting method and system
CN120676385A
A method and system for intelligent performance perception and prediction for Wi-Fi
CN120676385B
Geomagnetic gradient tensor depth representation learning method and system
CN120763607A
Geomagnetic gradient tensor depth representation learning method and system
CN120763607B
Class imbalance depth map learning method and system based on class incidence relation and cost sensitivity
CN121303190A