Deep compact feature representation network with dynamic rank correlation

Through feature diffusion enhancement, cluster consistency and rank-related dynamic mask modules, the pseudo-label transfer bias problem of teacher-student framework in semi-supervised learning is solved, the utilization efficiency of label-free data and the generalization ability of the model are improved, and more compact and robust feature representation and prediction are achieved.

CN120375025APending Publication Date: 2025-07-25GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510521207.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the existing semi-supervised learning method, the confirmation bias problem caused by pseudo-label transfer in the teacher-student framework is problematic, and the unlabeled data utilization is inefficient, making it difficult to achieve effective feature representation and prediction in category imbalance and low-density areas.

Method used

The feature diffusion enhancement module, cluster consistency module and rank-related dynamic mask module are adopted to break the global and local dependence of the image through pixel-by-pixel noise perturbation and adaptive metric learning, guide the transition of features to high-density areas, and construct dynamic masks through rank correlation coefficients to optimize feature distribution and prediction consistency.

Benefits of technology

It significantly improves the generalization performance and robustness of the model in semi-supervised learning tasks, effectively alleviates confirmation bias and error accumulation, and achieves more compact and accurate feature representation and prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375025A_ABST
    Figure CN120375025A_ABST
Patent Text Reader

Abstract

The invention provides a deep compact feature representation network model based on a rank correlation induced dynamic mask mechanism, which is used for relieving confirmation deviation generated by distillation training due to excessive unlabeled data in semi-supervised learning and effectively compressing distribution of features in a potential space. According to the method, global and local connection between pixels is broken through a feature diffusion enhancement module, a cluster sensing neighborhood consistency module is introduced, and intra-class compactness and inter-class separability are enhanced by using adaptive metric learning to realize pixel-by-pixel feature representation; a rank correlation induction dynamic mask module is adopted, a category correlation mode is reserved in a rank matching mode, and high-probability errors are filtered out, so that the problem of rank relation loss caused by semantic consistency weakening probability alignment randomness is ensured. Evaluation on a plurality of data sets shows that knowledge is distilled to the student network through the teacher network, meanwhile, the correlation between the feature distribution geometric structure and the prediction space is optimized, and the method is obviously superior to an existing method, so that the generalization performance of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a semi-supervised learning method based on deep compact features and dynamic rank correlation guidance. This method integrates feature diffusion, clustering consistency, and dynamic masking strategies to achieve efficient utilization of unlabeled data and knowledge distillation. Background Art

[0002] In semi-supervised learning, to optimize the representation of the latent feature space, researchers have proposed various methods to enhance the compactness of features. For example, using the teacher graph to detect neighborhood relationships to smooth the predictions of the student model, but this method relies heavily on the accuracy of the teacher graph and is easily affected by noise or insufficient neighborhood information; or by penalizing the distance between adjacent samples to improve intra-class compactness, but it has limited effect on the problems in low-density regions. At the same time, denoising diffusion probabilistic models and gradient-matching-based models introduce controlled noise injection and denoising processes to obtain more robust feature representations, but these methods usually do not provide a clear processing mechanism for specific class low-density regions. For this reason, subsequent work began to explore strategies based on graph diffusion and geometric deep learning to capture the local geometric structure of the data, although there are still challenges in issues such as class imbalance.

[0003] In terms of the knowledge distillation and alignment problems at the prediction level, early methods mostly adopted self-training, co-training, or pseudo-label strategies, using the class with the highest prediction confidence as the pseudo-label to utilize unlabeled data, but this way often led to error accumulation. To alleviate this problem, researchers introduced strategies such as consistency regularization, temporal ensembling, virtual adversarial training, etc., aiming to maintain the stability of model predictions under various perturbations. In addition, traditional knowledge distillation directly matching the outputs of the teacher and the student is prone to overfitting the wrong information of the teacher. Therefore, in recent years, improved methods based on dynamic masking, contrastive representation distillation, label smoothing, and logarithmic regularization have emerged. These methods, while filtering wrong predictions, strive to retain the fine-grained ranking relationships between classes, thereby achieving more accurate teacher-student knowledge alignment. Summary of the Invention

[0004] The object of the present invention is to provide a semi-supervised learning method based on deep compact feature representation and dynamic rank correlation guidance to solve the confirmation bias problem caused by pseudo-label transfer in the existing teacher-student framework and improve the utilization efficiency of unlabeled data.

[0005] To achieve the above object, the present invention provides a semi-supervised learning method, including the following:

[0006] Construct a semi-supervised learning model, where the semi-supervised learning model includes a student network and a teacher network, and the student network and the teacher network share some weights;

[0007] Feature diffusion enhancement is introduced. This module uses a forward noise injection and reverse denoising reconstruction process based on stochastic differential equations to break the global and local dependencies among pixels in the image pixel by pixel, thereby obtaining a more compact feature representation. Under the guidance of the graph convolutional network, this process ensures that the addition and elimination of noise can accurately restore the intrinsic features of each pixel and gradually guide the low-density regions to the high-density state;

[0008] By calculating the Euclidean distance between pixels in the feature space and adopting a similarity weighting strategy based on radial basis functions, a clustering loss function is constructed. While reducing the distance between similar samples, it promotes the differences between different samples, making the feature distribution more compact and distinguishable, thus overcoming the adverse effects brought by low-density regions;

[0009] A dynamic mask is generated based on the prediction confidence of the teacher network output, effectively filtering out high-probability incorrect predictions; Subsequently, the predictions of the teacher and student network outputs are rank-ordered. By calculating the relative ranking difference between the two and using the rank correlation coefficient to construct a rank correlation loss, the precise alignment of the relative structure of the network outputs is achieved, thereby retaining the high-confidence semantic information in the teacher network and reducing the transmission of incorrect information;

[0010] The optimization process is repeated until the preset conditions are met, and the optimized model is used to implement semi-supervised distillation learning.

[0011] Optionally, the feature diffusion enhancement layer uses a noise injection method based on stochastic differential equations. Through the pixel-by-pixel noise scrambling and denoising process, the global and local dependencies among pixels in the image are broken, and then a compact representation of features is achieved.

[0012] The clustering consistency layer includes an adaptive metric learning module that calculates the similarity of neighboring samples in the feature space. By minimizing the distance between similar samples and increasing the gap between different samples, the intra-class compactness and inter-class separability are improved;

[0013] The rank correlation dynamic mask layer calculates the rank difference between the prediction values of the teacher network and the student network and generates a dynamic mask based on this difference to suppress high-confidence incorrect predictions, thereby guiding the student network to better learn the structured information in the teacher network;

[0014] Optionally, the noise injection method gradually increases the noise through the forward diffusion process and restores the original features through the reverse denoising process, thereby ensuring a high-density representation of the feature representation in the latent feature space.

[0015] Optionally, the adaptive metric learning module uses radial basis functions to calculate the similarity between samples in the feature space and optimizes the clustering effect in combination with a weight weighting mechanism.

[0016] Optionally, the rank - correlation dynamic mask layer realizes the alignment of prediction consistency between the teacher and student networks through a ranking mechanism based on the rank - correlation coefficient, further enhancing the generalization ability of the student network.

[0017] The technical effects of the present invention are as follows:

[0018] The present invention analyzes and models from the perspectives of feature representation and prediction correlation. During the process of using the knowledge of the teacher model to guide the learning of the student model through the deep compact feature representation network, the feature diffusion enhancement module, the clustering - aware neighborhood consistency module, and the rank - correlation - induced dynamic masking module are innovatively introduced. The diffusion enhancement module focuses on breaking the global - local spatial connection between pixels, decoupling and enhancing feature representation at the pixel level, and guiding the smooth transition of features from low - density regions to high - density regions. The clustering - aware neighborhood consistency module further imposes clustering - aware neighborhood consistency constraints in the feature space, prompting the features of similar samples to be close to each other and the features of different samples to be far from each other, improving the intra - class compactness and inter - class separability of feature representation. The rank - correlation - induced dynamic masking module innovatively introduces a rank - correlation - coefficient dynamic masking mechanism during the knowledge distillation stage. While suppressing the high - confidence mispredictions of the teacher model, it effectively retains and transfers the inter - class rank correlation in the prediction results of the teacher model, preventing the student model from falling into the quagmire of confirmation bias. The present invention combines pixel - level noise perturbation to guide feature diffusion enhancement, adaptive metric learning to achieve clustering neighborhood consistency, and rank - correlation matching for dynamic masking, forcing the model to co - optimize the geometric structure of the feature distribution and the rank correlation of the prediction space during the semi - supervised learning process. Finally, the model can obtain a more compact and robust feature representation and more accurate and consistent prediction results with limited labeled data, effectively alleviating the class - specific low - density region problem and error accumulation phenomenon caused by confirmation bias in traditional semi - supervised learning methods, and significantly improving the generalization performance and reliability of the model in semi - supervised learning tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0020] Figure 1 is a schematic diagram of the overall framework of an embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0022] It should be noted that the steps shown in the framework diagram of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0023] Embodiment 1

[0024] As Figure 1 shown, in this embodiment, a method for a deep compact feature representation network is provided, including:

[0025] Dataset preparation

[0026] The present invention has been comprehensively evaluated and verified on multiple image classification benchmark datasets. To test the performance of the model under different data scales and class complexities, this study adopted three classic image classification datasets, namely SVHN, CIFAR-10, and CIFAR-100. The SVHN dataset, originating from street view house number images in real scenes, contains 600,000 images in 10 categories and has practical application significance; the CIFAR-10 dataset, as a classic benchmark in the field of computer vision, contains 60,000 color images in 10 categories, with rich image content and complex backgrounds; the CIFAR-100 dataset, on the basis of CIFAR-10, is further extended to 100 categories, with a total of 60,000 color images with a resolution of 32x32. The increase in the number of categories and the improvement of inter-class similarity pose higher challenges to the feature representation ability of the model.

[0027] Data preprocessing

[0028] Since the training data is limited, in order to increase the amount of training data and allow the model to be fully trained, the embodiment of the present invention introduces data augmentation operations. We adjust all input data to 32×32, randomly crop and fill the pixels around the edges of the image, and then expand the data volume by data augmentation methods such as horizontal flipping with a random probability.

[0029] Parameter setting

[0030] The proposed method is implemented in the PyTorch framework running on a GPU (Tesla GV100). The momentum of the mini-batch stochastic gradient descent (SGD) optimizer is set to μ = 0.9, and the scaling factor ε and the temperature parameter β are set to 6 and 4 respectively.

[0031] On the basis of the above embodiment, a student network for final training and a teacher network for auxiliary training are built. The teacher network is a stable and slow-updating version of the student network, and the two share network weights. On these two networks, co-training is carried out using labeled data, unlabeled data, and pseudo-labels of the teacher network.

[0032] In the first step, the RGB image first enters the initial convolution module. This module consists of a 3x3 convolutional layer, a batch normalization layer, and a ReLU activation function, aiming to extract basic features such as shallow edges and textures of the image. The convolution operation initially establishes the spatial correlation between pixels. The ReLU activation function introduces non-linearity, and the batch normalization layer accelerates training and improves the model's generalization ability. The output of the module is a shallow feature map, which retains the basic information of the image and prepares for subsequent deep feature extraction.

[0033] In the second step, the shallow feature map sequentially enters four residual module groups for layer-by-layer extraction of deep features. In the residual module group, the number of channels of the feature map gradually increases, and the size gradually decreases, so as to extract feature representations of different scales and different abstraction levels. It effectively alleviates the problem of gradient disappearance and ensures the effective training of the deep network. After being processed layer by layer by the four residual module groups, the low-level features of the input image are gradually abstracted into high-level semantic features.

[0034] In the third step, a feature representation module based on feature diffusion enhancement is constructed. This module aims to break the spatial connection between pixels, guide the feature representation to transition to high-density regions, and realize the mining and enhancement of pixel-level essential features. The core of constructing a feature representation module based on feature diffusion enhancement lies in decoupling pixel-level spatial dependence and guiding the feature flow to high-density regions, aiming to overcome the over-reliance on spatial context in traditional network feature representation learning, improve the essentiality of features, and enhance the robustness to complex scenarios. Feature diffusion enhancement includes the following:

[0035] Forward diffusion and noise injection strategy: Drawing on the idea of diffusion models, the forward diffusion stochastic differential equation is used to gradually inject noise into the feature matrix Z, and the formula is as follows:

[0036]

[0037] where, dw~N(0,1) represents the Wiener process of random noise, which is a Gaussian noise term with randomness. dt is a continuous variable representing a small change in time. β(t) is a weight parameter that changes uniformly with time. When t = 0 initially, Z(t) is not affected by noise; when t approaches 1, the noise gradually increases, the perturbation gradually accumulates, and Z(t) gradually approaches Gaussian noise.

[0038] Graph convolutional neural network-guided scoring function and feature enhancement: The noise at each moment during the noise addition process has randomness, and this kind of random noise cannot effectively guide the denoising process to obtain a stable feature recovery result. To obtain a more stable denoising effect, the gradient of the overall likelihood function changing from 0 to t moments is used To provide the optimal direction from the noisy feature Z(t) to the noise-free feature Z(0), ensuring that the denoising result under this gradient guidance can not only maintain the distribution characteristics of the original data but also improve the distribution density of similar targets in the feature space to a certain extent, enhancing the intra-class feature similarity. However, due to the randomness introduced by noise affecting the probability density gradient distribution of feature Z(t) under the condition of Z(0), the uncertainty in the cumulative distribution in the feature space increases significantly. This randomness leads to large fluctuations in gradient estimation during the denoising process, which is thus used as a reference for supervising the learning of the scoring function. Therefore, this paper constructs a scoring function ρ(Z(t),t) based on a graph convolutional neural network to regress the gradient and designs a loss function that can obtain a stable scoring function, ultimately providing the direction for unlabeled samples to gradually aggregate towards labeled samples of the same class. The loss function is defined as follows:

[0039]

[0040] where, logp 0|t (Z(t)|Z(0)) represents the gradient of the log-likelihood function of the conditional probability density function with respect to Z(t), which represents the gradient of the probability density distribution of feature Z(t) at time t with respect to Z(t). This gradient can be understood as the "guiding direction". E Z(t)|Z(0) [·] represents taking the expectation under the conditional probability distribution of Z(t) given the initial feature Z(0). Under the guidance of Bayes' theorem, further taking the expectation of E Z(t)|Z(0) [·] given the condition of Z(0) to ensure that the estimated scoring function can be free from the interference of random noise and initial offsets and obtain a stable gradient estimation; finally, taking the expectation of the losses at all times to enable the scoring function to adapt to different degrees of noise addition within the entire time range.

[0041] Reverse diffusion and revealing feature structure: To achieve feature denoising, the inverse stochastic differential equation is used to gradually eliminate noise and restore the original feature state under the guidance of the scoring function gradient. The inverse stochastic differential equation consists of two key parts: time reverse evolution and a correction term based on the scoring function, which can be expressed as follows:

[0042]

[0043] The model iteratively goes backward from time t to 0 step by step, and at each step, it estimates the denoising gradient direction of the feature in the current state by calculating the scoring function ρ(Z(t),t). Through iterative denoising operations, we focus on restoring the inherent representation of each pixel rather than relying on a globally coupled spatial mapping, effectively breaking the inherent spatial dependence between pixels. This pixel-gradient-based method, combined with a block-based neural learning strategy, enhances the high-density representation of each class in the latent feature space.

[0044] In the fourth step, to guide the pixel-by-pixel learning of high-density compact representations in the first step, a clustering loss function is constructed based on the neighborhood consistency assumption to reduce intra-class differences and thus improve inter-class separability. This guidance enables the student model to learn more robustly from complex and noisy images. The RBF function is used to estimate the similarity between two latent feature representations z i and z j :

[0045]

[0046] where z i and z j represent the feature representations of samples x i and x j by the feature extraction network, ||z i -z j ||2 represents the Euclidean distance between samples x i and x j in the feature space, reflecting their similarity. σ represents the decay rate controlling the similarity. τ is the distance threshold, and when the distance between samples exceeds τ, they are considered dissimilar, and at this time ω i,j = 0.

[0047] To further improve the compact representation of sample data, the following clustering-based loss function is designed:

[0048]

[0049] In the fifth step, to improve the distillation effect at the level of model prediction results, a dynamic masking strategy based on the confidence level of the teacher model is proposed. By masking the probabilities of some categories in the teacher model output prediction, the masking operation is implemented through the mask M T as follows:

[0050]

[0051] For the teacher probability distribution p T , mask out those categories for which the prediction probability of the teacher model for that category is greater than the probability of the true category, and only pass on the categories for which the prediction probability of the category is less than or equal to the probability of the true category to the student model. After obtaining the mask, the following formula is used to calculate the new probability distribution:

[0052]

[0053] After obtaining the mask guidance for distillation, the distillation loss is defined as follows:

[0054] L DM = ψ 2 ·KL(pT , p S )

[0055] Step 6. For each sample x i Rank the output score coefficients of the distillation model to obtain the corresponding rankings of the models respectively. Subsequently, calculate the ranking differences between the predictions of the distillation models. To avoid the cancellation of positive and negative rankings, square the ranking differences for each sample. Subsequently, define an index to measure the consistency of the output rankings of the two, and construct a loss function through it to maximize the ranking correlation between the distillation models. By minimizing the ranking differences between the distillation models, drive the student model to gradually learn the sorting pattern of the teacher model. This loss function is defined as:

[0056]

[0057] Step 7. Repeat the above training steps until the preset number of iterations, and output the image classification result.

[0058] As described above, only the preferred specific embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A semi-supervised learning method based on deep compact feature representation and dynamic rank correlation guidance, characterized in that include: Constructing a semi-supervised learning model, the model comprising a teacher network and a student network, the teacher network is used to generate pseudo labels and guidance information, and the student network and the teacher network share some parameter structures; Use labeled data to input the network for preliminary training to obtain the output prediction corresponding to the labeled samples, and use the labeled and unlabeled data to generate the predicted probability distribution through the teacher network; Based on the data, deep features are subjected to noise injection and denoising processing through a feature diffusion module to break the global or local coupling relationship between pixels and obtain a high-density compact feature representation; The clustering consistency module is used to bring similar samples in the feature space together through adaptive metrics, which can simultaneously reduce the feature differences between labeled and unlabeled samples and enhance the intra-class compactness. Based on the dynamic mask and rank-related module, the prediction distribution of the teacher network output is first dynamically masked to filter out high-confidence erroneous predictions. Then, the rank order difference between the teacher and student outputs is minimized to further guide the student network to finely learn the classification tendency of the teacher model. Repeat steps (3) to (5) until the preset convergence conditions are met, and use the optimized student network to achieve semi-supervised feature representation and classification tasks.

2. The method according to claim 1, wherein The feature extraction of the model adopts an improved ResNet structure to obtain an intermediate feature representation with decoupling capability.

3. The method according to claim 1, characterized in that, The feature diffusion module adopts a forward noise injection and reverse denoising and reconstruction process, wherein the forward process gradually introduces Gaussian noise according to a stochastic differential equation, and the reverse process gradually restores the original feature structure under the guidance of the noise gradient to achieve pixel-level dispersion and reconstruction.

4. The method according to claim 1, characterized in that, The method further includes constructing a clustering consistency module, which calculates the Euclidean distance between each sample in the feature space and weights the sample similarity based on the radial basis function, thereby forming a clustering loss based on the neighborhood consistency constraint, so as to shorten the feature distance between samples of the same type and expand the interval between samples of different types.

5. The method according to claim 1, wherein In the clustering consistency module, the clustering loss includes two parts: one is to minimize the Euclidean distance between the labeled data and its adjacent unlabeled data, and the other is to minimize the Euclidean distance between adjacent samples in the unlabeled data, so as to achieve tight intra-class and inter-class separation.

6. The method according to claim 1, wherein The generation of the dynamic mask and rank-related module includes: after calculating the one-hot encoding of the data prediction probability based on the teacher network, performing convolution downsampling and softmax normalization processing to obtain the prediction probability of each image area block; the predicted probability is sorted by information entropy, and a fixed proportion of area blocks are selected as candidate areas, and then a dynamic mask is generated by random sampling of Bernoulli distribution based on preset transformation coefficients; the obtained dynamic mask is used to screen the output of the teacher network, and the prediction distribution of the remaining areas is compared with the output of the student network for rank sorting consistency.

7. The method according to claim 6, characterized in that, The dynamic mask and rank correlation module further includes: by performing rank sorting on the predictions output by the teacher network and the student network, calculating the sum of the squares of the rank differences between the two for each category, and then constructing a rank correlation loss based on the rank correlation coefficient to quantify the difference in the prediction consistency between the teacher and the student networks, thereby guiding the student network to capture the relative ranking information output by the teacher.

8. The method according to claim 1, wherein The overall optimization objective of the model is: L = L CE + λ S L S + λ lc L lc + λ DM L DM + λ RC L RC Among them, L CE is the cross-entropy loss used to supervise the labeled data; L S is the gradient regression loss introduced by feature diffusion; L lc is the clustering neighborhood consistency loss; L DM is the dynamic mask loss; L RC is the rank correlation-based consistency loss.

9. The method according to claim 8, wherein The cross-entropy loss L CE is respectively for the true labels of the labeled data and for the teacher network pseudo-labels in the dynamic mask screening process to implement the teacher mask.

10. The method according to claim 1, characterized in that, Through the collaborative work of the above-mentioned modules, the efficient utilization of unlabeled data in semi-supervised learning is achieved, that is, while performing feature diffusion and noise reduction processing and enhancing the compactness of sample clustering, the interference of incorrect predictions of the teacher network is effectively suppressed through the dynamic mask and rank correlation strategy, thereby overall improving the classification accuracy and robustness of the model under data-scarce conditions.

Citation Information

Cited By

  • Self-adaptive asymmetric image steganography method and system based on diffusion model

    CN121151516A

  • An Adaptive Asymmetric Image Steganography Method and System Based on a Diffusion Model

    CN121151516B