Distillation diffusion model-based open set hyperspectral image instance map construction method

The example diagram of open set hyperspectral image is constructed through the distillation diffusion model, which solves the problem of missing prior information of unknown classes in cross-domain open set recognition, realizes accurate estimation of unknown class probability and learning of known class boundary relationships, and improves the recognition accuracy and robustness of the model.

CN120298880APending Publication Date: 2025-07-11CHINA UNIV OF MINING & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510289986.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the cross-domain open set recognition task, the existing image recognition model is difficult to effectively learn the complex boundary relationship between unknown classes and known classes due to the lack of prior information of unknown classes in the cross-domain open set recognition task. The traditional uncertainty quantization mechanism relies on fixed thresholds and lacks flexibility, making it difficult to adapt to complex real-life scenarios.

Method used

Using a method based on distillation diffusion model, the uncertainty of model prediction is quantified through uncertainty feature acquisition, uncertainty distribution modeling and graph isomorphic knowledge distillation, and an open set hyperspectral image example diagram is constructed to realize the accurate estimation of the probability of unknown classes and the learning of the boundary relationship between known classes and unknown classes.

Benefits of technology

Accurate quantification of model prediction uncertainty is achieved, the accuracy and robustness of cross-domain open set image recognition are improved, and the knowledge transfer efficiency and classifier consistency are improved through the graph isomorphic knowledge distillation optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298880A_ABST
    Figure CN120298880A_ABST
Patent Text Reader

Abstract

The invention discloses an open set hyperspectral image instance graph construction method based on a distillation diffusion model, and the method comprises the following steps: extracting the spatial and spectral features of a target domain HSI through a spatial encoder and a spectral encoder, executing forward diffusion in a diffusion classifier by taking a real label as a starting point, and carrying out the condition reverse denoising, and finally, closed set category prediction is obtained. A single sample prediction set is generated through Monte Carlo sampling, a prediction variance is calculated, uncertainty distribution is constructed, an unknown class probability is calculated by using a cumulative distribution function, and finally, a known class probability and an unknown class probability are fused to generate teacher open set prediction through a mapping module. A K neighbor graph structure of a sample is constructed in a feature space, teacher open set prediction and student open set prediction are respectively used as nodes to construct a teacher / student instance graph, a Wasserstein distance of prediction probability distribution of the teacher / student instance graph and the student open set prediction is minimized through a graph isomorphic knowledge distillation module, and knowledge distillation based on prediction manifold alignment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a method for constructing an open-set hyperspectral image instance graph based on a diffusion model. Background Art

[0002] Cross-domain open-set image recognition can accurately identify unknown classes in the target domain, which is of great significance for improving the security and reliability of image recognition models in real-world scenarios. However, the difficulty in achieving cross-domain open-set recognition lies in that existing image recognition models need to minimize their empirical risk on training data during the training process, that is, minimize the difference between their prediction results and the true labels of training samples. It should be noted that in the cross-domain open-set recognition task, since unknown classes are unknown during the model training process, the lack of prior knowledge of unknown classes during training poses a severe challenge to the learning of unknown class recognition strategies.

[0003] Currently, the mainstream solution is to consider that the predictions of models for unknown classes often have high uncertainty. By establishing a quantization mechanism for model prediction uncertainty and combining a threshold-based scheme, unknown class recognition is achieved. Based on this, the uncertainty quantization mechanism can determine whether a sample belongs to an unknown class by calculating the entropy value of the predicted probability distribution or based on a confidence threshold. However, this uncertainty quantization mechanism often relies on prior assumptions about the prediction accuracy of the model. In the context of cross-domain open-set image recognition, the model may make incorrect predictions due to the presence of noise patterns similar to known classes in unknown classes, which limits the effectiveness of such methods. In addition, the threshold-based scheme lacks flexibility, and the model performance highly depends on the setting of the prior threshold. In practical applications, the threshold needs to be finely tuned for different tasks, which is usually time-consuming, laborious, and costly. The noise samples in the target domain or the ambiguity of the class boundary may further amplify the uncertainty of the model, making the threshold-based scheme difficult to adapt to complex real-world scenarios.

[0004] To address the above problems, first of all, it is urgent to develop a more refined uncertainty quantification mechanism to achieve an accurate estimate of the probability that a sample belongs to an unknown class. In uncertainty reasoning, the cognitive process of humans provides important theoretical inspiration: in the face of uncertainty, human speculation often has differences, that is, for the same problem, different backgrounds, experiences, and knowledge will lead to diverse prediction results. When there are significant differences among multiple speculations, it reflects that the understanding of this problem is not sufficient, indicating a high degree of uncertainty; on the contrary, when the results of different speculations tend to be consistent, it means that the uncertainty is low and the cognition is more reliable. In model prediction, this difference can be used to quantify the prediction uncertainty by evaluating the range of changes in the results when the model makes multiple predictions on the same instance. This method not only provides a "point estimate" for the model but also allows the model to generate an accurate estimate of the prediction credibility or confidence interval. However, traditional classifiers usually adopt a fixed prediction mode and often only output a single and definite prediction result for the same sample, and cannot effectively model the uncertainty reflected by this speculation difference. Summary of the Invention

[0005] Aiming at the problems existing in the above-mentioned background technology, the purpose of the present invention is to provide a method for constructing an instance graph of an open-set hyperspectral image based on a diffusion model. Aiming at the problem that it is difficult for the model to learn the complex boundary relationship between unknown classes and known classes due to the unknown prior information of unknown classes during training in the cross-domain open-set image recognition task, a solution based on uncertainty distribution modeling and graph isomorphism knowledge distillation is provided.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A method for constructing an instance graph of an open-set hyperspectral image based on a diffusion model, characterized in that it includes the following steps:

[0008] Step 1, Uncertainty feature acquisition: Extract the spatial and spectral features of the target domain HSI through a spatial encoder and a spectral encoder, perform forward diffusion with the true label as the starting point in the diffusion classifier, and perform conditional reverse denoising based on the sample features to finally obtain the closed-set class prediction;

[0009] Step 2, Uncertainty distribution modeling: Generate a prediction set for a single sample through Monte Carlo sampling and calculate the prediction variance, construct an uncertainty distribution based on the prediction variances of all samples, use the cumulative distribution function to calculate the probability of unknown classes, and finally fuse the probabilities of known classes and unknown classes to generate a teacher open-set prediction through a mapping module;

[0010] Step 3, Graph Isomorphism Knowledge Distillation: Construct a K-nearest neighbor graph structure of samples in the feature space, and construct teacher / student instance graphs with the teacher open-set prediction and the student open-set prediction as nodes respectively. Minimize the Wasserstein distance between the prediction probability distributions of the two through the graph isomorphism knowledge distillation module to achieve knowledge distillation based on the alignment of prediction manifolds.

[0011] Further, the said Step 1 includes:

[0012] Step 11, Input the target domain HSI into the feature extractor, and use the spatial encoder and the spectral encoder to capture the spatial dependence between pixels and the spectral dependence between bands to obtain sample features;

[0013] Step 12, Input the sample features into the diffusion classifier. During the forward diffusion process, the diffusion classifier gradually adds Gaussian noise to the true label to obtain y T ;

[0014] Step 13, During the reverse denoising process, the diffusion classifier uses the sample features as conditional signals to gradually denoise the noise distribution to obtain the closed-set class prediction

[0015] Further, the said Step 2 includes:

[0016] Step 21, Obtain the prediction set for a single sample by performing Monte Carlo sampling in the noise distribution, and calculate the prediction variance in the set;

[0017] Step 22, Regard the variance of the single-sample prediction as a point in the uncertainty distribution, and model the uncertainty distribution from the prediction distributions of a group of samples. By calculating the cumulative distribution function (CDF) of each sample in the uncertainty distribution, obtain its cumulative distribution probability (CDP), and further estimate the unknown class probability of the sample;

[0018] Step 23, Input the known class probability and the known class probability of the diffusion classifier into the probability mapping module to obtain the teacher open-set prediction (TOP).

[0019] Further, the said Step 3 includes:

[0020] Step 31, Search for the K-nearest neighbors of the samples in the feature space to obtain the edges between the samples;

[0021] Step 32, Regard the sample and the teacher open-set prediction corresponding to its K-nearest neighbors as nodes, and establish edge connections between them to construct the teacher instance graph;

[0022] Step 33, Also perform open-set prediction on these samples based on the open-set student classifier, and construct the student instance graph based on the prediction results;

[0023] Step 34: Input the student instance graph and the teacher instance graph into the graph isomorphism knowledge distillation module. By minimizing the Wasserstein distance between them, constrain the consistent prediction manifolds of the student classifier and the teacher classifier to achieve knowledge distillation between the student classifier and the teacher classifier.

[0024] Further, in the above Step 1, the forward diffusion process includes:

[0025] The forward diffusion process of the diffusion classifier is disassembled as:

[0026]

[0027] where p(y 0:T ) represents the joint probability of all states of label y occurring from time 0 to T, q(y0) is the prior probability of state y0, and q(y t |y t-1 ) represents the conditional probability of the current state y t-1 given the previous state y t ;

[0028] In the forward diffusion process, given an observed label y0, the following conditional probability is obtained:

[0029]

[0030] where p(y 1:T |y0) represents the joint probability of all states of label y occurring from time 1 to T given y0 is known, and q(y1|y0) represents the conditional probability of state y1 given y0 is known;

[0031] In the framework of the diffusion model, the encoder q(y t |y t-1 ) at each step is fixed to a linear Gaussian transformation, that is, q(y t |y t-1 ) is a Gaussian distribution with as the mean and as the variance. Therefore, using the reparameterization trick, y t-1 is used to obtain y t :

[0032]

[0033] where α t is a hyperparameter used to control the strength of the noise, ε is a random variable following a standard normal distribution (mean 0, variance 1), represents the standard normal distribution, and I represents a variance of 1;

[0034] Based on the above recurrence relation, starting from y0, directly calculate y at any step t :

[0035]

[0036] Among them, It can be seen from this that due to α i ∈(0, 1), y t gradually approaches the standard Gaussian distribution as the time step increases.

[0037] Furthermore, in the above step 1, the reverse denoising process includes:

[0038] The goal of the reverse denoising process is to start from the standard Gaussian distribution and the decoder performs step-by-step denoising on y T conditioned on the sample feature z to recover the true label distribution. Therefore, the joint probability p(y 0:T ) is decomposed as:

[0039]

[0040] Among them, the probability density of p(y T ) is known and is a standard Gaussian distribution, that is However, p(y t |y t+1 , z) is difficult to directly calculate. Therefore, here a neural network is used to learn p θ (y0|y1, z) to approximate the true distribution, where θ is the neural network parameter.

[0041] Furthermore, in the above step 1, the learning objectives in the uncertainty feature acquisition process include:

[0042] The goal of the diffusion classifier is to model p(y0|z), where: y0 is the true label; z = f(x), f(·) is the feature extractor, and its label generation process is divided into two stages. In the forward diffusion stage, noise is gradually added to the label y0 and finally transformed into the standard Gaussian distribution; in the reverse denoising stage, through the conditional input z and the neural network-parameterized reverse process, denoising is gradually performed from the Gaussian distribution to recover to the target label y0;

[0043] The learning objective of the diffusion classifier is to maximize the log-likelihood of the observed data by optimizing the following evidence lower bound to learn the parameter θ:

[0044]

[0045] Among them, the first term is intended to minimize the reconstruction loss of the true label conditioned on the sample feature z, and the second term DKL (p(y T )||q(y T |y0)) aims to constrain the distribution of the end point y T of the forward diffusion to be close to the standard Gaussian distribution. The third term D KL (p θ (y t-1 |y t ,z)||q(y t-1 |y t ,y0)) aims to minimize the difference between the distribution p θ (y t |y t+1 ,z) learned by the parameter θ and the true distribution q(y t-1 |y t ,y0); Since the second term, the forward diffusion process, is defined as a linear Gaussian transformation, this term has no optimizable parameters during training and focuses on the optimization of the first and third terms;

[0046] First, considering that p θ (y0|y1,z) is a conditional Gaussian distribution depending on y1 and z. Let its mean μ θ be a parametric function of y1 and z, denoted as μ θ (y1,z,t = 1), and the variance is a constant, denoted as Σ; Expand lnp θ (y0|y1,z) to get:

[0047]

[0048] where exp represents the natural exponential function and n represents the dimension of the data;

[0049] Therefore, maximizing is equivalent to minimizing the mean squared error between the output of the diffusion classifier and y0;

[0050] Then, considering that in the third term, q(y t-1 |y t ,y0) conforms to a Gaussian distribution with mean μ q (y t ,y0) and variance Σ q (t); To make p θ (y t |y t+1 ,z) better fit q(y t-1 |y t ,y0), here p θ (y t |y t+1 ,z) is also regarded as a Gaussian distribution, using μ θ (y t, z) represents the mean of this distribution. Since q(y t-1 |y t , y0), the variance term is only related to t. Therefore, let p θ (y t |y t+1 , z) correspond to the variance Σ θ (t) = Σ q (t);

[0051] Next, obtain the KL divergence between p θ (y t |y t+1 , z) and q(y t-1 |y t , y0):

[0052]

[0053] Among them, represents the variance of the q distribution corresponding to time t;

[0054] It can be seen from the above formula that minimizing the third term is equivalent to minimizing the difference between μ θ (y t , z) and μ q (y t , y0);

[0055] Next, to reduce the learning difficulty of the model, μ θ (y t , z) needs to be parameterized in the form of μ q (y t , y0), and μ q (y t , y0) is expressed in the following form:

[0056]

[0057] Among them, α t is a hyperparameter used to control the strength of the noise,

[0058] Therefore, μ θ (y t , z) is expressed as:

[0059]

[0060] Among them, is the output of the parameterized neural network;

[0061] Therefore, substitute μ q (y t , y0) and into DKL (p θ (y t-1 |y t ,z)||q(y t-1 |y t ,y0))), we get:

[0062]

[0063] Among them, ε represents the random Gaussian noise added to y during the forward diffusion process from the (t - 1)-th time step to the t-th time step; t-1 of.

[0064] In summary, the final optimization objective of the diffusion classifier is expressed as:

[0065]

[0066] Among them, γ is the trade-off coefficient. The first term of this loss hopes that the predicted label after reverse denoising is consistent with the true label, and the second term expects the diffusion classifier to gradually recover the true label distribution from the noise distribution at each time step.

[0067] Furthermore, in step 2, the uncertainty distribution modeling process includes:

[0068] Considering that due to the lack of supervision information for unknown classes, the diffusion classifier can only give closed-set predictions for samples. Therefore, through uncertainty distribution modeling, the closed-set predictions are converted into open-set predictions;

[0069] First, since in the diffusion classifier, each step of the reverse denoising process contains randomness, especially the process of gradually denoising starting from standard Gaussian noise; therefore, using Monte Carlo sampling, by sampling M reverse denoising trajectories calculate the variance of the class prediction probability distribution:

[0070]

[0071] Among them, is the indicator function, with a value of 1 indicating belonging to class c, otherwise 0; μ c represents the mean of the prediction probability for class c, represents the variance of the prediction probability for class c, represents the predicted label of the i-th sample;

[0072] Next, based on the variances of the predictions for each class of feature z, quantify the overall uncertainty of the classifier:

[0073]

[0074] Among them, C represents the total number of classes, and c refers to a specific class;

[0075] Next, obtain the uncertainty variance values {U(z1),..., U(z q ,..., z Q )} for each sample from a set of sample feature sets {z1,..., z q ),..., U(z Q )}; assume that these variance values follow a Gaussian distribution, and use these values to fit the parameters of the uncertainty distribution to obtain the uncertainty distribution N(μ u , (σ u ) 2 ):

[0076]

[0077] where Q represents the total number of a set of sample feature sets, and q represents a specific sample feature;

[0078] After that, for each sample feature z q , by calculating the cumulative probability of the uncertainty variance of this sample in the uncertainty distribution, obtain the probability that it belongs to an unknown class:

[0079]

[0080] where ξ q represents a random integral variable, μ u represents the mean value obtained by normalizing the sample feature z q , and σ u represents the variance;

[0081] Finally, use the probability mapping module to convert the closed-set prediction to an open-set prediction:

[0082]

[0083] where is the open-set prediction, PMM(·) is the probability mapping module, represents the sum of the predicted means for each class; the probability mapping operation in the probability mapping module is expressed as:

[0084]

[0085] where || is the concatenation operation, and Softmax is an activation function;

[0086] Finally, to ensure that the uncertainty variance truly reflects the uncertainty of the model, use the following uncertainty calibration regularization to constrain the model to have low uncertainty on the correct classes and high uncertainty on the wrong classes. For a single sample, its uncertainty calibration regularization can be calculated by the following formula:

[0087]

[0088] Among them, is an indicator function, with a value of 1 indicating that the corresponding condition is met, otherwise 0.

[0089] Furthermore, in step 3, the instance graph construction in the graph isomorphism knowledge distillation process includes:

[0090] Let the sample feature set in the feature space be Z = {z1,..., z n ,..., z N}, and each sample feature z n corresponds to a class prediction, where N is the number of samples; it is desired to search for the K-nearest neighbor set of each sample feature z n in the feature space and construct an instance graph G = (V, E); among them, the node set V contains the sample features and the class predictions of their K-nearest neighbors, and the edge set E reflects the association attributes between the samples and their K-nearest neighbors;

[0091] First, in the feature space, the clustering between node v i and node v j is defined as:

[0092] d(v i , v j ) = ||v i - v j ||2

[0093] For each node v i , its K-nearest neighbor set is expressed as:

[0094]

[0095] Among them, are the indices of the nearest K samples sorted in ascending order of d(v i , v j );

[0096] Then, an edge is established between v i and each node v j ∈ N K (v i ), and the edge set is defined as:

[0097]

[0098] Finally, based on the sample features and the class predictions of each sample by the teacher classifier and the student classifier, where the teacher classifier is a closed-set diffusion classifier and the student classifier is an open-set classifier; construct the instance graph G corresponding to the teacher classifier based on the above graph construction principleTe = (V Te , E Te ) and the instance graph G corresponding to the student classifier Stu = (V Stu , E Stu ), where V Te represents the set of nodes in the teacher instance graph, and E Te represents the set of edges. V Stu represents the set of nodes in the student instance graph, and E Stu represents the set of edges.

[0099] Furthermore, in step 3, in the graph isomorphism knowledge distillation process, a graph isomorphism knowledge distillation module based on the instance graph is proposed. The structural associations between samples are characterized by the instance graph, and the Wasserstein distance is further introduced to achieve graph isomorphism, transforming the graph structure optimization problem into a distribution alignment problem;

[0100] For the sets of nodes V Te and V Stu in the teacher instance graph and the student instance graph, their distributions are defined as:

[0101]

[0102] where and represent the node features in the teacher instance graph and the student instance graph respectively. δ(·) is the Dirac function, used to represent the discrete distribution, and K represents the number of samples;

[0103] The Wasserstein distance between the teacher instance graph distribution and the student instance graph distribution is expressed as:

[0104]

[0105] where is the set of all joint distributions satisfying the marginal distributions and , expressed as:

[0106]

[0107] A Wasserstein distance metric Γ ω (·) based on neural network parameterization is introduced to approximate φ(·), where φ(·) is the learnable neural network parameter. Therefore, the optimization objective is expressed as:

[0108]

[0109] where represents the variable v TeSubject to a distribution When, through the function Γ ω (·) The calculated expectation Indicates that the variable v Stu Subject to a distribution When, through the function Γ ω (·) The calculated expectation

[0110] To ensure that Γ ω (·) Satisfies the 1-Lipschitz condition, the following gradient penalty term is introduced:

[0111]

[0112] Among them, Is the distribution generated by interpolation , α ∼ U(0,1) Represents the variable Subject to a distribution When, the expectation of the function calculation result

[0113] Therefore, the graph isomorphism knowledge distillation loss is expressed as:

[0114]

[0115] Among them, λ is the gradient penalty weight

[0116] Finally, in the graph isomorphism knowledge distillation process, by minimizing the Wasserstein distance loss between distributions and imposing the constraint of the 1-Lipschitz condition, the student distribution Gradually approaches the teacher distribution Close; in this way, the fine-grained knowledge about the sample structure association hidden in the teacher classifier is distilled into the student classifier to ensure the consistent prediction manifold of the student classifier and the teacher classifier

[0117] Beneficial effects: Aiming at the limitation that the fixed prediction of the traditional classification mechanism is difficult to accurately quantify the model prediction uncertainty, a label generation mechanism based on the conditional diffusion model is developed. By directly modeling the prediction probability distribution, the accurate quantification of the prediction uncertainty of each category of the model is realized, providing a new paradigm for the uncertainty quantification of the model. A graph isomorphism optimization mechanism is proposed, which models the knowledge distillation process as a graph isomorphism optimization problem of instance graphs. By aligning the prediction manifolds of the teacher and student classifiers, the efficiency and robustness of knowledge transfer between the teacher and student classifiers are improved Brief description of the drawings

[0118] Figure 1 Is the principle block diagram of the method of the present invention Detailed implementation manners

[0119] The present invention will be further described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0120] The method for constructing an open-set hyperspectral image instance graph based on the distilled diffusion model provided by the present invention has the following specific principle. Figure 1 As shown, first, a diffusion classifier is designed. By performing a Monte Carlo sampling strategy on the noise distribution during the reverse denoising process to quantify the model prediction uncertainty, and through uncertainty distribution modeling and cumulative distribution function calculation, the unknown class probability of the sample is obtained. Then, an open-set recognition framework based on knowledge distillation is established. The diffusion classifier of the closed set is regarded as the teacher classifier, and the complex boundary relationship between the known class and the unknown class is guided for the student classifier to learn through the way of knowledge distillation. Finally, the knowledge distillation process is formalized as a graph isomorphism optimization problem, and the knowledge transfer between classifiers is promoted by constraining the consistent prediction manifolds of the student classifier and the teacher classifier.

[0121] The method for constructing an open-set hyperspectral image instance graph based on the distilled diffusion model of the present invention specifically includes the following steps:

[0122] Step 1, obtaining uncertainty features: Extract the spatial and spectral features of the target domain HSI through the spatial encoder and the spectral encoder. Perform forward diffusion in the diffusion classifier starting from the true label, and perform conditional reverse denoising based on the sample features to finally obtain the closed-set class prediction. Specifically, it includes:

[0123] Step 11, input the target domain HSI into the feature extractor, and use the spatial encoder and the spectral encoder to capture the spatial dependence relationship between pixels and the spectral dependence relationship between bands to obtain sample features;

[0124] Step 12, input the sample features into the diffusion classifier. During the forward diffusion process, the diffusion classifier gradually adds Gaussian noise to the true label to obtain y T ;

[0125] Step 13, during the reverse denoising process, the diffusion classifier uses the sample features as conditional signals to gradually denoise the noise distribution to obtain the closed-set class prediction

[0126] Among them, the forward diffusion process includes:

[0127] The forward diffusion process of the diffusion classifier is disassembled into:

[0128]

[0129] where p(y 0:T ) represents the joint probability of all states of label y occurring from time 0 to T, q(y0) is the prior probability of state y0, and q(y t |y t-1 ) represents the conditional probability of the current state y t-1 given the previous state y t ;

[0130] In the forward diffusion process, given a label observation y0, the following conditional probabilities are obtained:

[0131]

[0132] where p(y 1:T |y0) represents the joint probability of all states of label y occurring from time 1 to T given y0, and q(y1|y0) represents the conditional probability of state y1 given y0;

[0133] In the framework of the diffusion model, the encoder q(y t |y t-1 ) at each step is fixed to a linear Gaussian transformation, that is, q(y t |y t-1 ) is a Gaussian distribution with as the mean and as the variance. Therefore, using the reparameterization trick, y t-1 is used to obtain y t :

[0134]

[0135] where α t is a hyperparameter used to control the strength of the noise, ε is a random variable of a standard normal distribution, represents the standard normal distribution, and I represents a variance of 1;

[0136] Based on the above recurrence relation, starting from y0, y t at any step is directly calculated:

[0137]

[0138] where It can be seen from this that since α i ∈(0,1), y t gradually approaches the standard Gaussian distribution as the time step increases.

[0139] Among them, the reverse denoising process includes:

[0140] The goal of the reverse denoising process is to start from the standard Gaussian distribution At the beginning, the decoder conditions on the sample feature z and progressively denoises y T to recover the true label distribution. Therefore, the joint probability p(y 0:T ) is factorized as:

[0141]

[0142] where the probability density of p(y T ) is known and is a standard Gaussian distribution, i.e., However, p(y t |y t+1 ,z) is difficult to directly calculate. Therefore, here a neural network is used to learn p θ (y0|y1,z) to approximate the true distribution, where θ are the neural network parameters.

[0143] Among them, the learning objectives in the uncertainty feature acquisition process include:

[0144] The goal of the diffusion classifier is to model p(y0|z), where: y0 is the true label; z = f(x), f(·) is the feature extractor, and its label generation process is divided into two stages. In the forward diffusion stage, noise is gradually added to the label y0 and finally transformed into a standard Gaussian distribution; in the reverse denoising stage, through the conditional input z and the neural network-parameterized reverse process, denoising is gradually performed from the Gaussian distribution to recover to the target label y0;

[0145] The learning objective of the diffusion classifier is to maximize the log-likelihood of the observed data by optimizing the following evidence lower bound to learn the parameter θ:

[0146]

[0147] Among them, the first term aims to minimize the reconstruction loss of the true label conditional on the sample feature z, and the second term D KL (p(y T )||q(y T |y0)) aims to constrain the distribution of the end point y T to be close to the standard Gaussian distribution, and the third term D KL (p θ (y t-1 |y t ,z)||q(y t-1 |y t ,y0)) aims to minimize the distribution p θ (y t |y t+1 ,z) learned by the parameter θ and the true distribution q(y t-1 |y t, the difference between y0); since the second positive diffusion process is defined as a linear Gaussian transformation, there are no optimizable parameters in this term during training, and the focus is on optimizing the first and third terms;

[0148] First, considering p θ (y0|y1,z) is a conditional Gaussian distribution that depends on y1 and z. Let its mean be μ θ which is a parametric function of y1 and z, denoted as μ θ (y1,z,t = 1), and the variance is a constant, denoted as Σ; expanding lnp θ (y0|y1,z) gives:

[0149]

[0150] where exp represents the natural exponential function and n represents the dimension of the data;

[0151] Therefore, maximizing is equivalent to minimizing the mean squared error between the output of the diffusion classifier and y0;

[0152] Then, considering that in the third term q(y t-1 |y t ,y0) conforms to a Gaussian distribution with mean μ q (y t ,y0) and variance Σ q (t); to make p θ (y t |y t+1 ,z) better fit q(y t-1 |y t ,y0), here p θ (y t |y t+1 ,z) is also regarded as a Gaussian distribution, and the mean of this distribution is represented by μ θ (y t ,z). Since in q(y t-1 |y t ,y0), the variance term is only related to t, so here let p θ (y t |y t+1 ,z) correspond to the variance Σ θ (t) = Σ q (t);

[0153] Next, obtain the KL divergence between p θ (y t |y t+1 ,z) and q(y t-1 |y t ,y0):

[0154]

[0155] Among them, represents the variance of the q-distribution corresponding to time t;

[0156] It can be seen from the above formula that minimizing the third term is equivalent to minimizing μ θ (y t , z) and the difference between μ q (y t , y0);

[0157] Next, to reduce the learning difficulty of the model, μ θ (y t , z) needs to be parameterized in the form of μ q (y t , y0), and μ q (y t , y0) is expressed in the following form:

[0158]

[0159] Among them, α t is a hyperparameter used to control the strength of the noise,

[0160] Therefore, μ θ (y t , z) is expressed as:

[0161]

[0162] Among them, is the output of the parameterized neural network;

[0163] Therefore, substituting μ q (y t , y0) and into D KL (p θ (y t-1 |y t , z)||q(y t-1 |y t , y0)), we get:

[0164]

[0165] Among them, ε represents the random Gaussian noise added to y t-1 during the forward diffusion process from the (t - 1)-th time step to the t-th time step;

[0166] In summary, the final optimization objective of the diffusion classifier is expressed as:

[0167]

[0168] Among them, γ is a trade-off coefficient. The first term of this loss hopes that the predicted label after reverse denoising is consistent with the true label, and the second term expects the diffusion classifier to gradually recover the true label distribution from the noise distribution at each time step.

[0169] Step 2, Uncertainty distribution modeling: Generate a prediction set for a single sample through Monte Carlo sampling and calculate the prediction variance. Construct an uncertainty distribution based on the prediction variances of all samples, use the cumulative distribution function to infer the unknown class probability, and finally fuse the known class and unknown class probabilities to generate a teacher open-set prediction through a mapping module. Specifically, it includes:

[0170] Step 21, Obtain a prediction set for a single sample by performing Monte Carlo sampling in the noise distribution and calculate the prediction variance in the set;

[0171] Step 22, Consider the variance of the single-sample prediction as a point in the uncertainty distribution, and model the uncertainty distribution from the prediction distributions of a group of samples. By calculating the cumulative distribution function (CDF) of each sample in the uncertainty distribution, obtain its cumulative distribution probability (CDP), and then estimate the unknown class probability of the sample;

[0172] Step 23, Input the known class probability and the known class probability of the diffusion classifier into the probability mapping module to obtain the teacher open-set prediction (TOP).

[0173] Among them, the uncertainty distribution modeling process includes:

[0174] Considering that due to the lack of supervision information for unknown classes, the diffusion classifier can only give a closed-set prediction of samples. Therefore, through uncertainty distribution modeling, convert the closed-set prediction into an open-set prediction;

[0175] First, since in the diffusion classifier, each step of the reverse denoising process contains randomness, especially the process of gradually denoising starting from standard Gaussian noise; therefore, use Monte Carlo sampling to sample M reverse denoising trajectories Calculate the variance of the category prediction probability distribution:

[0176]

[0177] Among them, is an indicator function, with a value of 1 indicating belonging to category c, otherwise 0; μ c represents the mean of the prediction probabilities of category c, represents the variance of the prediction probabilities of category c, represents the predicted label of the i-th sample;

[0178] Next, based on the variance of the predictions for each class of the feature z, the overall uncertainty of the classifier is quantified:

[0179]

[0180] where C represents the total number of classes, and c refers to a specific class;

[0181] Next, obtain the uncertainty variance values {U(z1),..., U(z q ,..., z Q} for each sample from a set of sample feature sets {z1,..., z q ),..., U(z Q )}; assume that these variance values follow a Gaussian distribution, and use these values to fit the parameters of the uncertainty distribution to obtain the uncertainty distribution N(μ u , (σ u ) 2 ):

[0182]

[0183] where Q represents the total number of a set of sample feature sets, and q represents a specific sample feature;

[0184] After that, for each sample feature z q , by calculating the cumulative probability of the uncertainty variance of this sample in the uncertainty distribution, obtain the probability that it belongs to an unknown class:

[0185]

[0186] where ξ q represents a random integral variable, μ u represents the mean value obtained by normalizing the sample feature z q , and σ u represents the variance;

[0187] Finally, use the probability mapping module to convert the closed-set prediction to an open-set prediction:

[0188]

[0189] where is the open-set prediction, PMM(·) is the probability mapping module, represents the sum of the means of the predictions for each class; the probability mapping operation in the probability mapping module is expressed as:

[0190]

[0191] where || is the concatenation operation, and Softmax is an activation function;

[0192] Finally, to ensure that the uncertainty variance truly reflects the uncertainty of the model, the following uncertainty calibration regularization is used to constrain the model to have low uncertainty for correct classes and high uncertainty for incorrect classes. For a single sample, its uncertainty calibration regularization can be calculated by the following formula:

[0193]

[0194] where is the indicator function, with a value of 1 indicating that the corresponding condition is satisfied, otherwise 0.

[0195] Step 3, Graph Isomorphism Knowledge Distillation: Construct the K-nearest neighbor graph structure of samples in the feature space. Respectively, use the teacher open-set prediction and the student open-set prediction as nodes to construct the teacher / student instance graph. Minimize the Wasserstein distance between the prediction probability distributions of the two through the graph isomorphism knowledge distillation module to achieve knowledge distillation based on the alignment of the prediction manifolds. Specifically, it includes:

[0196] Step 31, Search for the K-nearest neighbors of samples in the feature space to obtain the edges between samples;

[0197] Step 32, Consider the teacher open-set prediction corresponding to the sample and its K-nearest neighbors as nodes, and establish edge connections between them to construct the teacher instance graph;

[0198] Step 33, Based on the open-set student classifier, also perform open-set predictions on these samples, and construct the student instance graph based on the prediction results;

[0199] Step 34, Input the student instance graph and the teacher instance graph into the graph isomorphism knowledge distillation module. By minimizing the Wasserstein distance between them, constrain the consistent prediction manifolds of the student classifier and the teacher classifier to achieve knowledge distillation between the student classifier and the teacher classifier.

[0200] Among them, the construction of the instance graph in the graph isomorphism knowledge distillation process includes:

[0201] Let the sample feature set in the feature space be \(Z = \{z_1,...,z n ,...,z N \}\), and each sample feature \(z n \) corresponds to a class prediction, where \(N\) is the number of samples; it is desired to search for the K-nearest neighbor set of each sample feature \(z n \) in the feature space and construct the instance graph \(G=(V, E)\); among them, the node set \(V\) contains the class predictions of the sample features and their corresponding K-nearest neighbors, and the edge set \(E\) reflects the association attributes between the samples and their K-nearest neighbors;

[0202] First, in the feature space, the node v i and node v j The clustering between is defined as:

[0203] d(v i ,v j )=||v i -v j ||2

[0204] For each node v i , its K nearest neighbor set is expressed as:

[0205]

[0206] in, is to press d(v i ,v j ) The most recent K sample indexes sorted from small to large;

[0207] Next, in v i and each node v in its K set j ∈N K (v i ) and the edge set is defined as:

[0208]

[0209] Finally, based on the sample features and the teacher classifier and student classifier, the category of each sample is predicted, where the teacher classifier is a closed set diffusion classifier and the student classifier is an open set classifier; based on the above composition principle, the instance graph G corresponding to the teacher classifier is constructed Te =(V Te ,E Te ) and the instance graph G corresponding to the student classifier Stu =(V Stu ,E Stu ), where V Te represents the node set in the teacher instance graph, E Te represents the set of edges, V Stu represents the node set in the student instance graph, E Stu Represents a collection of edges.

[0210] In the graph isomorphism knowledge distillation process, a graph isomorphism knowledge distillation module based on instance graphs is proposed. The instance graphs are used to describe the structural associations between samples, and the Wasserstein distance is further introduced to achieve graph isomorphism, transforming the graph structure optimization problem into a distribution alignment problem.

[0211] For the node set V in the teacher instance graph and the student instance graph Te and V Stu, its distribution is defined as:

[0212]

[0213] where, and represent the node features in the teacher instance graph and the student instance graph respectively, δ(·) is the Dirac function, which is used to represent the discrete distribution, and K represents the number of samples;

[0214] The Wasserstein distance between the teacher instance graph distribution and the student instance graph distribution is expressed as:

[0215]

[0216] where, is the set of all joint distributions that satisfy the marginal distributions of and , and is expressed as:

[0217]

[0218] A Wasserstein distance metric Γ ω (·) based on neural network parameters is introduced to approximate φ(·), where φ(·) are learnable neural network parameters; thus, the optimization objective is expressed as:

[0219]

[0220] where, represents the expectation calculated by the function Γ Te (·) when the variable v follows the distribution ω (·), represents the expectation calculated by the function Γ Stu (·) when the variable v follows the distribution ω (·).

[0221] To ensure that Γ ω (·) satisfies the 1-Lipschitz condition, the following gradient penalty term is introduced:

[0222]

[0223] where, is the distribution generated by interpolation , α~U(0,1), represents the expectation of the function calculation result when the variable follows the distribution ;

[0224] Therefore, the graph isomorphism knowledge distillation loss is expressed as:

[0225]

[0226] where λ is the gradient penalty weight;

[0227] Finally, during the graph isomorphism knowledge distillation process, by minimizing the Wasserstein distance loss between distributions and imposing the constraint of the 1-Lipschitz condition, the student distribution gradually approaches the teacher distribution ; in this way, the fine-grained knowledge about the sample structure association implicit in the teacher classifier is distilled into the student classifier to ensure a consistent prediction manifold for the student classifier and the teacher classifier. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An open-set hyperspectral image instance graph construction method based on a diffusion model, characterized in that: It includes the following steps: Step 1, Uncertainty feature acquisition: Extract the spatial and spectral features of the target domain HSI through a spatial encoder and a spectral encoder. Perform forward diffusion in the diffusion classifier starting from the true label, and perform conditional reverse denoising based on the sample features. Finally, obtain the closed-set class prediction; Step 2, Uncertainty distribution modeling: Generate a prediction set for a single sample through Monte Carlo sampling and calculate the prediction variance. Construct an uncertainty distribution based on the prediction variances of all samples, use the cumulative distribution function to infer the unknown class probability, and finally fuse the known class and unknown class probabilities to generate the teacher open-set prediction through a mapping module; Step 3, Graph isomorphism knowledge distillation: Construct a K-nearest neighbor graph structure of samples in the feature space. Construct teacher / student instance graphs with the teacher open-set prediction and the student open-set prediction as nodes respectively. Minimize the Wasserstein distance between the prediction probability distributions of the two through the graph isomorphism knowledge distillation module to achieve knowledge distillation based on the alignment of the prediction manifolds.

2. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1, wherein: The said Step 1 includes: Step 11, Input the target domain HSI into the feature extractor, and use the spatial encoder and the spectral encoder to capture the spatial dependence between pixels and the spectral dependence between bands to obtain sample features; Step 12: Input the sample features into the diffusion classifier. During the forward diffusion process, the diffusion classifier gradually adds Gaussian noise to the true label to obtain y T ; Step 13, during the reverse denoising process, the diffusion classifier uses the sample features as conditional signals to gradually denoise the noise distribution to obtain closed-set class predictions 3. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1, wherein: The said Step 2 includes: Step 21, Obtain the prediction set for a single sample by performing Monte Carlo sampling in the noise distribution and calculate the prediction variance in the set; Step 22, Regard the variance of the single-sample prediction as a point in the uncertainty distribution, and model the uncertainty distribution from the prediction distributions of a group of samples. Calculate the cumulative distribution function of each sample in the uncertainty distribution to obtain its cumulative distribution probability, and further estimate the unknown class probability of the sample; Step 23, Input the known class probability and the known class probability of the diffusion classifier into the probability mapping module to obtain the teacher open-set prediction.

4. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1, characterized in that: The said Step 3 includes: Step 31, Search for the K-nearest neighbors of the samples in the feature space to obtain the edges between the samples; Step 32, Regard the sample and the teacher open-set prediction corresponding to its K-nearest neighbors as nodes, and establish edge connections between them to construct the teacher instance graph; Step 33, Based on the open-set student classifier, also perform open-set prediction on these samples and construct the student instance graph based on the prediction results; Step 34, Input the student instance graph and the teacher instance graph into the graph isomorphism knowledge distillation module. By minimizing the Wasserstein distance between them, constrain the prediction manifolds consistent between the student classifier and the teacher classifier to achieve knowledge distillation between the student classifier and the teacher classifier.

5. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1 or 2, characterized in that: In the said Step 1, the forward diffusion process includes: The forward diffusion process of the diffusion classifier is disassembled into: where p(y 0:T ) represents the joint probability of all states of label y occurring from time 0 to T, q(y0) is the prior probability of state y0, and q(y t |y t-1 ) represents the conditional probability of the current state y t-1 given the previous state y t ; In the forward diffusion process, given a label observation y0, the following conditional probability is obtained: where p(y 1:T |y0) represents the joint probability of all states of label y occurring at times 1 to T given y0, and q(y1|y0) represents the conditional probability of state y1 given y0; Under the framework of the diffusion model, the encoder q(y t |y t-1 ) is fixed to a linear Gaussian transformation, that is, q(y t |y t-1 ) is a Gaussian distribution with as the mean and as the variance. Therefore, using the reparameterization trick, y t-1 is used to obtain y t : where α t is a hyperparameter for controlling the noise intensity, ε is a random variable following a standard normal distribution, denotes the standard normal distribution, and I denotes a variance of 1; Based on the above recurrence relation, starting from y0, directly calculate y at any step t : Among them, it can be seen that, due to α i ∈(0,1), y t gradually approaches the standard Gaussian distribution as the time step increases.

6. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1 or 2, characterized in that: In the said Step 1, the reverse denoising process includes: The goal of the reverse denoising process is to start from the standard Gaussian distribution and the decoder gradually denoises y conditioned on the sample feature z to recover the true label distribution. Therefore, the joint probability p(y T ) is factorized as: 0:T ) Among them, the probability density of p(y T ) is known and is a standard Gaussian distribution, that is However, p(y t |y t+1 , z) is difficult to calculate directly. Therefore, here a neural network is used to learn p θ (y0|y1, z) to approximate the true distribution, where θ is the neural network parameter.

7. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1 or 2, characterized in that: In the said Step 1, the learning objectives in the uncertainty feature acquisition process include: The goal of the diffusion classifier is to model p(y0|z), where: y0 is the true label; z = f(x), and f(·) is the feature extractor. The label generation process is divided into two stages. In the forward diffusion stage, noise is gradually added to the label y0 and finally transformed into a standard Gaussian distribution. In the reverse denoising stage, through the conditional input z and the reverse process parameterized by the neural network, the Gaussian distribution is gradually denoised and restored to the target label y0; The learning objective of the diffusion classifier is to maximize the log-likelihood of the observed data by optimizing the following evidence lower bound to learn the parameter θ: Among them, the first term aims to minimize the reconstruction loss of the true label conditioned on the sample feature z. The second term D KL (p(y T )||q(y T |y0)) aims to constrain the distribution of the end point y T of the forward diffusion to be close to the standard Gaussian distribution. The third term D KL (p θ (y t-1 |y t ,z)||q(y t-1 |y t ,y0)) aims to minimize the difference between the distribution p θ (y t |y t+1 ,z) learned by the parameter θ and the true distribution q(y t-1 |y t ,y0); Since the second term, the forward diffusion process is defined as a linear Gaussian transformation, therefore, there are no optimizable parameters in this term during training, and the focus is on the optimization of the first term and the third term; First, considering p θ (y0|y1,z) is a conditional Gaussian distribution that depends on y1 and z, and let its mean be μ θ which is a parametric function of y1 and z, denoted as μ θ (y1,z,t = 1), and the variance is a constant, denoted as Σ; expanding lnp θ (y0|y1,z), we get: where exp represents the natural exponential function and n represents the dimension of the data; Therefore, maximizing is equivalent to minimizing the mean squared error between the diffusion classifier output and y0; Then, considering that q(y t-1 |y t ,y0) in the third term follows a Gaussian distribution with mean μ q (y t ,y0) and variance Σ q (t); to make p θ (y t |y t+1 ,z) better fit q(y t-1 |y t ,y0), here p θ (y t |y t+1 ,z) is also regarded as a Gaussian distribution, and μ θ (y t ,z) is used to represent the mean of this distribution. Since in q(y t-1 |y t ,y0), the variance term is only related to t, therefore, here it is set that the variance Σ θ (y t |y t+1 ,z) of the corresponding Gaussian distribution is Σ θ (t) = Σ q (t); Next, obtain p θ (y t |y t+1 ,z) and the KL divergence between q(y t-1 |y t ,y0): Among them, represents the variance of the q-distribution corresponding to time t; As can be seen from the above equation, minimizing the third term is equivalent to minimizing μ θ (y t , z) and the difference between μ q (y t , y0); Next, to reduce the learning difficulty of the model, it is necessary to parameterize μ θ (y t , z) in the form of μ q (y t , y0), and μ q (y t , y0) is expressed in the following form: Among them, α t is a hyperparameter used to control the strength of noise, Therefore, μ θ (y t , z) is expressed as: Among them, is the output of a parameterized neural network; Therefore, substituting μ q (y t , y0) and into D KL (p θ (y t-1 |y t , z)||q(y t-1 |y t , y0)), we get: where ε represents the random Gaussian noise added to y during the forward diffusion process from the (t - 1)-th time step to the t-th time step t-1 ; In summary, the final optimization objective of the diffusion classifier is expressed as: where γ is the trade-off coefficient. The first term of this loss hopes that the predicted label after reverse denoising is consistent with the true label, and the second term expects the diffusion classifier to gradually recover the true label distribution from the noise distribution at each time step.

8. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1 or 3, characterized in that: In step 2, the uncertainty distribution modeling process includes: Considering that due to the lack of supervision information for unknown classes, the diffusion classifier can only give closed-set predictions for samples. Therefore, through uncertainty distribution modeling, the closed-set predictions are converted into open-set predictions; First, since each step of the reverse denoising process in the diffusion classifier involves randomness, especially when starting from standard Gaussian noise and gradually denoising to generate the process; therefore, using Monte Carlo sampling, sample M reverse denoising trajectories Calculate the variance of the class prediction probability distribution: wherein, is an indicator function, with a value of 1 indicating belonging to class c and 0 otherwise; μ c represents the mean of the predicted probabilities for class c, represents the variance of the predicted probabilities for class c, represents the predicted label of the i-th sample; Next, based on the variance of the predictions for each class of the feature z, the overall uncertainty of the classifier is quantified: where C represents the total number of classes and c refers to a specific class; Next, obtain the uncertainty variance values {U(z1),..., U(z q ,..., z Q )} for each sample from a set of sample feature sets {z1,..., z q ),..., U(z Q )}; assume that these variance values follow a Gaussian distribution, and use these values to fit the parameters of the uncertainty distribution to obtain the uncertainty distribution N(μ u , (σ u ) 2 ): where Q represents the total number of sets of sample features and q represents a specific sample feature; After that, for each sample feature z q , by calculating the cumulative probability of the uncertainty variance of the sample in the uncertainty distribution, the probability that it belongs to an unknown class is obtained: Among them, ξ q represents a random integral variable, μ u represents the mean for normalizing the sample feature z q and σ u represents the variance; Finally, the probability mapping module is used to convert the closed-set predictions into open-set predictions: Among them, is open-set prediction, and PMM(·) is the probability mapping module, represents the cumulative sum of the predicted means for each category; the probability mapping operation in the probability mapping module is expressed as: where || is the concatenation operation and Softmax is an activation function; Finally, to ensure that the uncertainty variance truly reflects the uncertainty of the model, the following uncertainty calibration regularization is used to constrain the model to have low uncertainty for the correct class and high uncertainty for the wrong class. For a single sample, its uncertainty calibration regularization can be calculated by the following formula: Among them, is an indicator function, with a value of 1 indicating that the corresponding condition is met, otherwise 0.

9. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1 or 4, characterized in that: In step 3, the instance graph construction in the graph isomorphism knowledge distillation process includes: Suppose the feature space has a sample feature set Z = {z1,..., z n ,..., z N}, and each sample feature z n corresponds to a class prediction, where N is the number of samples; it is desired to search for the K-nearest neighbor set of each sample feature z n in the feature space and construct an instance graph G = (V, E); where the node set V contains the sample features and the class predictions of their corresponding K-nearest neighbors, and the edge set E reflects the association attributes between the samples and their K-nearest neighbors; First, in the feature space, the clustering between node v i and node v j is defined as: d(v i ,v j ) = ||v i -v j ||² For each node v i , its K-nearest neighbor set is denoted as: Among them, are the indices of the K nearest samples sorted in ascending order of d(v i , v j ); Next, between v i and each node v in its K-set j ∈ N K (v i ), an edge is established, and the edge set is defined as: Finally, based on the sample features and the class predictions of each sample by the teacher classifier and the student classifier, where the teacher classifier is a closed-set diffusion classifier and the student classifier is an open-set classifier; construct the instance graph G corresponding to the teacher classifier based on the above graph construction principle Te =(V Te , E Te ) and the instance graph G Stu =(V Stu , E Stu ), where V Te represents the set of nodes in the teacher instance graph, E Te represents the set of edges, V Stu represents the set of nodes in the student instance graph, and E Stu represents the set of edges.

10. The method for constructing an open-set hyperspectral image instance graph based on a diffusion model according to claim 1 or 4, characterized in that: In step 3, in the graph isomorphism knowledge distillation process, a graph isomorphism knowledge distillation module based on the instance graph is proposed. The structural associations between samples are characterized by the instance graph, and the Wasserstein distance is further introduced to achieve graph isomorphism, converting the graph structure optimization problem into a distribution alignment problem; For the node sets V in the teacher instance graph and the student instance graph Te and V Stu , their distribution is defined as: Among them, and respectively represent the node features in the teacher instance graph and the student instance graph. δ(·) is the Dirac function, which is used to represent the discrete distribution, and K represents the number of samples; The Wasserstein distance between the teacher instance graph distribution and the student instance graph distribution is expressed as: Among them, is the set of all joint distributions that satisfy the marginal distributions and and is denoted as: A Wasserstein distance metric Γ parameterized based on neural network parameters is introduced ω (·) to approximate φ(·), where φ(·) are learnable neural network parameters; thus, the optimization objective is expressed as: wherein, denotes the variable v Te subject to the distribution when it is passed through the function Γ ω (·), the calculated expectation denotes the variable v Stu subject to the distribution when it is passed through the function Γ ω (·), the calculated expectation. To ensure that Γ ω (·) satisfies the 1-Lipschitz condition, the following gradient penalty term is introduced: Among them, is a distribution generated by interpolation where α ∼ U(0, 1), denotes the variable follows the distribution and is the expectation of the function calculation result when... Therefore, the graph isomorphism knowledge distillation loss is expressed as: where λ is the gradient penalty weight; Finally, in the process of graph isomorphism knowledge distillation, by minimizing the Wasserstein distance loss between distributions and imposing the constraint of the 1-Lipschitz condition, the student distribution gradually approaches the teacher distribution ; in this way, the fine-grained knowledge about the structural association of samples implicit in the teacher classifier is distilled into the student classifier to ensure a consistent prediction manifold for the student classifier and the teacher classifier.

Citation Information

Cited By

  • Scalable expert foundry system using hierarchical supervisory networks and geometric manifold architectures for multi-domain cognitive processing

    US12572748B1

  • Evolutionary thought caching for multi-stage language model systems

    US12585882B1

  • Persistent cognitive machine with curated long term memory

    US12602549B1

  • Geometric Multi-Modal Sensor Fusion for Infrastructure Monitoring

    US20260228438A1