Method and system for detecting out-of-distribution samples based on self-adaptive distribution ball frame
Through dynamic distribution boundary adjustment, multi-scale feature fusion and energy optimization of the adaptive distribution ball framework, the adaptability and robustness of deep neural networks in sample detection outside the distribution is solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510840204.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the detection of out-of-distributed samples, existing deep neural networks have problems such as static threshold mechanisms that are not adapted to the dynamic data environment, limited feature representation capabilities and lack of dynamic optimization closed loops, resulting in a decrease in detection accuracy, especially in complex multimodal data and adversarial attack scenarios.
Adaptive distribution ball framework is adopted, including dynamic distribution boundary adjustment module, dynamic multi-scale feature fusion module and dynamic energy optimization module, to generate enhanced samples through adversarial generation network, dynamically adjust feature boundaries and optimize classification thresholds, forming an end-to-end closed-loop optimization mechanism.
The accuracy and scene adaptability of out-of-distribution sample detection are improved, the ability to adapt to complex distribution offsets is enhanced, and the robustness and efficiency of detection are improved.
Smart Images

Figure CN120375097A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and machine learning, and more specifically, to an out-of-distribution sample detection method and system based on an adaptive distribution sphere framework. Background Art
[0002] Currently, in the field of out-of-distribution (OOD) sample detection, typical method systems represented by the maximum soft probability (MSP), ODIN, OpenMax, etc. have been formed for deep neural networks. MSP constructs a detection threshold through the maximum probability value of the output layer. ODIN introduces temperature scaling and input perturbation to enhance the confidence discrimination. While the Mahalanobis distance method and energy-based model based on the statistical characteristics of the feature space improve the detection performance from the perspective of feature distribution modeling. These methods show certain effectiveness on standard benchmark datasets and have formed preliminary applications in scenarios such as autonomous driving perception modules and medical image assisted diagnosis.
[0003] However, the existing technologies have three key defects: First, there is an essential contradiction between the static threshold mechanism and the dynamic data environment. MSP and ODIN rely on preset fixed thresholds and cannot adapt to the continuously changing modal distributions and adversarial sample interferences in real scenarios. Second, the feature representation ability is limited. Traditional methods mostly use global semantic features or single-scale statistics and are difficult to capture the local detailed features and cross-scale context associations of tiny lesions in medical images. Third, there is a lack of a dynamic optimization closed-loop. The offline parameter curing characteristics of methods such as OpenMax make them ineffective in dealing with the distribution drift of real-time data streams, and the high-dimensional computational complexity based on the Mahalanobis distance and the static threshold strategy of the energy model further restrict the actual deployment efficiency. These problems lead to a significant decline in the detection accuracy of existing systems in complex multi-modal data and adversarial attack scenarios, such as high-risk hazards like misjudging road signs in autonomous driving and missing early lesions in medical images.
[0004] Therefore, how to develop a solution that can dynamically adjust the feature boundaries, fuse multi-scale information, and achieve end-to-end optimization to improve the accuracy and scene adaptability of out-of-distribution detection is an urgent problem for those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides an out-of-distribution sample detection method and system based on an adaptive distribution sphere framework, which overcomes the above defects.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] An out-of-distribution sample detection method based on an adaptive distribution sphere framework, specifically:
[0008] Training steps:
[0009] Construct an out-of-distribution sample detection model based on an adaptive distribution sphere framework, where the adaptive distribution sphere framework includes a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module;
[0010] Input the training data into the dynamic distribution boundary adjustment module, and generate enhanced samples that meet global semantic alignment and local detail offset optimization through a generative adversarial network to expand the distribution boundary; the training data comes from a training dataset constructed by in-distribution data and auxiliary out-of-distribution data;
[0011] Input the enhanced samples into the dynamic multi-scale feature fusion module, extract global features and local features, dynamically calculate the fusion weights based on the distribution differences between the global features and the local features, and fuse the global features and the local features based on the fusion weights to generate fused features;
[0012] Input the fused features into the dynamic energy optimization module, calculate the energy value through an energy function and dynamically adjust the classification threshold using a closed-loop feedback mechanism; and apply the energy loss gradient to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module for parameter adjustment;
[0013] Detection steps:
[0014] Obtain the image to be detected;
[0015] Input the image to be detected into the out-of-distribution sample detection model for OOD sample recognition and output the recognition result.
[0016] Optionally, it further includes a preprocessing module for performing normalization and enhancement processing on the training data or the image to be detected to generate preprocessed data.
[0017] Optionally, global semantic alignment quantifies the global semantic difference using the Wasserstein distance, and its expression is:
[0018] ;
[0019] In the formula, is the probability distribution of in-distribution data; is the probability distribution of out-of-distribution data; is the set of all joint distributions that satisfy the marginal distributions of and ; is the expectation operation with respect to the joint distribution γ; is and The L2 norm distance between.
[0020] Optionally, local detail offset optimization enhances the sensitivity to fine-grained anomalies through the Sinkhorn distance, and its expression is:
[0021] ;
[0022] In the formula, is the probability distribution of in-distribution data; is the probability distribution of out-of-distribution data; is the set of all joint distributions that satisfy the marginal distributions of and ; γ is the joint distribution; is the expected value; is the cost function; is the entropy regularization coefficient; H(γ) is the entropy of the joint distribution γ.
[0023] Optionally, the dynamic distribution boundary adjustment module also adjusts the weights of global and local distribution modeling through a dynamic weight strategy to achieve hierarchical boundary modeling. The defined joint optimization objective is:
[0024] ,
[0025] In the formula, is the global semantic alignment loss, is the local detail optimization loss; is the weight coefficient, which adjusts the balance between global and local losses over time , and its expression is:
[0026] ;
[0027] In the formula, is the initial weight coefficient; is the maximum number of training steps; is the non-linear growth factor.
[0028] Optionally, the calculation formula of the fusion weight is:
[0029] ;
[0030] In the formula, is the activation function; is the Wasserstein distance between the global feature and the local feature ; is the Sinkhorn distance between the global feature and the local feature ; is a very small constant; is the regularization term weight coefficient; is the covariance consistency regularization term.
[0031] Optionally, hierarchical contrast learning is introduced in the dynamic multi-scale feature fusion module to optimize the discriminability of the feature space, including:
[0032] Global contrast loss, by aggregating the high-level semantics of in-class samples, enhances the compactness of the feature space, and its expression is:
[0033] ;
[0034] Local contrast loss, by amplifying the fine-grained differences, enhances the sensitivity to local anomalies:
[0035] ;
[0036] In the formula, is the number of samples in the batch; is the global feature vector of the th sample; is the global feature vector of the positive sample of the same class as ; is the global feature vector of the negative sample; is the temperature coefficient of the global contrast loss; is the cosine similarity calculation function; is the local feature vector of the th sample; is the local feature vector of the positive sample of the same class as ; is the local feature vector of the negative sample; is the temperature coefficient of the local contrast loss; is the number of negative samples in the contrast learning.
[0037] Optionally, the energy function is defined as:
[0038] ;
[0039] In the formula, is the confidence level output by the model; C is the total number of classes.
[0040] Optionally, the calculation formula of the classification threshold is:
[0041] ;
[0042] In the formula, is the exponentially weighted moving average of the energy of in-distribution samples; is the threshold adjustment coefficient; is the exponentially weighted moving average of the standard deviation of the energy of in-distribution samples.
[0043] An out-of-distribution sample detection system based on an adaptive distribution sphere framework, comprising:
[0044] A model construction module for constructing an out-of-distribution sample detection model based on an adaptive distribution sphere framework, the adaptive distribution sphere framework including a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module;
[0045] A sample generation module for inputting training data into the dynamic distribution boundary adjustment module, generating enhanced samples that meet global semantic alignment and local detail offset optimization through a generative adversarial network to expand the distribution boundary; the training data comes from a training data set constructed by in-distribution data and auxiliary out-of-distribution data;
[0046] A feature fusion module for inputting the enhanced samples into the dynamic multi-scale feature fusion module, extracting global features and local features, dynamically calculating fusion weights based on the distribution differences between the global features and the local features, and fusing the global features and the local features based on the fusion weights to generate fused features;
[0047] A decision and feedback module for inputting the fused features into the dynamic energy optimization module, calculating an energy value through an energy function and dynamically adjusting a classification threshold using a closed-loop feedback mechanism; and applying an energy loss gradient to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module for parameter adjustment;
[0048] A sample recognition module for obtaining an image to be detected and inputting the image to be detected into the out-of-distribution sample detection model for OOD sample recognition, and outputting a recognition result.
[0049] As can be seen from the above technical solutions, the present invention provides an out-of-distribution sample detection method and system based on an adaptive distribution sphere framework. Compared with the prior art, it has the following beneficial effects:
[0050] 1. The dynamic distribution boundary adjustment module (DBA) generates enhanced samples covering global semantic differences and local detail anomalies through a multi-scale distance collaborative optimization mechanism (dynamic weight allocation of Wasserstein and Sinkhorn distances), effectively alleviating the distribution deviation between auxiliary data and real OOD data; this module breaks through the limitations of traditional static boundary modeling through a staged boundary evolution strategy, gradually transitioning from global alignment to local refinement, and enhancing the model's adaptability to complex distribution offsets;
[0051] 2. The Dynamic Multi-scale Feature Fusion Module (MFF) first upsamples the local features to 128 dimensions, then dynamically allocates fusion weights through the ratio of the Wasserstein distance to the Sinkhorn distance, suppresses covariate shift by combining with the covariance consistency regularization term, and realizes complementary modeling of cross-scale features through the synergistic effect of the global contrast loss (high-level semantic aggregation) and the local contrast loss (low-temperature fine-grained optimization);
[0052] 3. The Dynamic Energy Optimization Module (EGO) dynamically adjusts the classification threshold through Exponential Weighted Moving Average (EMA) based on the real-time statistical characteristics (mean and standard deviation) of the ID sample energy, solving the problem of insufficient robustness of traditional static thresholds. At the same time, the energy gradient drives the weight allocation of the MFF module and the adversarial sample generation direction of the DBA module through the backpropagation path, forming an end-to-end collaborative optimization closed-loop. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0054] Figure 1 It is a schematic diagram of the method flow provided by the present invention.
[0055] Figure 2 It is a schematic diagram of the model training process provided by the present invention;
[0056] Figure 3 It is a schematic diagram of the data processing flow of the DBA module provided by the present invention;
[0057] Figure 4 It is a schematic diagram of the interaction flow between the adversarial generation network (DBA module) and the backbone network provided by the present invention;
[0058] Figure 5 It is a schematic diagram of the ADSF backbone network provided by the present invention;
[0059] Figure 6 It is a schematic diagram of the internal structure of the BasicBlock provided by the present invention;
[0060] Figure 7 It is a schematic diagram of the data processing flow of the MFF module provided by the present invention;
[0061] Figure 8 It is a schematic diagram of the data processing flow of the EGO module provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] An embodiment of the present invention discloses an out-of-distribution sample detection method based on an adaptive distribution sphere framework, as Figure 1 shown, including the following steps:
[0064] Specific steps in the training stage:
[0065] Step 1: Construct an out-of-distribution sample detection model based on the adaptive distribution sphere framework. The adaptive distribution sphere framework includes a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module;
[0066] Step 2: Input the training data into the dynamic distribution boundary adjustment module, and generate enhanced samples that meet global semantic alignment and local detail offset optimization through an adversarial generation network to expand the distribution boundary; the training data comes from a training data set constructed by in-distribution data and auxiliary out-of-distribution data;
[0067] Step 3: Input the enhanced samples into the dynamic multi-scale feature fusion module, extract global features and local features, dynamically calculate the fusion weights based on the distribution differences between the global features and the local features, and fuse the global features and the local features based on the fusion weights to generate fused features;
[0068] Step 4: Input the fused features into the dynamic energy optimization module, calculate the energy value through an energy function and dynamically adjust the classification threshold using a closed-loop feedback mechanism; and apply the energy loss gradient to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module for parameter adjustment;
[0069] Specific steps in the detection stage:
[0070] Step 5: Obtain the image to be detected;
[0071] Step 6: Input the image to be detected into the out-of-distribution sample detection model for OOD sample recognition, and output the recognition result.
[0072] In one embodiment, it further includes a preprocessing module for performing standardization and enhancement processing on the training data or the image to be detected to generate preprocessed data.
[0073] Further, in this embodiment, in-distribution (ID) data is obtained based on the CIFAR-10 / CIFAR-100 dataset, and auxiliary out-of-distribution (OOD) data is obtained based on 300K Random Images. The image to be processed is standardized and enhanced to obtain preprocessed data.
[0074] Among them, the specific steps of standardization are as follows: The input image is adjusted to a preset resolution (such as 32×32 pixels) using bilinear interpolation to ensure input consistency;
[0075] The pixel values are normalized, and the formula is:
[0076] ;
[0077] In the formula, is the original pixel value of the input image, that is, the data to be normalized; and are the mean and standard deviation respectively; is the normalized pixel value, that is, the standardized input.
[0078] Subsequently, random horizontal flipping (probability 50%) and random cropping (cropped to 32×32 after padding 4 pixels) are applied to enhance data diversity; finally, the ID data and the auxiliary OOD data are mixed in a ratio of 1:2, the ID batch size is 128, and the OOD batch size is 256.
[0079] Furthermore, in the test stage, a real OOD dataset (such as SVHN, Textures, LSUN_resize, iSUN, Places365) is used to verify the model generalization ability; the test data mixes ID and OOD samples, the batch size is 256, and data augmentation operations are disabled.
[0080] In one embodiment, the adaptive distribution sphere framework consists of a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module, where:
[0081] The dynamic distribution boundary adjustment module generates enhanced samples that satisfy global semantic alignment (constrained by the Wasserstein distance) and local detail offset optimization (constrained by the Sinkhorn distance) through an adversarial generation network, and adjusts the weights of global and local distribution modeling based on a dynamic weight strategy;
[0082] The dynamic multi-scale feature fusion module dynamically calculates the fusion weights of global and local features. The fusion weights are allocated based on the ratio of the global Wasserstein distance to the local Sinkhorn distance, and noise interference is suppressed through a covariance consistency regularization term;
[0083] The dynamic energy optimization module dynamically adjusts the classification threshold based on the sample energy score and applies gradient penalty to high-energy samples.
[0084] Furthermore, as Figure 2 shown, starting from the in-distribution data and auxiliary out-of-distribution data, first, the adversarial sample expansion distribution boundary is generated through the dynamic distribution boundary adjustment (DBA) module. Then, the global semantic difference is quantified using the Wasserstein distance and combined with the Sinkhorn distance to optimize the local details, generating enhanced samples. The global and local distribution modeling weights are dynamically adjusted to achieve hierarchical boundary modeling. Subsequently, through the dynamic multi-scale feature fusion (MFF) module, the feature weights are dynamically allocated based on the global and local distribution differences, and combined with hierarchical contrastive learning (high-temperature coefficient aggregates the in-class semantics, and low-temperature coefficient amplifies the fine-grained differences) to optimize the discriminability of the feature space. Finally, through the dynamic energy optimization (EGO) module, an energy function is defined to quantify the separability of samples, the classification threshold is dynamically adjusted based on real-time statistical characteristics, and the energy gradient is fed back to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module in the reverse direction, forming a "generation-fusion-decision-feedback" closed-loop optimization process, significantly improving the discrimination ability of in-distribution and out-of-distribution samples.
[0085] In one embodiment, the preprocessed data is input into the dynamic distribution boundary adjustment module to generate adversarial samples to expand the distribution boundary, and the DBA module realizes hierarchical boundary modeling through a multi-scale distance collaborative optimization mechanism.
[0086] In one embodiment, the processing flow of the DBA module is as Figure 3 shown, specifically:
[0087] Global semantic alignment: The overall distribution shift between in-distribution (ID) and out-of-distribution (OOD) samples is quantified using the Wasserstein distance, and its expression is:
[0088] ;
[0089] wherein, is the probability distribution of in-distribution data; is the probability distribution of out-of-distribution data; is the set of all joint distributions satisfying the marginal distributions of and ; γ is the joint distribution, belonging to ; is the expectation operation with respect to the joint distribution γ; is the L2 norm distance between x and y.
[0090] Local detail offset optimization: The sensitivity to fine-grained anomalies is enhanced through the Sinkhorn distance, and its expression is:
[0091] ;
[0092] Wherein, is the probability distribution of in-distribution data; is the probability distribution of out-of-distribution data; is the set of all joint distributions that satisfy the marginal distributions of and ; γ is the joint distribution; is the expected value; is the cost function, defined as the difference metric between and ; is the entropy regularization coefficient, used to balance the weights of the cost term and the entropy term; H(γ) is the entropy of the joint distribution γ, calculated as .
[0093] Dynamic weight strategy: Dynamically adjust the global and local distribution modeling weights, define the joint optimization objective, and its expression is:
[0094] ;
[0095] Wherein, is the global semantic alignment loss, is the local detail optimization loss; is the weight coefficient, which can increase dynamically with the training stage, and the formula is:
[0096] ;
[0097] Wherein, is the initial weight coefficient, , in this embodiment ; is the maximum number of training steps, ≤2000, in this embodiment =1000; is the non-linear growth factor, , in this embodiment at .
[0098] Adversarial sample generation: The dynamic distribution boundary adjustment module (DBA) generates samples through an independent adversarial generation network, and the adversarial generation network adopts a lightweight fully connected architecture to realize the generation of adversarial samples. As Figure 4As shown, the input of the generation network is a 128-dimensional high-level semantic feature vector of in-distribution (ID) samples extracted from the backbone network. First, feature compression is performed through the first fully connected layer (BN) (dimension 128→64), and the ReLU activation function is used to enhance the non-linear expression ability. Subsequently, further dimensionality reduction is carried out through the second fully connected layer (BN) (dimension 64→32), and the perturbation feature vector of the generated sample is output. The final output of the generation network limits the perturbation amplitude within 10% of the original feature vector through L2 norm constraint, ensuring that the generated adversarial samples effectively expand the distribution boundary while avoiding excessive deviation from the true data distribution. The generated perturbation features are linearly superimposed with the original ID features to form enhanced samples. Their global semantic characteristics are far from the ID distribution through Wasserstein distance constraint, and the local details are optimized by Sinkhorn distance to approximate the fine-grained abnormal patterns of the true OOD samples. The entire generation process is achieved through end-to-end adversarial training, and the parameters of the generation network and the backbone network are updated using an alternating optimization strategy.
[0099] Among them, the backbone network consists of an input layer, three stage blocks, and an output layer, as Figure 4 and Figure 5 shown. The input layer processes a 32×32 pixel RGB image using a 3×3 convolution kernel and outputs a 16-channel feature map. The network contains three main stages (block1, block2, block3), each stage consists of multiple improved BasicBlocs, and the number of channels is 32, 64, and 128 respectively. Among them, block2 and block3 perform 2-fold downsampling respectively, and its structure is as Figure 6As shown. Each BasicBloc integrates three attention mechanisms on the basis of the traditional residual structure (N stacked residual blocks (Network Block)): First, the CBAM (Convolutional Block Attention Module) module is applied to enhance the response of important features through channel attention and spatial attention; Second, the SE (Squeeze-and-Excitation) module is introduced to dynamically calibrate the channel importance using the squeeze-and-excitation mechanism; Finally, the ECA (Efficient Channel Attention) module is added to achieve cross-channel interaction with lightweight 1D convolution. This multi-attention collaborative design significantly improves the network's sensitivity to distribution shifts. The network output layer includes global average pooling and a fully connected layer, and finally outputs 128-dimensional high-level semantic features. After the convolutional layer of each BasicBlock, Dropout (dropRate = 0.3), batch normalization (Bn1), and the ReLU activation function are applied. Then, through the two-dimensional average pooling layer (AvgPool2d(8)) and the fully connected layer (fc layer), the 128-dimensional input is converted into a num-classes-dimensional output to obtain logits (confidence), where num-classes is the number of classes; Finally, the energy calculation operation logsumexp(logits) is used to calculate the energy. In particular, temperature scaling (temp = 2.0) is introduced in the classification layer to smooth the logits (confidence) output. When training the network, an energy-based constraint mechanism is adopted, and through the energy function to distinguish between in-distribution and out-of-distribution samples. In terms of feature extraction, the 128-dimensional global features output from block3 are used to capture high-level semantic information; the 64-dimensional local features output from block2 retain more spatial detail information.
[0100] When generating adversarial samples, energy constraints are combined to force the generated samples to move away from the ID distribution and approach the true OOD characteristics. The expression of the adversarial loss function is:
[0101] ;
[0102] In the formula, is the ID repulsion coefficient, which controls the degree to which the generated samples move away from the ID distribution; is the OOD attraction coefficient, which controls the degree to which the generated samples approach the auxiliary OOD distribution; is the Wasserstein distance between the generated sample ( ) and the ID data ( ); is the Wasserstein distance between the generated sample and the auxiliary OOD data ( ).
[0103] Furthermore, in this embodiment, the expression of the adversarial loss function is:
[0104] .
[0105] Where, the ID rejection coefficient α=-0.8 forces the generated samples to be away from the ID distribution (minimizing ) maximizes the difference between the generated sample and the ID data to avoid the generated sample being misjudged as ID. The OOD attraction coefficient δ = 0.2 guides the generated sample to approach the auxiliary OOD distribution (maximize ), so that the distribution of generated samples is close to the auxiliary OOD data, and the recognition boundary of the model for the real OOD is expanded. Experiments show that overfitting the auxiliary OOD data will lead to boundary overfitting, so a weight ratio of 4:1 is selected to prioritize the generated samples away from the ID distribution.
[0106] The optimizer uses Adam, with a learning rate of 0.001 and a weight decay of 1e-4.
[0107] In one embodiment, the features output by the DBA module are input into a dynamic multi-scale feature fusion module, and the discriminability is optimized by combining global and local features, such as Figure 7 shown.
[0108] In one embodiment, the MFF module processing flow is as follows:
[0109] Extract 128-dimensional high-level semantic features (global feature extraction) from the last convolution output of the backbone network, and then extract 64-dimensional detail features from the middle-level feature map of the backbone network (such as the output of the second stage);
[0110] Dynamically allocate fusion weight α based on the difference between global and local distribution:
[0111] ;
[0112] In the formula, is the activation function; Global feature With local features Wasserstein distance between them; Global feature With local features The Sinkhorn distance between is a very small constant (such as 1×10 −5 ), to prevent the denominator from being zero; is the weight coefficient of the regularization term (value range [5,15]); is the covariance consistency regularization term, defined as .
[0113] Furthermore, for global feature extraction: 128-dimensional high-level semantic features are extracted from the output of the last layer of the backbone network (WideResNet).
[0114] For local feature extraction: 64-dimensional detailed features are extracted from the output of the middle layer (the second stage) of the backbone network.
[0115] In this embodiment, λ = 10 to balance the contribution of the regularization term.
[0116] In one embodiment, in the autonomous driving scenario, the method dynamically adjusts the fusion weight α through the global-local feature mapping relationship of lidar point cloud data, specifically:
[0117] Global feature weight ;
[0118] Local feature weight ;
[0119] When a dynamic obstacle is detected, Automatically increase by 0.1.
[0120] In one embodiment, the covariance consistency regularization term suppresses noise interference, where the regularization term is:
[0121] .
[0122] Furthermore, the covariance consistency regularization term is only activated ; in the formula, F is the Frobenius norm.
[0123] In one embodiment, a hierarchical contrastive learning strategy is adopted to optimize the feature space, including:
[0124] Global contrast loss (high temperature coefficient = 1.0) enhances the compactness of the feature space by aggregating the high-level semantics of intra-class samples, and its expression is:
[0125] ;
[0126] Local contrast loss (low temperature coefficient = 0.1) amplifies the fine-grained differences and strengthens the sensitivity to local anomalies, and its expression is:
[0127] ;
[0128] In the formula, is the temperature coefficient of the global contrast loss (such as = 1.0), which controls the aggregation degree of intra-class samples; is the temperature coefficient of the local contrast loss (such as = 0.1), amplifying the fine-grained differences; is the number of samples in a batch; is the number of negative samples in contrastive learning; is the global feature vector of the -th sample; global feature vector of the positive sample of the same category as global feature vector of other samples (negative samples); is the local feature vector of the -th sample; local feature vector of the positive sample of the same category as local feature vector of the negative sample; is the cosine similarity calculation function.
[0129] In one embodiment, the dynamic multi-scale feature fusion module compresses the global feature dimension to 256 dimensions and the local feature dimension to 64 dimensions through knowledge distillation technology, and adopts a pruning strategy to reduce the calculation amount of fusion weights (parameter compression rate ≥ 50%, calculation power consumption ≤ 5W).
[0130] In one embodiment, the fused features output by the MFF module are input into the dynamic energy optimization module to calculate the energy value and dynamically adjust the classification threshold, as Figure 8 shown.
[0131] Among them, the energy function is defined as:
[0132] ;
[0133] In the formula, is the confidence output of the i-th class of the classification network. The energy value of the ID sample is low, and the energy value of the OOD sample is high; C is the number of classes.
[0134] In one embodiment, the closed-loop feedback mechanism includes:
[0135] The energy loss gradient acts on the gradient backpropagation path through the chain rule: the perturbation direction of the adversarial generation network, the weight calculation of the dynamic multi-scale feature fusion, and the parameter optimization of the classifier, driving the distribution of the generated samples to approximate the true OOD characteristics and enhancing the discriminability of the feature space.
[0136] Parameter update strategy: The exponential weighted moving average (EMA) is used to update the dynamic threshold, and the formula is to ensure smooth adjustment of the threshold.
[0137] Stability monitoring: The training stability is detected in real time through the variance of the loss function. If the variance exceeds the preset threshold, the early stopping mechanism is triggered.
[0138] Furthermore, the threshold T is dynamically updated based on the real-time statistical characteristics of the ID sample energy, and its expression is:
[0139] , ;
[0140] In the formula, is the mean value of the ID sample energy, is the standard deviation of the ID sample energy, is an adjustable hyperparameter.
[0141] In the formula, and are updated through exponential weighted moving average (EMA):
[0142] ;
[0143] In this embodiment, ;
[0144] ;
[0145] In the formula, is the mean value of the sample energy within the current batch distribution; is the standard deviation of the sample energy within the current batch distribution.
[0146] The energy gradient is backpropagated to the DBA module, which can adjust the direction of adversarial perturbation and balance the global and local generation strategies;
[0147] The energy gradient is backpropagated to the MFF module, which can optimize the fusion weight α and enhance the contribution of local features.
[0148] Furthermore, the threshold T is updated after each batch iteration during the training process, and the energy value of the OOD sample is constrained by a piecewise linear loss function:
[0149] ;
[0150] In the formula, is the energy constraint loss of the OOD sample; is the number of OOD samples in the current batch; is the set of OOD samples in the current batch; is the dynamic classification threshold; is the piecewise linear activation function.
[0151] In one embodiment, the model is trained using the SGD optimizer, with an initial learning rate set to 0.07, a momentum parameter of 0.9, a weight decay coefficient of 0.0005, and the learning rate is adjusted by the cosine annealing strategy, with the period set to 50 training epochs;
[0152] During the training process, the energy distribution curves of the ID and OOD samples, the FPR95 (the OOD false positive rate when the ID true positive rate is 95%), and the AUROC (the area under the receiver operating characteristic curve) are recorded for each round to monitor the model convergence status;
[0153] In the test stage, the same normalization process (data augmentation is disabled) is performed on the image to be detected, and the energy value is calculated by inputting it into the trained model If the energy value exceeds the dynamic threshold T, it is determined as an OOD sample; the dynamic threshold is updated based on the real-time statistics of the ID sample energy, and the fluctuations during the training process are smoothed through the exponential weighted moving average (EMA) strategy;
[0154] The auxiliary OOD data can be extended to datasets such as LSUN and iSUN to verify the cross-domain generalization ability, and the hyperparameters can be adjusted according to the scene complexity (for example, setting γ0 = 0.5 and λ = 5 to balance efficiency and accuracy).
[0155] On the other hand, this embodiment discloses an out-of-distribution sample detection system based on an adaptive distribution sphere framework, including:
[0156] A model construction module for constructing an out-of-distribution sample detection model based on the adaptive distribution sphere framework. The adaptive distribution sphere framework includes a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module;
[0157] A sample generation module for inputting the training data into the dynamic distribution boundary adjustment module and generating enhanced samples that meet the global semantic alignment and local detail offset optimization through an adversarial generation network to expand the distribution boundary; the training data comes from a training dataset constructed by in-distribution data and auxiliary out-of-distribution data;
[0158] A feature fusion module for inputting the enhanced samples into the dynamic multi-scale feature fusion module, extracting global features and local features, dynamically calculating the fusion weights based on the distribution differences between the global features and the local features, and fusing the global features and the local features based on the fusion weights to generate fused features;
[0159] A decision and feedback module for inputting the fused features into the dynamic energy optimization module, calculating the energy value through an energy function, dynamically adjusting the classification threshold through a closed-loop feedback mechanism, and applying the energy loss gradient to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module for parameter adjustment;
[0160] A sample recognition module for obtaining the image to be detected and inputting the image to be detected into the out-of-distribution sample detection model for OOD sample recognition, and outputting the recognition result.
[0161] Furthermore, the dynamic distribution boundary adjustment module: Deployed on a GPU accelerator (such as NVIDIA architecture), it supports batch parallel generation of adversarial samples (single-card throughput ≥ 500 samples / second);
[0162] The dynamic multi-scale feature fusion module: Integrates a high-performance inference engine (such as TensorRT), and supports real-time feature fusion (throughput ≥ 1000 samples / second, latency ≤ 1ms);
[0163] The dynamic energy optimization module: Achieves energy calculation through hardware acceleration (such as GPU / FPGA), supports FP16 / FP32 mixed precision, and the dynamic threshold update latency ≤ 2ms (including EMA smoothing time).
[0164] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for relevant parts.
[0165] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An out-of-distribution sample detection method based on an adaptive distributed sphere framework, characterized in that Specifically: Training steps: Construct an out-of-distribution sample detection model based on the adaptive distribution sphere framework, where the adaptive distribution sphere framework includes a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module; Input the training data into the dynamic distribution boundary adjustment module, and generate enhanced samples that meet global semantic alignment and local detail offset optimization through a generative adversarial network to expand the distribution boundary; the training data comes from a training dataset constructed by in-distribution data and auxiliary out-of-distribution data; Input the enhanced samples into the dynamic multi-scale feature fusion module, extract global features and local features, dynamically calculate the fusion weights based on the distribution differences between the global features and the local features, and fuse the global features and the local features based on the fusion weights to generate fused features; Input the fused features into the dynamic energy optimization module, calculate the energy value through an energy function and dynamically adjust the classification threshold using a closed-loop feedback mechanism; And apply the energy loss gradient to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module for parameter adjustment; Detection steps: Obtain the image to be detected; Input the image to be detected into the out-of-distribution sample detection model for OOD sample recognition and output the recognition result.
2. The out-of-distribution sample detection method based on an adaptive distribution sphere framework according to claim 1, wherein It also includes a preprocessing module for normalizing and enhancing the training data or the image to be detected to generate preprocessed data.
3. The out-of-distribution sample detection method based on an adaptive distributed spherical framework according to claim 1, wherein Global semantic alignment quantifies the global semantic difference using the Wasserstein distance, and its expression is: ; wherein, is the probability distribution of in-distribution data; is the probability distribution of out-of-distribution data; is the set of all joint distributions satisfying the marginal distributions of and ; is the expectation operation with respect to the joint distribution γ; is and the L2-norm distance between.
4. The out-of-distribution sample detection method based on an adaptive distributed spherical framework according to claim 1, wherein, Local detail offset optimization enhances the sensitivity to fine-grained anomalies through the Sinkhorn distance, and its expression is: ; Wherein, is the probability distribution of in-distribution data; is the probability distribution of out-of-distribution data; is the set of all joint distributions satisfying the marginal distributions of and ; γ is the joint distribution; is the expected value; is the cost function; is the entropy regularization coefficient; H(γ) is the entropy of the joint distribution γ.
5. The out-of-distribution sample detection method based on an adaptive distributed spherical framework according to claim 1, characterized in that, The dynamic distribution boundary adjustment module also adjusts the weights of global and local distribution modeling through a dynamic weight strategy to achieve hierarchical boundary modeling, and the defined joint optimization objective is: , Wherein, is the global semantic alignment loss, is the local detail optimization loss; is the weight coefficient, which adjusts the balance between the global and local losses over time The expression is as follows: ; wherein, is the initial weight coefficient; is the maximum number of training steps; is the non-linear growth factor.
6. The out-of-distribution sample detection method based on an adaptive distribution sphere framework according to claim 1, wherein The calculation formula for the fusion weight is: ; In the formula, is the activation function; Global feature With local features Wasserstein distance between them; Global feature With local features The Sinkhorn distance between is a very small constant; is the weight coefficient of the regularization term; is the covariance consistency regularization term.
7. A method for detecting out-of-distribution samples based on an adaptive distribution sphere framework according to claim 1, characterized in that, Hierarchical contrast learning is also introduced in the dynamic multi-scale feature fusion module to optimize the discriminability of the feature space, including: Global contrast loss, which enhances the compactness of the feature space by aggregating the high-level semantics of intra-class samples, and its expression is: ; Local contrast loss, which enhances the sensitivity to local anomalies by magnifying the fine-grained differences: ; Wherein, is the number of samples in a batch; is the global feature vector of the -th sample; is the global feature vector of the positive sample of the same category as ; is the global feature vector of the negative sample; is the temperature coefficient of the global contrast loss; sin(⋅) is the cosine similarity calculation function; is the local feature vector of the -th sample; is the local feature vector of the positive sample of the same category as ; is the local feature vector of the negative sample; is the temperature coefficient of the local contrast loss; is the number of negative samples in contrastive learning.
8. A method for detecting out-of-distribution samples based on an adaptive distribution sphere framework according to claim 1, wherein, The energy function is defined as: ; In the formula, is the confidence level output by the model; C is the total number of categories.
9. The out-of-distribution sample detection method based on an adaptive distributed spherical framework according to claim 1, wherein The calculation formula for the classification threshold is: ; In the formula, is the exponentially weighted moving average of the energy of in-distribution samples; is the threshold adjustment coefficient; is the exponentially weighted moving average of the standard deviation of the energy of in-distribution samples.
10. An out-of-distribution sample detection system based on an adaptive distribution sphere framework, characterized in that, Including: A model construction module for constructing an out-of-distribution sample detection model based on the adaptive distribution sphere framework, where the adaptive distribution sphere framework includes a dynamic distribution boundary adjustment module, a dynamic multi-scale feature fusion module, and a dynamic energy optimization module; A sample generation module for inputting the training data into the dynamic distribution boundary adjustment module and generating enhanced samples that meet global semantic alignment and local detail offset optimization through a generative adversarial network to expand the distribution boundary; the training data comes from a training dataset constructed by in-distribution data and auxiliary out-of-distribution data; A feature fusion module, which is used to input the enhanced samples into the dynamic multi-scale feature fusion module, extract global features and local features, dynamically calculate fusion weights based on the distribution differences between the global features and the local features, and fuse the global features and the local features based on the fusion weights to generate fused features; A decision-making and feedback module, which is used to input the fused features into the dynamic energy optimization module, calculate energy values through an energy function and dynamically adjust the classification threshold by using a closed-loop feedback mechanism; and apply the energy loss gradient to the dynamic distribution boundary adjustment module and the dynamic multi-scale feature fusion module for parameter adjustment; A sample recognition module, which is used to obtain a to-be-detected image, input the to-be-detected image into the out-of-distribution sample detection model for OOD sample recognition, and output a recognition result.
Citation Information
Patent Citations
New intention recognition method, training method and device of distributed external monitoring model, and electronic equipment
CN116702048A
Multi-modal data-driven education large model distribution external detection method and multi-modal data-driven education large model distribution external detection system
CN119339398A
Cited By
Defect detection method and device based on residual contrast learning and medium
CN121504918A