Semantic-based out-of-distribution detection method and device
By combining multimodal visual language model and singlemodal visual model and using text labels to train linear layers, the problem of difficulty in effectively utilizing text labels in the prior art is solved, and the out-of-distribution detection performance is significantly improved.
Patent Information
- Application Number
- CN202510109515.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively utilize text labels of data sets within and out of distribution in out of distribution detection, which limits the model's out of distribution detection performance.
The multimodal visual language model and a singlemodal visual model are used to map the output of the singlemodal visual model to the same dimensions of the multimodal visual language model by superimposing a linear layer, and the linear layer is trained using a loss function to enhance the out-of-distribution detection performance.
By utilizing text data, the performance of off-distribution detection is significantly improved, especially in large-scale off-distribution detection benchmarks, which outperform state-of-the-art methods.
Smart Images

Figure CN120014645A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of out-of-distribution detection in computer vision, and in particular relates to a semantic-based out-of-distribution detection method and device. Background Art
[0002] Machine learning models have become the cornerstone of modern computational intelligence in various fields, including natural language processing, computer vision, and data analysis. However, most of these models are based on closed-world assumptions.
[15] The training is done under the assumption that the test data is independently sampled from the same distribution as the training data, and when deployed to an open world scenario [1] (such as autonomous driving [7]
[12] , safety authentication system [2] and medical analysis
[13] ), the test data may deviate from the training data. In this case, the model also needs to reliably recognize and respond to samples that have not been seen in training, a task known as out-of-distribution detection that has attracted increasing attention in the research community. The standard approach to out-of-distribution detection is to derive a score from a model trained on an in-distribution dataset to distinguish between in-distribution and out-of-distribution data.
[0003] In the era of deep learning, Hendrycks and Gimpel [3] For the first time, they provided a comprehensive description of the out-of-distribution detection task. They proposed a method to detect out-of-distribution samples using the maximum softmax probability as a scoring function. Currently, the mainstream paradigm of out-of-distribution detection adopts a training-free strategy. They design a scoring function by obtaining the output of the pre-trained model to determine whether the test sample is an out-of-distribution sample. This scoring function is mainly based on logit [4][8]
[14] 、probability [3][5] or feature [6]
[17] , and assign higher scores to in-distribution samples and lower scores to out-of-distribution samples. In addition to these training-free strategies, there are also some methods that involve training models to enhance out-of-distribution detection performance. However, they usually require additional out-of-distribution image data. [4][9]
[10]
[11]
[16] To assist the fine-tuning process of the model, these methods utilize additional out-of-distribution data to promote the regularization of the decision boundary between in-distribution and out-of-distribution data throughout the training process. The additional out-of-distribution image data limits their practical use as the model performance may be affected by the correlation between the provided (extra) and true out-of-distribution samples.
[0004] Recently, the rise of multimodal vision-language models has made breakthrough progress in the field of computer vision. Inspired by these methods, we can try to leverage textual data, i.e., labels of in-distribution and out-of-distribution datasets, to enhance the performance of out-of-distribution detection. Considering that in real-world scenarios, we usually have access to the labels of some out-of-distribution samples, which can be classified as "known out-of-distribution labels" and the rest as "unknown out-of-distribution labels", we can leverage these real-world in-distribution and known out-of-distribution labels to enhance the performance of out-of-distribution detection. Taking the autonomous driving scenario as an example, pedestrians and vehicles constitute in-distribution labels. Meanwhile, animals that may appear on the road (e.g., dogs and deer) are considered as known out-of-distribution labels. However, there are countless objects that may be encountered while driving (e.g., poles, fire hydrants, stones), which obviously cannot be fully enumerated. Due to the lack of prior knowledge about what these objects may be, there is no label information available for them.
[0005] Therefore, considering only how to leverage “known” information, how to use text labels of both in-distribution and known out-of-distribution datasets to enhance the model’s out-of-distribution detection capability is currently an unsolved problem.
[0006] [1]Nick Drummond and Rob Shearer.The open world assumption.In eSIWorkshop,2006.
[0007] [2]Ivan Evtimov, Kevin Eykholt, Earlence Fernandes, Tadayoshi Kohno, BoLi, Atul Prakash, Amir Rahmati, and Dawn Song. Robust physical-world attacks on machine learning models. arXiv preprint arXiv:1707.08945, 2017.
[0008] [3]Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In ICLR, 2016.
[0009] [4]Dan Hendrycks,Mantas Mazeika,and Thomas Dietterich.Deep anomalydetection with outlier exposure.In ICLR,2018.
[0010] [5]Rui Huang and Yixuan Li.Mos:Towards scaling out-of-distributiondetection for large semantic space.In CVPR,2021.
[0011] [6]Rui Huang,Andrew Geng,and Yixuan Li.On the importance of gradientsfor detecting distributional shifts in thewild.In NeurIPS,2021.
[0012] [7]Sanghun Jung,Jungsoo Lee,Daehoon Gwak,Sungha Choi,and JaegulChoo.Standardized max logits:A simple yeteffective approach for identifyingunexpected road obstaclesin urban-scene segmentation.In ICCV,2021.
[0013] [8]Weitang Liu,Xiaoyun Wang,John Owens,and Yixuan Li.Energy-basedout-of-distribution detection.In NeurIPS,2020.
[0014] [9]Yifei Ming,Ying Fan,and Yixuan Li.Poem:Out-of-distributiondetection with posterior sampling.In ICML,2022.
[0015]
[10] Sina Mohseni,Mandar Pitale,JBS Yadawa,and ZhangyangWang.Self-supervised learning for generalizable out-of-distribution detection.In AAAI,2020.
[0016]
[11] Aristotelis-Angelos Papadopoulos,Mohammad Reza Rajati,NazimShaikh,and Jiamian Wang.Outlier exposure withconfidence control for out-of-distribution detection.Neurocomputing,2021.
[0017]
[12] Xuanchi Ren,Tao Yang,Li Erran Li,Alexandre Alahi,and QifengChen.Safety-aware motion prediction with unseenvehicles for autonomousdriving.In ICCV,2021.
[0018]
[13] Thomas Schlegl,Philipp Sebastian M Waldstein,UrsulaSchmidt-Erfurth,and Georg Langs.Unsupervisedanomaly detection with generativeadversarial networks toguide marker discovery.In IMPI,2017.
[0019]
[14] Haoqi Wang,Zhizhong Li,Litong Feng,and Wayne Zhang.Vim:Out-of-distribution with virtual-logit matching.InCVPR,2022.
[0020]
[15] Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. Generalized out-of-distribution detection: A survey. arXiv preprint arXiv:2110.11334, 2021.
[0021]
[16] Qing Yu and Kiyoharu Aizawa.Unsupervised out-of-distributiondetection by maximum classifier discrepancy.InICCV, 2019.
[0022]
[17] Yeonguk Yu, Sungho Shin, Seongju Lee, Changhyun Jun, and Kyoobin Lee. Block selection method for using featurenorm in out-of-distribution detection. In CVPR, 2023. Summary of the invention
[0023] The present invention is made to solve the above-mentioned problem, and its purpose is to provide a semantic-based out-of-distribution detection method and device.
[0024] The present invention provides a semantic-based out-of-distribution detection method, which has the following characteristics and includes the following steps: S10, inputting an image and text into a multimodal visual language model and outputting a logit, then inputting the same image into a unimodal visual model and outputting a logit of a dimension different from the output result of the multimodal visual language model; S20, mapping the logit output by the unimodal visual model to the same dimension as the logit output by the multimodal visual language model through a linear layer superimposed on the unimodal visual model, and then processing with a loss function to train the linear layer; S30, using the multimodal out-of-distribution detection module composed of the unimodal visual model, the linear layer, the multimodal visual language model and the loss function trained in step S20 as a test model, inputting an in-distribution image or an out-of-distribution image into the test model for testing, and realizing out-of-distribution detection by using the logit output by the test model, the probability information corresponding to the logit output by the test model and the output feature information of the internal layer of the test model.
[0025] The semantic-based out-of-distribution detection method provided by the present invention may also have the following features: wherein, in step S10, the multimodal visual language model maps the image and text into the same feature space, and then measures the distance between the image and text data in the feature space according to the cosine similarity, and finally takes the similarity between the data of a single image and multiple texts as the logit of a single image.
[0026] The semantic-based out-of-distribution detection method provided by the present invention may also have the following features: wherein, in step S10, the unimodal visual model is a pre-trained ViT model, and the multimodal visual language model is a pre-trained CLIP model, the CLIP model processes text and images through its text encoder and image encoder respectively, and the linear layer is randomly initialized.
[0027] The semantic-based out-of-distribution detection method provided by the present invention may also have the following characteristics: wherein, in step S20, the mapping formula is: is the logit output by the ViT model, for The weight of for The bias, for The corresponding logit after mapping.
[0028] The semantic-based out-of-distribution detection method provided by the present invention may also have the following characteristics: wherein the loss function is processed by And the logit output of the CLIP model is used to train the linear layer, and the loss function is: α represents the weight parameter, represents the knowledge distillation loss, represents the smooth L1 loss function, represents the logit output of the CLIP model, represents the classification loss, represents the cross entropy loss function, represents the within-distribution label.
[0029] The semantic-based out-of-distribution detection method provided by the present invention may also have the following feature: wherein, in step S20, the linear layer participates in back propagation for training, and the unimodal visual model and the multimodal visual language model do not participate in back propagation during training.
[0030] The semantic-based out-of-distribution detection method provided by the present invention may also have the following feature: in step S30, when the test model performs out-of-distribution detection, only the ViT model and the linear layer participate in the test.
[0031] The semantic-based out-of-distribution detection method provided by the present invention may also have the following features: wherein, in step S30, the test model implements out-of-distribution detection through the following sub-steps: S31, using strengthen The classification ability of the in-distribution data set is as follows: C in represents the number of labels in the distribution, θ represents the weight parameter; S32, Use softmax transformation to utilize probability information; S33, use the ViT model as a feature extractor to extract features of the training set and test samples, and calculate the average distance between the features of the test samples and the features of the nearest k samples in the training set Thus, the feature information is utilized; S34, and the final scoring function for out-of-distribution detection is obtained: p i is the probability of representing probability information, Represents the features extracted by the feature extractor.
[0032] The semantic-based out-of-distribution detection method provided by the present invention may also have the following feature: in step S34, for a given image, an image whose scoring function value is higher than a threshold is judged as an in-distribution image, and an image whose scoring function value is lower than the threshold is judged as an out-of-distribution image.
[0033] The present invention also provides a semantic-based out-of-distribution detection device, having the following characteristics: it uses any of the aforementioned semantic-based out-of-distribution detection methods, including: a teacher branch, which is a multimodal visual language model; a student branch, which includes a unimodal visual model and a linear layer superimposed thereon; a loss function, which is used to process the output results of the teacher branch and the student branch to train the linear layer; and a scoring function, which is used to score the logit output by the test model composed of the trained teacher branch, the student branch and the loss function, the probability information corresponding to the logit output by the test model, and the output feature information of the internal layer of the test model to achieve out-of-distribution detection.
[0034] Functions and Effects of the Invention
[0035] The present invention proposes a novel out-of-distribution detection method and device, comprising: (1) a multimodal out-of-distribution detection module (a multimodal visual language model, a unimodal visual model and a linear layer superimposed thereon, and a loss function) for learning text data to enhance out-of-distribution detection capabilities; and (2) a scoring function tailored for the multimodal out-of-distribution detection module for out-of-distribution detection using logit, probability, and feature.
[0036] The present invention performs well on a large-scale out-of-distribution detection benchmark, significantly surpassing the state-of-the-art methods, demonstrating the effectiveness of the out-of-distribution detection method proposed in the present invention.
[0037] Extensive ablation experiments are performed to explain the different modules of the proposed method and the impact of hyper-parameters is discussed. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic diagram of a distribution inner label, a known distribution outer label, and an unknown distribution outer label in an embodiment of the present invention;
[0039] Figure 2 is a schematic diagram of a semantic-based out-of-distribution detection method in an embodiment of the present invention;
[0040] Figure 3 is the result of using different hyperparameters θ in the test examples of the present invention;
[0041] Figure 4 is the result of using different hyperparameters k in the test examples of the present invention;
[0042] Figure 5 It is the result of using different hyperparameters α in the test examples of the present invention;
[0043] Figure 6 It is a comparison chart of the experimental results in the test example of the present invention and the advanced experimental results in the industry;
[0044] Figure 7 These are the results of using different known out-of-distribution data sets in the test examples of the present invention. DETAILED DESCRIPTION
[0045] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate a semantic-based out-of-distribution detection method and device of the present invention.
[0046] <Example>
[0047] This embodiment provides a semantic-based out-of-distribution detection device, including a teacher branch, a student branch, a loss function, and a scoring function.
[0048] The teacher branch is a multimodal visual language model, which takes pictures and texts as input and outputs logit. Specifically, in this embodiment, the multimodal visual language model is a pre-trained CLIP model.
[0049] The student branch consists of a unimodal vision model and a linear layer superimposed on it. The student branch takes the image as input and outputs logit. Specifically, in this embodiment, the unimodal visual model is a pre-trained ViT model; the linear layer It is randomly initialized and used to map the output of the student branch to the same dimension as the output of the teacher branch.
[0050] The loss function is used to process the output results of the teacher branch and the student branch to train the linear layer
[0051] The scoring function is used to score the logit output by the test model composed of the trained teacher branch, student branch and loss function, the probability information corresponding to the logit output by the test model, and the output feature information of the internal layer of the test model to achieve out-of-distribution detection.
[0052] Figure 1 It is a schematic diagram of the in-distribution labels, known out-distribution labels and unknown out-distribution labels in the embodiment of the present invention.
[0053] like Figure 1 As shown, in real application scenarios, there are in-distribution labels and known out-of-distribution labels available for use.
[0054] Figure 2 It is a schematic diagram of a semantic-based out-of-distribution detection method in an embodiment of the present invention.
[0055] like Figure 2 As shown, this embodiment also provides a semantic-based out-of-distribution detection method, which uses the semantic-based out-of-distribution detection device in this embodiment, and uses in-distribution labels and known out-of-distribution labels to enhance the out-of-distribution detection performance, including the following steps:
[0056] S10, input the image and text into the multimodal visual language model CLIP, and output logit after processing through the text encoder and image encoder The same image is then fed into the unimodal vision model ViT and the output is the same as Logit of different dimensions
[0057]
[0058] The multimodal visual language model CLIP maps images and texts into the same feature space, and then measures the distance between image and text data in the feature space based on cosine similarity. Finally, the similarity between a single image and multiple text data is used as the logit of a single image.
[0059] The models used in out-of-distribution detection are basically unimodal visual models that cannot process text information. Therefore, the text information must first be converted into a form that can be used by the unimodal visual model ViT using the text encoder of the multimodal visual language model CLIP. This embodiment uses the multimodal visual language model CLIP to convert the out-of-distribution knowledge in the text information into a logit form through its multimodal capabilities for use by the unimodal visual model ViT during training.
[0060] S20, through a linear layer superimposed on the unimodal visual model ViT The output of the unimodal visual model ViT Mapped to the output of the multimodal visual language model CLIP The same dimension is then processed using the loss function to train the linear layer During training, only the linear layer Participate in back-propagation training, the unimodal visual model ViT and the multimodal visual language model CLIP do not participate in back-propagation, and their parameters are not updated.
[0061] The mapping formula is:
[0062] in, for The weight of for The bias, for The corresponding logit after mapping.
[0063] The loss function is:
[0064] Among them, α is the weight parameter used for weighted averaging, Indicates that To approximate The knowledge distillation loss (the smooth L1 loss function is insensitive to outliers and abnormal values, the gradient changes are relatively small, and it is more robust during training), represents the smooth L1 loss function, represents the classification loss, represents the cross entropy loss function, represents the within-distribution label.
[0065] In the above formula, the classification loss Improved by cross entropy loss function The classification accuracy of the data within the distribution allows it to learn more knowledge about the data within the distribution and give More in-distribution knowledge to improve its out-of-distribution detection ability.
[0066] S30, the multimodal out-of-distribution detection module is composed of a teacher branch, a student branch and a loss function. The multimodal out-of-distribution detection module trained in step S20 is used as a test model, and an in-distribution image or an out-of-distribution image is input into the test model for testing. Only the student branch participates in the test process, including the following sub-steps S31 to S34:
[0067] S31, using the logit output of the last layer of the test model:
[0068] use strengthen The ability to classify data sets within the distribution of and To add, use The ability to classify data within a distribution is further enhanced The out-of-distribution detection capability of is because the two dimensions are different, and only the overlapping parts of the dimensions are added, as shown in the following formula:
[0069]
[0070] Among them, C in represents the number of labels in the distribution, and θ represents the weight parameter.
[0071] S32, post-processing results of logit output from the last layer of the test model:
[0072] Will Use softmax conversion to process it into probability p to utilize probability information.
[0073] S33, using the internal layer output of the test model:
[0074] The ViT model is used as a feature extractor to extract features of the training set and test samples, and the average distance between the features of the test samples and the features of the nearest k samples in the training set is calculated. So as to utilize feature information.
[0075] S34, obtain the final scoring function for out-of-distribution detection:
[0076]
[0077] Among them, p i is the probability of representing probability information, Represents the features extracted by the feature extractor.
[0078] For a given image, an image whose scoring function value is higher than a certain artificially determined threshold is judged as an in-distribution image, and an image whose scoring function value is lower than the threshold is judged as an out-of-distribution image.
[0079] <Test example>
[0080] This test example uses a semantics-based out-of-distribution detection method and device in the embodiment to perform corresponding tests.
[0081] The model accepts 224×224 fixed-size images as input for the teacher branch and the student branch. Images of various sizes are scaled to the above sizes through the same preprocessing operation before being input into the multimodal out-of-distribution detection module (consisting of the teacher branch, the student branch, and the loss function).
[0082] like Figure 1 As shown, the out-of-distribution detection method of the test example of the present invention uses a multimodal out-of-distribution detection module and a scoring function.
[0083] The number of labels in the distribution is C in , the number of labels outside the known distribution is C out .
[0084] For the teacher branch, the text encoder of the multimodal visual language model CLIP accepts the input text and obtains C in +C out The image encoder accepts the input image and obtains an image feature, which is consistent with C in +C out The cosine similarity is calculated for each text feature, and the length is C in +C out The vector of
[0085] For the student branch, it accepts the input image and obtains a length of C in Vector Then through the linear layer The mapping of The length is C in +C out ,and The lengths of the two branches are the same, so that they can better learn the knowledge of the out-of-distribution data contained in the teacher branch.
[0086] For the multimodal visual language model CLIP, this test case uses clip-vit-base-patch16 in the CLIP library; for the unimodal visual model ViT, this test case uses vit_base_patch16_224 in the timm library; for the linear layer For example, it consists of a fully connected layer and an activation function ReLU.
[0087] During training, the AdamW optimizer was used, with a base learning rate of 0.01, and then multiplied by a cosine decay factor in the range of [0.01, 1], so that the learning rate is reduced as the number of training steps increases to accelerate the convergence of the neural network. In-distribution labels and known out-of-distribution labels are converted into prompts, such as "a photo of a [CLS] in ID / OOD dataset", where [CLS] is the corresponding label. The hyperparameter α used to balance the loss function is 0.2. Finally, the multimodal out-of-distribution detection module was trained for 80 epochs.
[0088] During testing, it is used to strengthen The in-distribution knowledge of θ=0.2, and the k=2 used to extract feature knowledge.
[0089] Figure 3 is the result of using different hyperparameters θ in the test examples of the present invention; Figure 4 is the result of using different hyperparameters k in the test examples of the present invention; Figure 5 These are the results of using different hyperparameters α in the test examples of the present invention.
[0090] like Figure 3 to Figure 5 As shown, when the hyperparameters θ=0.5, k=2, and α=0.2, the out-of-distribution detection result of this test case is optimal.
[0091] Figure 6 It is a comparison chart of the experimental results in the test example of the present invention and the advanced experimental results in the industry.
[0092] like Figure 6 As shown, the experimental results in the test examples of the present invention are significantly better than the most advanced results in the industry, proving the effectiveness of the out-of-distribution detection method in the embodiment.
[0093] This test case uses AUROC (Area Under the Receiver Operating Characteristic curve) and FPR95 (FPR at 95% TPR, false positive rate when the true positive rate is 95%) to comprehensively measure the performance of the model's out-of-distribution detection, with the in-distribution samples predicted by the model as positive samples and the out-of-distribution samples as negative samples. AUROC reflects the performance of the classifier through the size of the area between the receiver operating characteristic curve and the coordinate axis. Its numerical meaning is that if a positive sample and a negative sample are randomly selected, the classifier correctly judges the probability that the value of the positive sample is higher than that of the negative sample, that is, the classifier gives a higher probability to the in-distribution sample than the out-of-distribution sample. Therefore, the maximum value of AUROC is 1 and the minimum value is 0. FPR95 is another evaluation indicator, which requires that when the true positive rate is 95%, the lower the false positive rate, the better.
[0094] The AUROC (Area Under the Receiver Operating Characteristic curve) and FPR95 (FPR at 95% TPR, false positive rate when the true positive rate is 95%) of this test case when the in-distribution dataset is ImageNet-1k and the out-of-distribution datasets are SUN, Places, OpenImage-O, and iNatualist are shown in Table 1 below.
[0095] Table 1 (AUROC and FPR95 when the in-distribution dataset is ImageNet-1k and the out-of-distribution dataset is SUN, Places, OpenImage-O and iNatualist)
[0096]
[0097] The classification performance of this test case on the in-distribution dataset is shown in Table 2 below.
[0098] Table 2 (Classification performance on in-distribution datasets)
[0099]
[0100] The performance of this test case when using different CLIP and ViT models is shown in Table 3 below.
[0101] Table 3 (Performance when using different CLIP and ViT models)
[0102]
[0103] As shown in Tables 1, 2, and 3, the test model in this test case performs well on the large-scale out-of-distribution detection benchmark and can significantly improve the performance of out-of-distribution detection.
[0104] Figure 7 These are the results of using different known out-of-distribution data sets in the test examples of the present invention.
[0105] like Figure 7 As shown, the semantic-based out-of-distribution detection method in the test case of the present invention is sufficiently robust to different known out-of-distribution datasets.
[0106] Functions and Effects of the Embodiments
[0107] This embodiment and test case use ImageNet-1k as the in-distribution dataset and four different commonly used out-of-distribution datasets for performance evaluation, following the large-scale out-of-distribution detection benchmark, and show leading performance on these datasets. It is worth emphasizing that this embodiment and test case not only achieve excellent performance on known out-of-distribution datasets with additional text information, but also on unknown out-of-distribution datasets without additional information during training. In addition, the test case performs extensive ablation analysis to gain a deeper understanding of the technical solution of the present invention.
[0108] This embodiment proposes a novel out-of-distribution detection method and apparatus, including: (1) a multimodal out-of-distribution detection module (a multimodal visual language model, a unimodal visual model and a linear layer superimposed thereon, and a loss function) for learning text data to enhance out-of-distribution detection capabilities. (2) a scoring function tailored for the multimodal out-of-distribution detection module that uses logit, probability, and feature for out-of-distribution detection.
[0109] This embodiment and test example performed well on a large-scale out-of-distribution detection benchmark, significantly surpassing the most advanced methods, proving the effectiveness of the out-of-distribution detection method proposed in this embodiment.
[0110] Extensive ablation experiments are performed on the examples and test cases to explain the different modules of the proposed method and discuss the impact of hyper-parameters.
[0111] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A semantic-based out-of-distribution detection method, characterized in that: The following steps are involved: S10, inputting the image and text into a multimodal visual language model and outputting a logit, and then inputting the same image into a unimodal visual model and outputting a logit of a different dimension from the output result of the multimodal visual language model; S20, mapping the logit output by the unimodal visual model to the same dimension as the logit output by the multimodal visual language model through a linear layer superimposed on the unimodal visual model, and then processing using a loss function to train the linear layer; S30, using the multimodal out-of-distribution detection module composed of the unimodal visual model trained in step S20, the linear layer, the multimodal visual language model and the loss function as a test model, inputting in-distribution images or out-of-distribution images into the test model for testing, and implementing out-of-distribution detection by using the logit output by the test model, the probability information corresponding to the logit output by the test model, and the output feature information of the internal layer of the test model.
2. The semantic-based out-of-distribution detection method according to claim 1, characterized in that: in, In step S10, the multimodal visual language model maps the image and the text into the same feature space, and then measures the distance between the image and text data according to the cosine similarity in the feature space, and finally takes the similarity between the single image and the multiple text data as the logit of the single image.
3. The semantic-based out-of-distribution detection method according to claim 1, characterized in that: in, In step S10, the unimodal visual model is a pre-trained ViT model. The multimodal visual language model is a pre-trained CLIP model. The CLIP model processes the text and the image through its text encoder and image encoder respectively. The linear layer is randomly initialized.
4. The semantic-based out-of-distribution detection method according to claim 3, characterized in that: in, In step S20, the mapping formula is: is the logit output by the ViT model, for The weight of for The bias of for The corresponding logit after mapping.
5. The semantic-based out-of-distribution detection method according to claim 4, characterized in that: in, The loss function is processed by and the logit output of the CLIP model to train the linear layer, The loss function is: α represents the weight parameter, represents the knowledge distillation loss, represents the smooth L1 loss function, represents the logit output of the CLIP model, represents the classification loss, represents the cross entropy loss function, represents the within-distribution label.
6. The semantic-based out-of-distribution detection method according to claim 1, characterized in that: in, In step S20, the linear layer participates in back propagation training. The unimodal visual model and the multimodal visual language model do not participate in back propagation during training.
7. The semantic-based out-of-distribution detection method according to claim 4, characterized in that: in, In step S30, when the test model performs out-of-distribution detection, only the ViT model and the linear layer participate in the test.
8. The semantic-based out-of-distribution detection method according to claim 7, characterized in that: in, In step S30, the test model implements out-of-distribution detection through the following sub-steps: S31, using strengthen The classification ability of the in-distribution data set is as follows: C in represents the number of labels in the distribution, and θ represents the weight parameter; S32, Use softmax transformation to utilize probability information; S33, using the ViT model as a feature extractor to extract features of the training set and test samples, and calculating the average distance between the features of the test samples and the features of the nearest k samples in the training set So as to utilize feature information; S34, obtain the final scoring function for out-of-distribution detection: p i is the probability of representing probability information, Represents the features extracted by the feature extractor.
9. The semantic-based out-of-distribution detection method according to claim 8, characterized in that: in, In step S34, for a given image, an image whose value of the scoring function is higher than a threshold is determined as an in-distribution image, and an image whose value of the scoring function is lower than the threshold is determined as an out-of-distribution image.
10. A semantic-based out-of-distribution detection device, characterized in that: The semantic-based out-of-distribution detection method according to any one of claims 1 to 9 is used, comprising: The teacher branch is the multimodal visual language model; a student branch, comprising the unimodal visual model and the linear layer superimposed thereon; A loss function, used to process the output results of the teacher branch and the student branch to train the linear layer; and A scoring function is used to score the logit output by the test model composed of the trained teacher branch, the student branch and the loss function, the probability information corresponding to the logit output by the test model, and the output feature information of the internal layer of the test model to achieve out-of-distribution detection.
Citation Information
Cited By
Intelligent urban road waterlogging depth classification method and system based on knowledge distillation and diffusion model
CN120852842A