Visual language model distribution external sample detection method based on multi-modal abnormal value synthesis

Frozen prototype vectors of images and texts are generated through a multimodal outlier synthesis method. Combined with visual language model training, a multimodal prototype matching mechanism is constructed, which solves the problem of insufficient modal utilization in multimodal out-of-distribution detection and improves the generalization ability and robustness of the model.

CN120687829APending Publication Date: 2025-09-23SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510671596.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing multimodal out-of-distribution detection methods suffer from insufficient modal utilization and incomplete negative sample construction in visual language models, resulting in unclear boundaries between in-distribution and out-of-distribution image-text alignment features, and unable to effectively and cost-effectively discriminate out-of-distribution data.

Method used

A method based on multimodal outlier synthesis is adopted to extract features through image encoder and text encoder, generate frozen prototype vectors and learnable vectors of images and texts, combine with visual language model training, construct a multimodal prototype matching mechanism, design out-of-distribution discrimination scores, and improve detection capabilities.

Benefits of technology

It improves the generalization and robustness of the visual language model in multimodal out-of-distribution detection tasks, and is able to generate high-quality out-of-distribution samples with a small number of in-distribution samples, thereby improving the ability to recognize out-of-distribution data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687829A_ABST
    Figure CN120687829A_ABST
Patent Text Reader

Abstract

The invention discloses a visual language model out-of-distribution sample detection method based on multi-modal abnormal value synthesis, which comprises the following steps of: based on a small number of in-distribution image samples, generating high-quality multi-modal out-of-distribution samples which are closely related to in-distribution semantics and are in the form of images and texts through semantic feature analysis and sampling of image contents; then constructing an image prototype and a text prototype of a fusion distribution inner sample and a synthesis distribution outer sample, in the reasoning process, adopting an image and text bimodal prototype matching mechanism, performing similarity calculation at the same time, and on this basis, proposing a multi-modal prototype matching score to comprehensively evaluate the similarity between the to-be-tested sample and the distribution inner category. According to the method, the out-of-distribution samples with the image and text labels can be automatically generated based on a small number of in-distribution samples, and the generalization ability and robustness of the model in a multi-modal out-of-distribution detection task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of out-of-distribution data detection, and specifically relates to a method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis. Background Art

[0002] Deep Neural Networks (DNNs) are increasingly being used in safety-sensitive fields such as autonomous driving, intelligent manufacturing, and medical diagnosis. The reliability of their decisions is directly related to system safety. However, the widespread presence of out-of-distribution (OOD) samples in real-world scenarios has become a key challenge hindering the implementation of deep models. OOD samples refer to input data that deviates from the training data distribution (such as untrained road signs in autonomous driving and rare lesion types in medical imaging). If the model cannot accurately identify such samples, it may output incorrect predictions with high confidence, leading to catastrophic consequences.

[0003] To address this issue, academia and industry have proposed a variety of OOD detection methods. Early research was mainly based on single-modal data (such as pure images or pure text), and achieved detection through the following two technical routes:

[0004] Posterior probability-driven methods use the model's output probability distribution characteristics (such as the maximum softmax probability threshold) to determine whether an OOD sample is present. Auxiliary abnormal sample introduction methods enhance the model's ability to perceive unknown distributions by introducing external abnormal data or generating adversarial samples. However, these methods rely heavily on the data characteristics of a single modality and suffer from low detection accuracy and insufficient generalization in complex multimodal scenarios (such as instruction comprehension involving both images and text).

[0005] In recent years, multimodal visual language models such as CLIP have demonstrated outstanding performance in image-text understanding, driving the development of multimodal OOD detection. Multimodal out-of-distribution detection methods based on visual language models mainly include zero-shot methods, small-shot methods, and fully supervised methods. The zero-shot method does not require ID samples but is limited by semantic shift. While fully supervised methods offer higher performance, they have a high data cost and suffer from insufficient generalization. Therefore, small-shot out-of-distribution detection methods are an excellent option. Existing multimodal out-of-distribution detection methods primarily construct out-of-distribution image or text samples for training. However, these methods focus on synthesizing out-of-distribution samples from a single modality and lack relevant technologies for fully synthesizing and utilizing out-of-distribution samples from both image and text modalities.

[0006] In summary, existing multimodal out-of-distribution detection systems and methods often only synthesize and introduce out-of-distribution samples of a single modality, resulting in insufficient modality utilization and incomplete negative sample construction. As a result, there is an unclear boundary between in-distribution and out-of-distribution image-text alignment features, which makes it impossible for visual language models to effectively and cost-effectively discriminate out-of-distribution data. Summary of the Invention

[0007] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis. The method automatically generates out-of-distribution samples with image and text labels based on a small number of in-distribution samples, thereby improving the generalization ability and robustness of the model in multimodal out-of-distribution detection tasks.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] In a first aspect, the present invention provides a method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis, comprising the following steps:

[0010] Select a multimodal information synthesis method based on different original in-distribution samples, and synthesize out-of-distribution samples. Combine the original in-distribution samples with the out-of-distribution samples. Each sample includes multiple categories of images and corresponding class labels to obtain an image set and a class label set.

[0011] An image encoder is used to extract image prototype features from an image set. After calculating the mean, frozen prototype vectors of images of various samples are obtained. Frozen prototype vectors of images of the same category are matched with learnable vectors of the same dimension, and after combining, various image prototypes are obtained. The class label and the corresponding learnable parameter prefix are combined into a text prompt. The text prompt is then subjected to feature extraction using a text encoder to obtain text features of various samples.

[0012] Various image prototypes and text features of various samples are used as common input, and a visual language model is trained to obtain the final detection results based on the out-of-distribution detection score.

[0013] As an optimal technical solution, the multimodal information synthesis method includes an out-of-distribution sample synthesis method based on recognition and an out-of-distribution sample synthesis method based on segmentation; the out-of-distribution sample synthesis method based on recognition is suitable for scenarios where the in-distribution image has clear background semantic information, and the out-of-distribution sample synthesis method based on segmentation is suitable for scenarios where the in-distribution image has a complex background or fuzzy semantics.

[0014] As a preferred technical solution, the identification-based out-of-distribution sample synthesis method includes:

[0015] Use the pre-trained image annotation model to perform object recognition on each in-distribution image and obtain the set of all recognizable semantic labels in the image;

[0016] Remove the labels in the semantic label set that overlap with the in-distribution labels, and use the remaining semantic labels as candidate out-of-distribution labels;

[0017] Perform several random cropping operations on each in-distribution image to generate multiple local image blocks, and construct text templates based on candidate out-of-distribution labels.

[0018] The local image segmentation and the corresponding text template are input into the pre-trained visual language model, the similarity score between the image and the text is calculated, and the local image segmentation with the highest score is selected as the out-of-distribution image. The out-of-distribution images and their corresponding out-of-distribution labels constitute the out-of-distribution sample set.

[0019] As a preferred technical solution, the segmentation-based out-of-distribution sample synthesis method includes:

[0020] Use an unsupervised pre-trained image segmentation model to perform full image segmentation on each in-distribution image, obtain image segments of multiple background regions, and unify the image segments;

[0021] The original in-distribution image and its category label are input into the object detection model to obtain the bounding box of the foreground object. The overlapping area between each segmented image and the foreground area is calculated. If the overlap ratio is greater than a set threshold, the image segment is removed. The remaining image segments are divided into several categories using a clustering algorithm, and the number of cluster categories is set to be the same as the number of in-distribution categories.

[0022] The average image feature value of each clustered image segment is calculated, and the image segments are matched with the noun or adjective candidate set in the WordNet lexicon. The word with the highest similarity is selected as the text label of the segment. If the image segments of different clustered segments generate the same text label, they are merged into the same segment.

[0023] The image segments of each cluster category are regarded as out-of-distribution images, and the corresponding text labels are regarded as corresponding out-of-distribution labels. The similarity between the out-of-distribution images and the original in-distribution images is calculated, and the categories with the lowest similarity are selected as the final out-of-distribution samples.

[0024] As a preferred technical solution, multiple image encoders are used, and parameters are shared among the image encoders.

[0025] As a preferred technical solution, the frozen prototype vectors of images of the same category are matched with the learnable vectors of the same dimension, specifically:

[0026] Let the learnable vector be The image frozen prototype vector is recorded as The initial values ​​of both are 0. The two are spliced ​​together to obtain the image prototype of category i, as shown in the following formula:

[0027] As a preferred technical solution, the method of combining the class label and the corresponding learnable parameter prefix into a text prompt includes:

[0028] Set the class label text to the non-learnable part, set the learnable parameter prefix to the learnable part, set the vector length of the prompt text, and initialize the learnable part with a Gaussian distribution;

[0029] In the above example, the in-distribution text prompts and out-of-distribution text prompts are input into the text encoder to obtain the text features of the in-distribution samples and the text features of the out-of-distribution samples.

[0030] As a preferred technical solution, the training of the visual language model includes the following steps:

[0031] Inputting training set data samples into an image encoder to obtain image feature vectors; the training set data samples include in-distribution samples and out-of-distribution samples generated based on the in-distribution samples;

[0032] Calculate the total loss of the common input and image feature vector until the parameters converge and the training is completed;

[0033] The total loss includes the in-distribution textual hint loss Out-of-distribution text hint loss and image prototype loss The total loss is as follows:

[0034]

[0035] Among them, λ1 and λ2 are the weight coefficients of the loss function.

[0036] As a preferred technical solution, it includes:

[0037] The in-distribution text prompt loss is as follows:

[0038]

[0039] Among them, M represents the number of categories of samples within the distribution, N1 represents the number of categories of synthesized out-of-distribution samples, represents the similarity between the training image embedding and the text category hint embedding within the m-th distribution, represents the similarity between the training image embedding and the n1th out-of-distribution text category hint embedding, s * represents the similarity between the training image embedding and the text hint embedding of the corresponding real category;

[0040] The out-of-distribution text hint loss is as follows:

[0041]

[0042] The image prototype loss is as follows:

[0043]

[0044] in, represents the similarity between the embedding of the training image and the embedding of the image prototype within the m-th distribution, represents the similarity between the training image embedding and the nth image prototype embedding, s′ * Indicates the similarity between the training image embedding and the image prototype embedding of the corresponding real category.

[0045] As a preferred technical solution, obtaining the final test result based on the out-of-distribution test score includes:

[0046] For each test image, we calculate the similarity between the image embedding obtained by the visual encoder and the text hint embedding and image prototype embedding of each category, and calculate the multimodal matching score S MPM As the final out-of-distribution detection score, the multimodal matching score S MPM As follows:

[0047]

[0048] Where M is the number of in-distribution categories and N is the number of synthetic out-of-distribution categories. i 、s j and s k is the similarity between the embedding of the test image and the embedding of the text prompt according to the multimodal matching score, s′i, s′j, s′k are the similarities between the embedding of the test image and the embedding of the image prototype;

[0049] According to the multimodal matching score S MPM The size of determines whether the test sample belongs to the sample within the distribution or the sample outside the distribution, as follows:

[0050]

[0051] Where 1 indicates that the test data belongs to the sample within the distribution, 0 indicates that the test data belongs to the sample outside the distribution, and λ represents the set threshold.

[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0053] (1) This paper proposes a new few-sample out-of-distribution detection framework, which can simultaneously synthesize out-of-distribution information of image modality and text modality based on a small number of in-distribution samples. Without relying on additional out-of-distribution data collection, it fully taps the multimodal potential, thereby improving the out-of-distribution detection capability of visual language models.

[0054] (2) This paper designs two multimodal information synthesis methods for different scenarios, depending on whether clear out-of-distribution labels can be extracted from in-distribution samples: a synthesis method based on recognition information and a synthesis method based on image segmentation. This design has good adaptability and can automatically select the appropriate strategy based on the task scenario, making it applicable to various types of few-sample out-of-distribution detection tasks.

[0055] (3) The present invention simultaneously introduces the category prototype information of the image modality and the text modality in the inference stage, constructs a multimodal prototype matching mechanism, and proposes a new out-of-distribution discrimination score based on this mechanism, which helps to further improve the model's ability to recognize out-of-distribution samples and overcome the judgment bias caused by relying only on a single modality in existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0057] Figure 1 This is a flowchart of a method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0059] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0060] See also Figure 1 This embodiment provides a method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis, comprising the following steps:

[0061] S1. Select a multimodal information synthesis method based on different original in-distribution samples, and synthesize out-of-distribution samples. Combine the original in-distribution samples with the out-of-distribution samples. Each sample includes multiple categories of images and corresponding class labels to obtain an image set and a class label set.

[0062] First, in order to more effectively synthesize high-quality multimodal out-of-distribution data from a small number of in-distribution samples and improve the out-of-distribution detection capability of the visual language model, based on a small number of in-distribution image samples, the semantic features of the image content are analyzed and sampled to generate high-quality multimodal out-of-distribution samples in the form of images and texts that are closely related to the in-distribution semantics. This process does not rely on any external pre-collected out-of-distribution data, which improves the quality and adaptability of the generated samples. This embodiment designs two multimodal information synthesis methods: ROS, a synthesis method based on recognition information, and SOS, a synthesis method based on image segmentation. According to the observation of the background characteristics of the images of different in-distribution data sets themselves, the ROS, SOS, or a mode shared by ROS and SOS can be switched according to actual needs.

[0063] The following describes two multimodal information synthesis methods in detail.

[0064] 1. Identification-based out-of-distribution sample synthesis method.

[0065] This method is particularly suited for scenarios where in-distribution images possess clear contextual semantics. It primarily leverages a pretrained object recognition model to extract labels from in-distribution images, removing labels belonging to in-distribution categories to construct a set of out-of-distribution labels. Combining image cropping with the image-text similarity matching mechanism within the visual language model, it selects image patches from local regions of the training image that semantically match the out-of-distribution labels as out-of-distribution images. Ultimately, out-of-distribution images are paired with out-of-distribution labels, forming a complete multimodal out-of-distribution sample.

[0066] The specific steps include:

[0067] (1) Label Extraction: Using the pre-trained image annotation model RAM (Recognize Anything Model), we perform object recognition on each in-distribution image and obtain the set of all recognizable semantic labels in the image. For example, an in-distribution image of the “car” category might be labeled “car,” “street,” “building,” “tree,” etc. by RAM.

[0068] (2) Out-of-distribution label screening: remove the labels that overlap with the in-distribution label set from the above identified label set, and only retain the labels that do not appear in the out-of-distribution labels as candidate out-of-distribution semantic labels.

[0069] (3) Image cropping and text construction: Each in-distribution image is randomly cropped several times (e.g., 64 times) to generate multiple local image slices. At the same time, a text template is constructed based on the candidate out-of-distribution labels, such as "a photo of a building" or "a photo of a tree."

[0070] (4) Matching scoring and selection: All cropped images and their corresponding text templates are fed into a pre-trained visual language model (e.g., CLIP) to calculate the similarity score between the image and the text. For each candidate out-of-distribution label, the image local patch with the highest score is selected as its corresponding out-of-distribution image.

[0071] (5) Out-of-distribution sample selection: The out-of-distribution images and their text labels constructed above are used to form an out-of-distribution sample set, which is paired with the original in-distribution sample set to train the out-of-distribution detection capability of the visual language model. For each class of out-of-distribution samples generated, the similarity between them and the in-distribution class prototype is calculated, and the classes with the lowest similarity (such as the top 20%) are selected as the final out-of-distribution samples for training to maximize their diversity.

[0072] 2. Segmentation-based out-of-distribution sample synthesis method.

[0073] This method is particularly suitable for scenarios where the in-distribution image background is complex or semantically ambiguous. First, an unsupervised image segmentation model is used to extract the background region of the in-distribution image. Then, an object detection model is used to remove foreground regions related to the in-distribution category, thereby obtaining background information that is as independent of the in-distribution category as possible. Subsequently, a clustering algorithm is used to cluster the background image. By calculating image-text similarity and matching it with an existing vocabulary, the corresponding text label is matched to each category, forming a multimodal out-of-distribution sample for image-text matching.

[0074] The specific steps include:

[0075] (1) Image segmentation: Use the unsupervised pre-trained image segmentation model SAM (Segment Anything Model) to perform full image segmentation on each in-distribution image, obtain image segments of multiple background areas, and uniformly fill them into rectangles.

[0076] (2) In-distribution foreground exclusion: To avoid in-distribution information contamination, the original image and its category labels are fed into the object detection model GroundingDINO to obtain the bounding box of the foreground object. Subsequently, the overlap area between each segmented image and the foreground region is calculated. If the overlap ratio is greater than a set threshold (e.g., 15%), the image segment is excluded.

[0077] (3) Candidate background clustering: All remaining background image segments are grouped into a set and divided into several categories using the K-means++ clustering algorithm. The number of clusters is generally set to the same as the number of categories in the distribution. For example, if the distribution has 100 categories, the number of clusters is set to 100.

[0078] (4) Text label generation: Calculate the average image feature value for each cluster category and perform text matching with the noun or adjective candidate set in the WordNet lexicon. Select the word with the highest similarity as the text label for that category. If different clusters generate the same text label, they can be merged and classified.

[0079] (5) Out-of-distribution sample selection: The out-of-distribution images and their text labels constructed above are used to form an out-of-distribution sample set, which is paired with the original in-distribution sample set to train the out-of-distribution detection capability of the visual language model. For each class of out-of-distribution samples generated, the similarity between them and the in-distribution class prototype is calculated, and the classes with the lowest similarity (such as the top 20%) are selected as the final out-of-distribution samples for training to maximize their diversity.

[0080] After synthesizing the out-of-distribution samples, this embodiment combines them with the original small sample in-distribution dataset and uses them together as training data to train the visual language model. In this embodiment, the visual language model includes multiple image encoders and text encoders, and the image encoders can share parameters. The training process is detailed in step S2.

[0081] S2. Use the image encoder to extract the image prototype features of the image set, calculate the mean and obtain the image frozen prototype vectors of various samples respectively, correspond the image frozen prototype vectors of the same category with the learnable vectors of the same dimension, and obtain various image prototypes after combination; combine the class label and the corresponding learnable parameter prefix into a text prompt, use the text encoder to extract features of the text prompt, and obtain the text features of various samples.

[0082] Before training the visual language model, this embodiment also involves the step of initializing the image and text. First, the text prompt is constructed and the trainable parameters are obtained. A learnable text prompt (Prompt) is trained for each in-distribution and out-of-distribution category to achieve image-text matching. Each text prompt is composed of a learnable parameter prefix and a category name text, wherein the category name text part is not learnable, and the learnable parameter prefix part is initialized with a Gaussian distribution with a mean of 0 and a standard deviation of 0.02. The length of each prompt vector is set to 77, which is consistent with the standard input length of the CLIP text encoder. During the training process, the in-distribution and out-of-distribution prompts are input into the text encoder of the visual language model to obtain text prompt embedding, and then the training image is input into the image encoder of the visual language model to obtain image embedding. Second, frozen image prototype extraction and trainable parameter initialization. The image prototype features of each in-distribution and out-of-distribution category are pre-extracted by the visual language model, and the average is calculated to obtain the frozen prototype vector of each type of image. The learnable vector is denoted as The image frozen prototype vector is recorded as The initial values ​​of both are 0. The two are spliced ​​together to obtain the image prototype of category i, as shown in the following formula:

[0083] S3. Take various image prototypes and text features of various samples as common input, train the visual language model, and obtain the final detection results based on the out-of-distribution detection score.

[0084] In the reasoning process, a dual-modal prototype matching mechanism of image and text is adopted, and similarity calculation is performed simultaneously to improve the detection accuracy. On this basis, a new out-of-distribution discrimination scoring index is proposed - multimodal prototype matching score S MPM , which is used to comprehensively evaluate the similarity between the sample to be tested and the categories within the distribution, so as to determine whether it is an out-of-distribution sample.

[0085] Specifically, in the visual language model inference training process, it is divided into three parts: the first is the optimization of image prototypes, the second is the optimization of in-distribution text prompts, and the third is the optimization of out-of-distribution text prompts.

[0086] During inference, the training set data samples are input into the image encoder to obtain the image feature vector; the training set data samples include in-distribution samples and out-of-distribution samples generated based on the in-distribution samples; the total loss of the common input and the image feature vector is calculated until the parameters converge and the training is completed.

[0087] While retaining the original in-distribution task capabilities, the recognition ability of out-of-distribution samples is significantly improved. The total loss of the three parts of training includes the in-distribution text prompt loss Out-of-distribution text hint loss and image prototype loss The total loss is as follows:

[0088]

[0089] Among them, λ1 and λ2 are the weight coefficients of the loss function.

[0090] The in-distribution text prompt loss is as follows:

[0091]

[0092] Among them, M represents the number of categories of samples within the distribution, N1 represents the number of categories of synthesized out-of-distribution samples, represents the similarity between the training image embedding and the text category hint embedding within the m-th distribution, represents the similarity between the training image embedding and the n1th out-of-distribution text category hint embedding, s * represents the similarity between the training image embedding and the text hint embedding of the corresponding real category;

[0093] The out-of-distribution text hint loss is as follows:

[0094]

[0095] The image prototype loss is as follows:

[0096]

[0097] in, represents the similarity between the embedding of the training image and the embedding of the image prototype within the m-th distribution, represents the similarity between the training image embedding and the nth image prototype embedding, s′ * Indicates the similarity between the training image embedding and the image prototype embedding of the corresponding real category.

[0098] After completing the training of the visual language model, it is necessary to test whether the acquired out-of-distribution data is accurate. This embodiment uses multimodal matching scores to calculate the accuracy. The steps are as follows:

[0099] For each test image, we calculate the similarity between the image embedding obtained by the visual encoder and the text hint embedding and image prototype embedding of each category, and calculate the multimodal matching score S MPM As the final out-of-distribution detection score, the multimodal matching score S MPM As follows:

[0100]

[0101] Where M is the number of in-distribution categories and N is the number of synthetic out-of-distribution categories.i 、s j and s k is the similarity between the embedding of the test image and the embedding of the text prompt according to the multimodal matching score, s′i, s′j, s′k are the similarities between the embedding of the test image and the embedding of the image prototype;

[0102] According to the multimodal matching score S MPM The size of determines whether the test sample belongs to the sample within the distribution or the sample outside the distribution, as follows:

[0103]

[0104] Where 1 indicates that the test data belongs to the sample within the distribution, 0 indicates that the test data belongs to the sample outside the distribution, and λ represents the set threshold.

[0105] In order to better illustrate the technical effect of this embodiment, this embodiment adopts the CLIP model of ViT-B / 16 architecture, and is completed in practice using two NVIDIA RTX 3090 (24G) GPUs. The SGD optimizer is used with a momentum of 0.9, and 30 epochs of training are performed. The cosine learning rate scheduler is used, the learning rate gradually decays from 0.002 to 0, the weight decay is 0.0005, and the batch size is 32. In terms of data, the ImageNet-1k dataset is used as the in-distribution data, and the images of the four datasets iNaturalist, SUN, Places365, and Texture are used as out-of-distribution data. iNaturalist is a real-world dataset containing 8142 fine-grained species. SUN is a scene recognition dataset containing 397 categories. Places365 consists of images for scene recognition and contains 365 scene categories. Texture is a dataset of images describing textures containing 47 categories. The results show that this method significantly outperforms the current mainstream methods in multiple benchmark tests, and reaches the optimal or suboptimal level in evaluation indicators such as accuracy, AUROC, and FPR95.

[0106] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.

[0107] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0108] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis, characterized in that: include: Select a multimodal information synthesis method based on different original in-distribution samples, and synthesize out-of-distribution samples. Combine the original in-distribution samples with the out-of-distribution samples. Each sample includes multiple categories of images and corresponding class labels to obtain an image set and a class label set. An image encoder is used to extract image prototype features from an image set. After calculating the mean, frozen prototype vectors of images of various samples are obtained. Frozen prototype vectors of images of the same category are matched with learnable vectors of the same dimension, and after combining, various image prototypes are obtained. The class label and the corresponding learnable parameter prefix are combined into a text prompt. The text prompt is then subjected to feature extraction using a text encoder to obtain text features of various samples. Various image prototypes and text features of various samples are used as common input, and a visual language model is trained to obtain the final detection results based on the out-of-distribution detection score.

2. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 1, characterized in that: The multimodal information synthesis method includes an out-of-distribution sample synthesis method based on recognition and an out-of-distribution sample synthesis method based on segmentation; the out-of-distribution sample synthesis method based on recognition is suitable for scenarios where the in-distribution image has clear background semantic information, and the out-of-distribution sample synthesis method based on segmentation is suitable for scenarios where the in-distribution image has a complex background or ambiguous semantics.

3. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 2, characterized in that: The identification-based out-of-distribution sample synthesis method includes: Use the pre-trained image annotation model to perform object recognition on each in-distribution image and obtain the set of all recognizable semantic labels in the image; Remove the labels in the semantic label set that overlap with the in-distribution labels, and use the remaining semantic labels as candidate out-of-distribution labels; Perform several random cropping operations on each in-distribution image to generate multiple local image blocks, and construct text templates based on candidate out-of-distribution labels. The local image segmentation and the corresponding text template are input into the pre-trained visual language model, the similarity score between the image and the text is calculated, and the local image segmentation with the highest score is selected as the out-of-distribution image. The out-of-distribution images and their corresponding out-of-distribution labels constitute the out-of-distribution sample set.

4. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 2, characterized in that: The segmentation-based out-of-distribution sample synthesis method includes: Use an unsupervised pre-trained image segmentation model to perform full image segmentation on each in-distribution image, obtain image segments of multiple background regions, and unify the image segments; The original in-distribution image and its category label are input into the object detection model to obtain the bounding box of the foreground object. The overlapping area between each segmented image and the foreground area is calculated. If the overlap ratio is greater than a set threshold, the image segment is removed. The remaining image segments are divided into several categories using a clustering algorithm, and the number of cluster categories is set to be the same as the number of in-distribution categories. The average image feature value of each clustered image segment is calculated, and the image segments are matched with the noun or adjective candidate set in the WordNet lexicon. The word with the highest similarity is selected as the text label of the segment. If the image segments of different clustered segments generate the same text label, they are merged into the same segment. The image segments of each cluster category are regarded as out-of-distribution images, and the corresponding text labels are regarded as corresponding out-of-distribution labels. The similarity between the out-of-distribution images and the original in-distribution images is calculated, and the categories with the lowest similarity are selected as the final out-of-distribution samples.

5. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 1, characterized in that: Multiple image encoders are used, and parameters are shared between image encoders.

6. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 1, characterized in that: The frozen prototype vectors of images of the same category are matched with the learnable vectors of the same dimension, specifically: Let the learnable vector be The image frozen prototype vector is recorded as The initial values ​​of both are 0. The two are spliced ​​together to obtain the image prototype of category i, as shown in the following formula:

7. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 1, characterized in that: The class label and the corresponding learnable parameter prefix are combined into a text prompt, including: Set the class label text to the non-learnable part, set the learnable parameter prefix to the learnable part, set the vector length of the prompt text, and initialize the learnable part with a Gaussian distribution; In the above example, the in-distribution text prompts and out-of-distribution text prompts are input into the text encoder to obtain the text features of the in-distribution samples and the text features of the out-of-distribution samples.

8. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 1, characterized in that: The training of the visual language model comprises the following steps: Inputting training set data samples into an image encoder to obtain image feature vectors; the training set data samples include in-distribution samples and out-of-distribution samples generated based on the in-distribution samples; Calculate the total loss of the common input and image feature vector until the parameters converge and the training is completed; The total loss includes the in-distribution textual hint loss Out-of-distribution text hint loss and image prototype loss The total loss is as follows: Among them, λ1 and λ2 are the weight coefficients of the loss function.

9. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 8, characterized in that: include: The in-distribution text prompt loss is as follows: Among them, M represents the number of categories of samples within the distribution, N1 represents the number of categories of synthesized out-of-distribution samples, represents the similarity between the training image embedding and the text category hint embedding within the m-th distribution, represents the similarity between the training image embedding and the n1th out-of-distribution text category hint embedding, s * represents the similarity between the training image embedding and the text hint embedding of the corresponding real category; The out-of-distribution text hint loss is as follows: The image prototype loss is as follows: in, represents the similarity between the embedding of the training image and the embedding of the prototype image in the mth distribution, represents the similarity between the training image embedding and the nth image prototype embedding, s′ * Indicates the similarity between the training image embedding and the image prototype embedding of the corresponding real category.

10. The method for detecting out-of-distribution samples in a visual language model based on multimodal outlier synthesis according to claim 1, characterized in that: Obtaining the final test result based on the out-of-distribution test score includes: For each test image, we calculate the similarity between the image embedding obtained by the visual encoder and the text hint embedding and image prototype embedding of each category, and calculate the multimodal matching score S MPM As the final out-of-distribution detection score, the multimodal matching score S MPM As follows: Where M is the number of in-distribution categories and N is the number of synthetic out-of-distribution categories. i 、s j and s k is the similarity between the embedding of the test image and the embedding of the text prompt according to the multimodal matching score, s′i, s′j, s′k are the similarities between the embedding of the test image and the embedding of the image prototype; According to the multimodal matching score S MPM The size of determines whether the test sample belongs to the sample within the distribution or the sample outside the distribution, as follows: Where 1 indicates that the test data belongs to the sample within the distribution, 0 indicates that the test data belongs to the sample outside the distribution, and λ represents the set threshold.