Out-of-distribution detection method, server, storage medium and program product
By obtaining the description context and using pre-trained model to generate feature vectors, the problem of low accuracy of out-of-distribution detection in the prior art is solved, and a higher accuracy of out-of-distribution sample recognition is achieved.
Patent Information
- Application Number
- CN202410092701.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
The existing off-distribution detection scheme only uses category names as the characteristics of the category, which leads to the inaccurate distinction between in-distribution samples and out-distribution samples, with low accuracy.
By obtaining the description context of the samples to be detected and the target category, the text encoder and sample encoder of the pretrained model are used to code and generate feature vectors, and determine whether the sample is an out-of-distribution sample based on the similarity of the feature vector.
It improves the accuracy of off-distribution detection, can more accurately identify off-distribution samples, and provides data basis for algorithm/model optimization.
Smart Images

Figure CN120372227A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and in particular to an out-of-distribution detection method, a server, a storage medium, and a program product. Background Art
[0002] The purpose of out-of-distribution (OOD) sample detection is to detect whether a sample is an out-of-distribution sample that does not belong to any category of interest, and it is an important way to perform adaptive evaluation of the performance of algorithms / models in the adaptive learning link. Through out-of-distribution detection, it can be determined whether there are out-of-distribution samples in the output results of the algorithm / model, so as to obtain the situation of misdetection of the algorithm / model in this scenario, providing important data basis for further optimizing the algorithm / model.
[0003] The current out-of-distribution detection scheme only uses the category name as all the text information describing the category features, which limits the discriminability of the similarity between potential sample feature information and text features, and cannot accurately distinguish in-distribution samples and out-of-distribution samples, resulting in low accuracy of out-of-distribution detection. Summary of the Invention
[0004] This application provides an out-of-distribution detection method, a server, a storage medium, and a program product to solve the problem of low accuracy of out-of-distribution detection.
[0005] In a first aspect, this application provides an out-of-distribution detection method, including:
[0006] Obtain a sample to be detected, at least one given target category, and the description context of each target category, where the description context contains the sample feature information of the target category, and the description contexts of different target categories are different; input the category name and description context of each target category into the text encoder of the pre-trained model for encoding to obtain the first text feature vector of each target category, and input the sample to be detected into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the sample to be detected; determine whether the sample to be detected is an out-of-distribution sample according to the similarity between the feature vector of the sample to be detected and the first text feature vector of each target category.
[0007] In a second aspect, this application provides an out-of-distribution detection method, including:
[0008] The receiving - end device receives a detection request for the segmentation image of the target object in the target detection result, obtains at least one interested category of the target detection task and the description context of each interested category, where the description context includes sample feature information of the interested category, and the description contexts of different interested categories are different; inputs the category names and description contexts of each interested category into the text encoder of the pre - trained model for encoding to obtain the third text feature vector of each interested category, and inputs the segmentation image into the image encoder of the pre - trained model for encoding to obtain the feature vector of the segmentation image; determines whether the segmentation image is an out - of - distribution sample according to the similarity between the feature vector of the segmentation image and the third text feature vectors of each interested category to obtain an out - of - distribution detection result; and returns the out - of - distribution detection result to the end - side device.
[0009] In a third aspect, the present application provides a server, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the server is enabled to execute the method described in the first aspect or the second aspect.
[0010] In a fourth aspect, the present application provides a computer - readable storage medium, in which computer - executable instructions are stored, and when a processor executes the computer - executable instructions, the method described in the first aspect or the second aspect is implemented.
[0011] In a fifth aspect, the present application provides a computer program product, including a computer program, which when executed by a processor, implements the method described in the first aspect or the second aspect.
[0012] The out - of - distribution detection method, server, storage medium and program product provided by the present application, by pre - obtaining the description context of each target category, where the description context of the target category describes the sample feature information of the target category, and the first text feature vector encoded from the category name and description context of the target category also contains rich sample feature information of the target category, determines whether the sample to be detected is an out - of - distribution sample based on the similarity between the first text feature vector and the feature vector of the sample, realizes OOD detection, and compared with the OOD detection based only on the similarity between the text feature vector of the category name and the feature vector of the sample, can greatly improve the accuracy of OOD detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0014] Figure 1Schematic diagram of an example system architecture applicable to this application;
[0015] Figure 2 Flowchart of an out-of-distribution detection method provided by an exemplary embodiment of this application;
[0016] Figure 3 Flowchart of an out-of-distribution detection method based on perceptual context and fallacy context provided by an exemplary embodiment of this application;
[0017] Figure 4 Flowchart of a method for training the descriptive context of each target class provided by an exemplary embodiment of this application;
[0018] Figure 5 Flowchart of a method for training perceptual context and fallacy context provided by an exemplary embodiment of this application;
[0019] Figure 6 Framework diagram of perceptual context and fallacy context training provided by an exemplary embodiment of this application;
[0020] Figure 7 Flowchart of a method for out-of-distribution detection in an object detection scenario provided by an exemplary embodiment of this application;
[0021] Figure 8 Flowchart of an out-of-distribution detection method based on perceptual context and fallacy context in an object detection scenario provided by an exemplary embodiment of this application;
[0022] Figure 9 Schematic diagram of the structure of a server provided by an embodiment of this application.
[0023] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0024] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0025] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0026] First, the terms involved in this application are explained:
[0027] In-distribution (ID) classification: also known as in-distribution sample classification, based on several pre-given categories of interest, the sample classification is predicted to be one of several categories of interest.
[0028] Out-of-distribution (OOD) detection: also known as out-of-distribution sample detection, detects whether a sample is an out-of-distribution (OOD) sample, that is, whether the sample does not belong to any category of interest.
[0029] Contrastive Language-Image Pre-training (CLIP) model: It is a visual language model (Vision-Language Models, VLM) for image-text correlation matching. It is a multimodal pre-training model across text and image modalities.
[0030] Context: In this case, it refers to the parameters to be learned in the cue fine-tuning training.
[0031] Description context: In this solution, it specifically refers to a set of parameters determined by prompt fine-tuning training and used to describe the sample features of the category of interest. The description context of different categories is different. The description context of any category includes multiple learnable word vectors. The description context of each category can be determined through prompt fine-tuning training.
[0032] Prompt Tuning: It is a visual language model (VLM) training method.
[0033] Visual question answering task: Given an input image and a question, determine the answer to the question from the visual information of the input image.
[0034] Image description task: Generate description text for an input image.
[0035] Visual entailment task: predict the semantic relevance of an input image and text, i.e., entailment, neutral, or contradiction.
[0036] Referential Expression and Comprehension Task: Locate the image region corresponding to the input text in the input image according to the input text.
[0037] Image Generation Task: Generate an image based on the input descriptive text.
[0038] Text-based Sentiment Classification Task: Predict the sentiment classification information of the input text.
[0039] Text Summarization Task: Generate summary information of the input text.
[0040] Multi-modal Task: It refers to downstream tasks where the input and output data involve multiple modal data such as images and texts, such as Visual Question Answering Task, Image Captioning Task, Visual Entailment Task, Referential Expression and Comprehension Task, Image Generation Task, etc.
[0041] Multi-modal Pre-trained Model: It refers to a pre-trained model where the input and output data involve multiple modal data such as images and texts, and can be applied to multi-modal task processing after fine-tuning training.
[0042] Pre-trained Language Model: A pre-trained model obtained by pre-training a large-scale language model (abbreviated as LLM).
[0043] A large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, or even hundreds of billions of model parameters. A large model can also be called a Foundation Model (FM). Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with over hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large-scale language models (LLMs), multi-modal pre-trained models, etc.
[0044] When a large model is actually applied, only a small number of samples are needed to fine-tune the pre-trained model and then it can be applied to different tasks. Large models can be widely applied in fields such as Natural Language Processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Captioning (IC), image generation, etc., and tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0045] Adaptive learning refers to learning using new data in the operating environment and quickly adapting to changes in real-world situations that are unforeseen during the development process. Adaptive learning includes links such as self-evaluating model / algorithm performance, optimizing the model / algorithm, and scenario migration. The purpose of self-evaluation is to automatically perceive the change / decrease in model / algorithm performance when the specific application scenario changes, so as to periodically discover problems independently, optimize the model / algorithm independently, evaluate the effect independently, update the model / algorithm automatically, and then conduct long-term adaptive management of the model / algorithm. And one of the important ways to achieve self-evaluation is out-of-distribution (OOD) detection.
[0046] For example, taking the self-evaluation performance of an object detection algorithm as an example, out-of-distribution detection is performed on images in which the object detection algorithm detects objects of predefined classes of interest, and it is judged whether they are out-of-distribution samples that do not belong to the classes of interest, so as to evaluate whether the object detection algorithm outputs more false detection results and provide data basis for optimizing the object detection algorithm.
[0047] Aiming at the problem of low accuracy of existing out-of-distribution detection schemes, this application provides an out-of-distribution detection method. By obtaining a sample to be detected, a given plurality of target classes, and the description context of each target class, the description context of the target class describes the in-distribution sample feature information and boundary feature information of the target class; encoding the class name and description context of each target class into the text encoder of the pre-trained model to obtain the first text feature vector of each target class, and encoding the sample to be detected into the sample encoder of the pre-trained model to obtain the feature vector of the sample to be detected; determining whether the sample to be detected is an out-of-distribution sample according to the similarity between the feature vector of the sample to be detected and the first text feature vector of the target class, so as to achieve out-of-distribution (OOD) detection.
[0048] Among them, the sample to be detected can be data of various modalities such as images, voices, videos, texts, etc. For samples to be detected of different modalities, the pre-trained models used for OOD detection can be different.
[0049] For example, in the scenario where the sample to be detected is an image, the pre-trained model used for OOD detection of the image can be a vision-language model (VLM) that includes a text encoder and an image encoder, such as a contrastive language-vision pre-training (CLIP) model, or other multi-modal pre-trained models that cross text and images.
[0050] For example, in a scenario where the sample to be detected is speech, the pre-trained model used for out-of-distribution (OOD) detection of speech can be a multi-modal pre-trained model that includes a text encoder and a speech encoder, such as a multi-modal pre-trained model across text and speech, or a tri-modal pre-trained model across text, image, and speech.
[0051] In the method of this embodiment, by obtaining in advance the descriptive context of each target category, where the descriptive context of the target category describes the in-distribution sample feature information and boundary feature information of the target category, and encoding the category name and descriptive context of the target category into a first text feature vector, which richly contains the in-distribution sample feature information and boundary feature information of the target category, and determining whether the sample to be detected is an out-of-distribution sample based on the similarity between the first text feature vector and the feature vector of the sample, OOD detection is achieved. Compared with the detection based only on the similarity between the text feature vector of the category name and the feature vector of the sample, the accuracy of OOD detection can be greatly improved.
[0052] Figure 1 FIG. [X] is a schematic diagram of an example system architecture applicable to the present application. As Figure 1 shown, the system architecture includes a server and an edge device. Among them, there is a communicable communication link between the server and the edge device, which can realize the communication connection between the server and the edge device.
[0053] Among them, the edge device can be an electronic device running a downstream application, specifically a hardware device with network communication function, computing function, and information display function, including but not limited to smart phones, tablets, desktop computers, local servers, cloud servers, etc. When the edge device runs the downstream application, it generates a sample to be detected that needs to be subjected to OOD detection. Given multiple target categories of interest, the sample to be detected and the multiple target categories are transmitted to the server.
[0054] The server is an electronic device with computing power deployed in the cloud or locally, such as a cloud cluster, etc. The server stores a pre-trained model, and the pre-trained model is a multi-modal pre-trained model that includes a text encoder and a sample encoder. Among them, the sample encoder encodes the sample to be detected, and for samples of different modalities, different pre-trained models can be used. The server can train and obtain the descriptive context of each target category based on the pre-trained model and the given multiple target categories, and the descriptive context can accurately describe the in-distribution sample feature information and boundary feature information of the corresponding target category; and use the descriptive context of each target category to implement OOD detection of the given sample to be detected.
[0055] Specifically, the server receives the sample to be detected and multiple target categories transmitted by the end-side device, obtains the description context of each target category based on the multiple target categories, and the description context describes the in-distribution sample feature information and boundary feature information of the target category; inputs the category name and description context of each target category into the text encoder of the pre-trained model for encoding to obtain the text feature vector of each target category, and inputs the sample to be detected into the sample encoder of the pre-trained model for encoding to obtain the sample feature vector of the sample to be detected; determines whether the sample to be detected is an out-of-distribution sample according to the similarity between the sample feature vector of the sample to be detected and the text feature vector of the target category, and obtains the out-of-distribution detection result. Further, the server returns the out-of-distribution detection result to the end-side device.
[0056] In an example scenario, the end-side device can be a device running an object detection task. Specifically, based on the target categories of interest in the current task scenario, the original image is input into the object detection algorithm for object detection, and the positions and categories of the objects belonging to the target categories of interest that appear in the original image are output. Based on the positions of the target objects in the object detection result, the end-side device segments the image of the region where the target object is located from the original image to obtain the segmented image of the target object. In order to evaluate the performance of the object detection algorithm, the end-side device needs to obtain the out-of-distribution detection result of the segmented image of the target object to determine the misdetection situation of the object detection algorithm. Specifically, the end-side device sends a detection request for the segmented image of the target object in the object detection result to the server, and the detection request includes the segmented image of the target object and multiple categories of interest in the current object detection task.
[0057] The server receives the detection request for the segmented image of the target object in the object detection result from the end-side device, and obtains the segmented image of the target object and multiple categories of interest in the current object detection task. The server uses the segmented image of the target object as the sample to be detected, and uses multiple categories of interest in the current object detection task as the target categories to perform out-of-distribution sample detection. Specifically, the server obtains the description context of each category of interest (target category) obtained through training based on multiple categories of interest (target categories) in the object detection task, and the description context accurately describes the in-distribution sample feature information and boundary feature information of the corresponding category of interest. Further, the server inputs the category name and description context of each category of interest into the text encoder of the pre-trained model for encoding to obtain the third text feature vector of each category of interest, and inputs the segmented image into the image encoder of the pre-trained model for encoding to obtain the feature vector of the segmented image; determines whether the segmented image is an out-of-distribution sample according to the similarity between the feature vector of the segmented image and the third text feature vector of each category of interest, and obtains the out-of-distribution detection result. Further, the server returns the out-of-distribution detection result to the end-side device.
[0058] The edge device receives the out-of-distribution detection result of the segmented image of the target object returned by the server. Based on the out-of-distribution detection result of the segmented image of the target object, for the segmented image that is an out-of-distribution sample, it can be considered as a misdetection result of the target detection algorithm. The edge device can calculate the misdetection rate of the target detection algorithm based on the out-of-distribution detection results of the segmented images of all target objects determined by the target detection results of the target detection algorithm, and further diagnose / optimize the target detection algorithm according to the misdetection rate of the target detection algorithm to improve the accuracy of the target detection algorithm.
[0059] It should be noted that the out-of-distribution detection method in this embodiment can be applied not only to the aforementioned target detection scenario, but also to the performance evaluation of the image classification model in the image classification scenario, or to perform out-of-distribution detection in various scenarios of other voice and video data processing. This embodiment does not make specific limitations here.
[0060] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the drawings.
[0061] Figure 2 It is a flowchart of the out-of-distribution detection method provided by an exemplary embodiment of the present application. The execution subject of this embodiment is the server in the aforementioned system architecture. As Figure 2 shown, the specific steps of the method are as follows:
[0062] Step S201, obtain the sample to be detected, at least one given target category, and the description context of each target category. The description context includes the sample feature information of the target category, and the description contexts of different target categories are different.
[0063] Among them, the given target category is one or more categories that the user is interested in configured in a specific task scenario. The sample to be detected refers to the data that needs to perform out-of-distribution detection for the given target category. The sample to be detected can be various types of data such as images, voices, videos, and texts.
[0064] Exemplarily, in the target detection scenario, at least one given target category can be the category of interest configured in the target detection task, and the sample to be detected can be the segmentation image of the target object in the target detection result. By performing out-of-distribution detection on the segmentation image of the target object, it can be determined whether the segmentation image of the target object is an out-of-distribution sample that does not belong to any given category of interest. If the segmentation image of a certain target object is an out-of-distribution sample, it indicates that there is an error in the segmentation image of the target object and it is a result of misdetection, thereby evaluating whether the target detection algorithm gives more misdetection results.
[0065] In this embodiment, based on at least one given target category, the description context of each target category is obtained. The description context is obtained by training with training data with category annotations and can accurately describe the sample feature information of the corresponding target category. Specifically, the description context can be composed of multiple word vectors. Each word vector in the description context can be randomly initialized and then trained with training data with category annotations to obtain the description context of different categories.
[0066] Optionally, the server can use the description context of each target category as learnable parameters and use a pre-trained model to train and determine the description context of each target category.
[0067] Optionally, the server can pre-train and store the description context of each target category. In this step, the server obtains the stored description context of each target category.
[0068] Specifically, the server maintains a description context library. For different task scenarios, the server can train and determine the description context of the known categories of interest in each task scenario according to the known categories of interest and store them in the description context library. In this step, for the category name of any given target category, if it is determined that the description context of the target category has been stored in the description context library, the description context of the target category is directly obtained from the description context library. For target categories that do not exist in the description context library, the description context of each target category can be trained using a pre-trained model. Further, the trained description context of the target category is added to the description context library. When subsequent out-of-distribution detection involves the same target category, the description context in the description context library can be reused without repeated training.
[0069] It should be noted that the description context library stores one or more categories of interest in a task scenario, as well as the category names and description contexts of each category of interest determined through training. When constructing the description context library, the description contexts of the categories of interest for a single task scenario can be trained and determined respectively to avoid mutual interference between different task scenarios. The description contexts of categories of interest with the same name in different task scenarios can be different.
[0070] Step S202: Input the category names and description contexts of each target category into the text encoder of the pre-trained model for encoding to obtain the first text feature vectors of each target category.
[0071] In this embodiment, an out-of-distribution (OOD) detection is implemented using a multi-modal pre-trained model that includes a text encoder and a sample encoder. For samples to be detected in different modalities, different sample encoders are used, and different pre-trained models are used for OOD detection. The pre-trained model can be selected according to the modality of the samples to be detected in the current task scenario.
[0072] For example, in a scenario where the sample to be detected is an image, the pre-trained model used for OOD detection of the image can be a vision-language model (VLM) that includes a text encoder and an image encoder, such as a contrastive language-vision pre-training (CLIP) model, or other multi-modal pre-trained models that cross text and image.
[0073] For example, in a scenario where the sample to be detected is speech, the pre-trained model used for OOD detection of the speech can be a multi-modal pre-trained model that includes a text encoder and a speech encoder, such as a multi-modal pre-trained model that crosses text and speech, or a multi-modal pre-trained model that crosses text, image, and speech.
[0074] In this step, after concatenating the category names and description contexts of each target category, they are input into the text encoder of the pre-trained model for encoding to obtain the text feature vectors of each target category (denoted as the first text feature vectors). The first text feature vectors of the target categories contain rich sample feature information of the target categories.
[0075] Step S203: Input the sample to be detected into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the sample to be detected.
[0076] In this step, the sample to be detected is input into the sample encoder of the pre-trained model, and the feature vector of the sample to be detected is obtained through encoding by the sample encoder.
[0077] In this embodiment, steps S202 and S203 can be executed in parallel or in any order successively, and no specific limitation is imposed on the execution order between these two steps here.
[0078] Step S204: Determine whether the sample to be detected is an out-of-distribution sample according to the similarity between the feature vector of the sample to be detected and the first text feature vectors of each target category.
[0079] In this step, calculate the similarity between the feature vector of the sample to be detected and the first text feature vectors of each target category as the classification confidence of the sample to be detected belonging to each target category; perform normalization processing on the classification confidence of the sample to be detected belonging to each target category, so that after normalization, the classification confidence of the sample to be detected belonging to each target category takes values in the interval [0, 1] and the sum of the classification confidences is equal to 1. Further, take the negative of the maximum value among the normalized classification confidences of the sample to be detected belonging to each target category as the confidence that the sample to be detected is an out-of-distribution sample. If the confidence that the sample to be detected is an out-of-distribution sample is greater than or equal to the confidence threshold, determine that the sample to be detected is an out-of-distribution sample. If the confidence that the sample to be detected is an out-of-distribution sample is less than the confidence threshold, determine that the sample to be detected is not an out-of-distribution sample.
[0080] Exemplarily, use r k to represent the classification confidence that the sample to be detected belongs to the k-th target category. Then the confidence g(X) that the sample to be detected X is an out-of-distribution sample is: g(X) = -max{softmax(r k )}, where softmax is the normalization method and max represents taking the maximum value. Use r' to represent the confidence threshold. If g(X) > r', then the sample to be detected X is an out-of-distribution sample; if g(X) ≤ r', then the sample to be detected X is not an out-of-distribution sample. Among them, the confidence threshold r' can be determined according to the number of target categories in the actual task scenario and is negatively correlated with the number of target categories, such as 2 / the number of target categories, 0.5 / the number of target categories, etc. The specific value can be configured and adjusted according to the requirements and empirical values of the actual task scenario, and no specific limitation is made here.
[0081] In an alternative embodiment, this step can also calculate the confidence g(X) that the sample to be detected X is an out-of-distribution sample in the following way: where max represents taking the maximum value, r k represents the classification confidence that the sample to be detected X belongs to the k-th target category, r j represents the classification confidence that the sample to be detected X belongs to the j-th target category, and the value of j is an integer in the interval [1, C], where C represents the number of target categories. It should be noted that in this embodiment and subsequent embodiments, the similarity between any two feature vectors can be the cosine similarity of the two feature vectors. In addition, the similarity between two feature vectors can also be calculated by means of Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc. This embodiment does not make specific limitations in this regard.
[0082] In the method of this embodiment, by obtaining the description context of each target category in advance, the description context of the target category can accurately describe the sample feature information of the target category, and the description contexts of different target categories are different. The first text feature vector encoded by the category name and description context of the target category contains rich sample feature information of the target category. Based on the similarity between the first text feature vector of each target category and the feature vector of the sample, it is determined whether the sample to be detected is an out-of-distribution sample, realizing OOD detection. Compared with the detection based on the similarity between the text feature vector of the category name only and the feature vector of the sample, the accuracy of OOD detection can be greatly improved.
[0083] In an alternative embodiment, the description context of each target category includes a perceptual context and a spurious context. Among them, the perceptual context contains the feature information of the in-distribution samples of the target category, and the spurious context contains the feature information of the spurious samples outside the distribution of the target category (denoted as spurious OOD samples). The perceptual context and the spurious context are used to hierarchically describe the boundaries of each category, so as to accurately detect the out-of-distribution samples that fall outside the boundaries of the target category.
[0084] Figure 3 It is a flowchart of the out-of-distribution detection method based on the perceptual context and the spurious context provided by the embodiment of the present application. As Figure 3 shown, the specific steps of this method are as follows:
[0085] Step S301, obtain the sample to be detected, a given plurality of target categories, and the description context of each target category, where the description context includes a perceptual context and a spurious context.
[0086] In this embodiment, the implementation manner of obtaining the sample to be detected, a given plurality of target categories, and the description context of each target category is similar to that in the foregoing step S201. For specific reference, see the relevant description of the foregoing step S201. The difference is that in this embodiment, the description context of the target category specifically includes a perceptual context and a spurious context.
[0087] Among them, the perception context contains the feature information of in-distribution samples of the target category, and the fallacy context contains the feature information of fallacy samples (denoted as fallacy OOD samples) outside the distribution of the target category. The perception context can perceive the differences between different target categories in the current task scenario. The fallacy context models the fallacy samples in the out-of-distribution samples of each target category (samples that are similar to the in-distribution samples of the target category but do not belong to the target category). The perception context and fallacy context of the target category hierarchically construct an accurate feature description of the target category. Based on the perception context, the target category to which the sample belongs can be roughly predicted, and then based on the combination of the perception context and the fallacy context, it can be finely identified whether the sample is a true in-distribution sample or an OOD sample.
[0088] For example, for the object detection task of detecting cats and apples in an image, cats and apples are two different target categories. Through the perception context of the two target categories, the differences between the two categories of cats and apples can be perceived. For a black panther appearing in the image, which is similar to a cat but does not belong to a cat, the black panther can be regarded as a fallacy OOD sample of the target category of cats. For a peach appearing in the image, which is similar to an apple but does not belong to an apple, the peach can be regarded as a fallacy OOD sample of the target category of apples. Since the peach is not similar to the cat, although the peach is an out-of-distribution sample, it cannot be regarded as a fallacy OOD sample of the target category of cats. Since the black panther is not similar to the apple, although the black panther is an out-of-distribution sample, it cannot be regarded as a fallacy OOD sample of the target category of apples.
[0089] Since the description context includes these two types of context information, namely the perception context and the fallacy context, in the foregoing step S202, the category name and description context of each target category are input into the text encoder of the pre-trained model for encoding to obtain the first text feature vector of each target category, which can be specifically implemented through steps S302 and S303.
[0090] In this embodiment, OOD detection uses a pre-trained model. For specific details, refer to the relevant description in the foregoing step S202, which will not be elaborated here.
[0091] Step S302: Concatenate the category name of each target category with the perception context and then input it into the text encoder of the pre-trained model for encoding to obtain the first perception feature vector of each target category.
[0092] In this step, after concatenating the category name of each target category with the perception context and inputting it into the text encoder of the pre-trained model for encoding, the obtained text feature vector of each target category is denoted as the first perception feature vector. The first perception feature vector of the target category contains rich feature information of the in-distribution samples of the target category.
[0093] Step S303: Concatenate the category names of each target category with the fallacy context and input them into the text encoder for encoding to obtain the first fallacy feature vectors of each target category.
[0094] In this step, after concatenating the category names of each target category with the fallacy context and inputting them into the text encoder of the pre-trained model for encoding, the obtained text feature vectors of each target category are denoted as the first fallacy feature vectors. The first fallacy feature vectors of the target category contain rich feature information of the out-of-distribution (OOD) fallacy samples of the target category.
[0095] In this embodiment, steps S302 and S303 can be performed simultaneously (encoding separately using two text encoders or encoding together after concatenation), or executed sequentially in any order. Here, no specific limitation is imposed on the execution order between these two steps.
[0096] Step S304: Input the sample to be detected into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the sample to be detected.
[0097] In this step, input the sample to be detected into the sample encoder of the pre-trained model, and obtain the feature vector of the sample to be detected through encoding by the sample encoder.
[0098] Among them, the text encoding in steps S302 - S303 and the sample encoding in S304 can be performed in parallel or executed sequentially in any order. Here, no specific limitation is imposed on the execution order between these two steps.
[0099] Since the description context includes two types of context information, namely the perception context and the fallacy context, two first text feature vectors of each target category are obtained through steps S302 and S303, including the first perception feature vector and the first fallacy feature vector. In this embodiment, steps S305 and S306 can be used to implement the determination in step S204 of whether the sample to be detected is an out-of-distribution sample based on the similarity between the feature vector of the sample to be detected and the first text feature vectors of each target category.
[0100] Step S305: For any target category, calculate the first similarity between the feature vector of the sample to be detected and the first perception feature vector of the target category, and the second similarity between the feature vector of the sample to be detected and the first fallacy feature vector of the target category, and calculate the classification confidence of the sample to be detected belonging to the target category based on the first similarity and the second similarity.
[0101] Specifically, for any target category, calculate the first similarity between the feature vector of the sample to be detected and the first perceptual feature vector of the target category, and the second similarity between the feature vector of the sample to be detected and the first fallacy feature vector of the target category. Then, calculate the classification confidence of the sample to be detected belonging to the target category based on the first similarity and the second similarity, so as to obtain the classification confidence of the sample to be detected belonging to each target category.
[0102] Specifically, when calculating the classification confidence of the sample to be detected belonging to the target category according to the first similarity and the second similarity, specifically, the second similarity can be used as an inhibitory term for the first similarity, and the first confidence of the sample to be detected belonging to the target category can be calculated; calculate the product of the first similarity and the first confidence to obtain the classification confidence of the sample to be detected belonging to the target category.
[0103] Optionally, use r k to represent the classification confidence of the sample to be detected belonging to the k-th target category, and r k can be calculated and determined by the following formula (1):
[0104]
[0105] where x represents the feature vector of the sample to be detected X, represents the first perceptual feature vector of the k-th target category, represents the first fallacy feature vector of the k-th target category, e represents the natural constant, represents the first similarity between and x, represents the second similarity between and x, and
[0106] represents the first confidence of the sample to be detected belonging to the target category. k Optionally, the classification confidence r
[0107]
[0108] of the sample to be detected belonging to the k-th target category can also be calculated and determined by the following formula (2):
[0109] where the meanings of the symbols in formula (2) are the same as those in formula (1), and will not be elaborated here.In this embodiment, the first similarity represents the similarity between the sample X to be detected and the perceptual context of the k-th target category, reflecting the possibility that the sample X to be detected belongs to the k-th target category. The second similarity is used as an inhibitory term in the first confidence. Its function is to suppress the classification confidence that the sample X to be detected belongs to the k-th target category by the similarity between the sample X to be detected and the fallacy context when the similarity between the perceptual context and the out-of-distribution sample is too large. The multiplication of these two items integrates the classification confidence that the sample X to be detected belongs to the k-th target category.
[0110] In other alternative embodiments, the first similarity and the first confidence can also be weighted and summed as the classification confidence that the sample X to be detected belongs to the k-th target category. Among them, the weight coefficients of the first similarity and the first confidence can be configured and adjusted according to the needs of the actual application scenario, and no specific limitation is made here.
[0111] Step S306: Determine the confidence that the sample to be detected is an out-of-distribution sample according to the classification confidence that the sample to be detected belongs to each target category.
[0112] After obtaining the classification confidence that the sample to be detected belongs to each target category, in this step, the confidence that the sample to be detected is an out-of-distribution sample is determined according to the classification confidence that the sample to be detected belongs to each target category.
[0113] Specifically, the classification confidence that the sample to be detected belongs to each target category is normalized so that the classification confidence that the sample to be detected belongs to each target category after normalization takes values in the interval [0, 1] and the sum of the classification confidences is equal to 1. The negative value of the maximum value among the classification confidences that the sample to be detected belongs to each target category after normalization is used as the confidence that the sample to be detected is an out-of-distribution sample.
[0114] Exemplarily, the confidence g(X) that the sample X to be detected is an out-of-distribution sample can be determined in the following way: g(X) = -max{softmax(r k )}, where softmax is the normalization method and max represents taking the maximum value.
[0115] Step S307: Determine whether the sample to be detected is an out-of-distribution sample according to the confidence that the sample to be detected is an out-of-distribution sample.
[0116] In this step, it is judged whether the confidence that the sample to be detected is an out-of-distribution sample is greater than or equal to the confidence threshold. If the confidence that the sample to be detected is an out-of-distribution sample is greater than or equal to the confidence threshold, it is determined that the sample to be detected is an out-of-distribution sample. If the confidence that the sample to be detected is an out-of-distribution sample is less than the confidence threshold, it is determined that the sample to be detected is not an out-of-distribution sample.
[0117] Exemplarily, use rk Denote the classification confidence that the sample to be detected belongs to the k-th target category. Then the confidence g(X) that the sample to be detected X is an out-of-distribution sample is: g(X) = -max{softmax(r k )}, where softmax is a normalization method and max represents taking the maximum value. Denote the confidence threshold as r'. If g(X) > r', then the sample to be detected X is an out-of-distribution sample; if g(X) ≤ r', then the sample to be detected X is not an out-of-distribution sample. Among them, the confidence threshold r' can be determined according to the number of target categories in the actual task scenario, and is negatively correlated with the number of target categories, such as 2 / the number of target categories, 0.5 / the number of target categories, etc. The specific value can be configured and adjusted according to the requirements and empirical values of the actual task scenario, and no specific limitation is made here.
[0118] In this embodiment, the description context of each target category includes a perception context and a fallacy context. The perception context contains the feature information of the in-distribution samples of the target category, and the fallacy context contains the features of the fallacy samples (fallacy OOD samples) outside the distribution of the target category. The perception context can perceive the differences between different target categories in the current task scenario, and the fallacy context models the fallacy samples in the out-of-distribution samples of each target category. The perception context and fallacy context of the target category hierarchically construct an accurate feature description of the target category, so as to accurately detect the out-of-distribution samples that fall outside the boundary of the target category and improve the accuracy of OOD detection.
[0119] In another alternative embodiment, the description context of the target category may only include the perception context. The OOD detection process is as follows: Input the category name and perception context of each target category into the text encoder of the pre-trained model for encoding to obtain the first text feature vector (i.e., the first perception feature vector) of each target category. Input the sample to be detected into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the sample to be detected. Further, determine whether the sample to be detected is an out-of-distribution sample according to the similarity between the feature vector of the sample to be detected and the first text feature vectors of each target category. For the specific implementation method, refer to the relevant content of step S204, which will not be elaborated here.
[0120] Figure 4 This is a flowchart of a method for training the description context of each target category provided by an exemplary embodiment of the present application. On the basis of any of the foregoing embodiments, the training process of the description context of each target category is described in detail in this embodiment. As Figure 4 shown, the specific process of training the description context of each target category is as follows:
[0121] Step S401, obtain the training samples corresponding to each target category.
[0122] Among them, the training samples at least include in-distribution samples of each target category. The in-distribution samples of each target category can be obtained by collecting samples generated in the task scenario, or using samples labeled with the corresponding categories in existing labeled datasets.
[0123] Optionally, the training samples can also include out-of-distribution samples of each target category. The out-of-distribution samples corresponding to each target category are obtained by collecting samples in each task scenario or using large models to generate similar samples of in-distribution samples, and then manually screening out out-of-distribution samples with similar target categories.
[0124] Step S402: Use the description context of each target category as a parameter to be trained, and initialize the description context of each target category.
[0125] Among them, the description context of any target category includes multiple learnable word vectors, and these learnable word vectors are trained through prompt fine-tuning. During the training process, keep the parameters of the pre-trained model fixed and only train the description context of each target category.
[0126] In this step, when initializing the description context of each target category, each word vector in the description context of each target category can be randomly initialized.
[0127] Step S403: Input the category name and the current description context of each target category into the text encoder of the pre-trained model for encoding to obtain the second text feature vector of each target category.
[0128] In this embodiment, a multi-modal pre-trained model including a text encoder and a sample encoder is used to train the description context of each target category. The pre-trained model used here can be the same as the pre-trained model used for OOD detection. For the specific selection of the pre-trained model used, refer to the relevant description in step S202, which will not be elaborated here.
[0129] In this step, after concatenating the category name and the description context of each target category, input them into the text encoder of the pre-trained model for encoding to obtain the text feature vector of each target category (denoted as the second text feature vector). For the sake of distinction, here the second text feature vector is used to refer to the text feature vector of the target category generated during the training process.
[0130] Step S404: Input the training samples into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the training samples.
[0131] In this step, input the training samples into the sample encoder of the pre-trained model, and obtain the feature vector of the sample to be detected through the encoding of the sample encoder.
[0132] In this embodiment, steps S202 and S203 can be performed in parallel or in any order successively, and no specific limitation is imposed on the execution order between these two steps here.
[0133] Step S405: Train the description context of each target category according to the similarity between the second text feature vectors of each target category and the feature vectors of the training samples.
[0134] In this step, according to the similarity between the second text feature vectors of each target category and the feature vectors of the training samples, calculate the loss function value, and update the description context of each target category by means of backpropagation of the reverse gradient according to the loss function value.
[0135] In an alternative embodiment, the training samples include in-distribution samples, and the description context includes perceptual context. Specifically, the loss function value can be calculated in the following manner:
[0136] Calculate the first cross-entropy loss according to the similarity between the feature vectors of the in-distribution samples and the second text feature vectors of each target category; take the target category to which the training sample belongs as the positive sample, take the target categories that the training sample does not belong to as the negative samples, and calculate the first contrastive loss according to the similarity between the feature vectors of the training samples and the second text feature vectors of the positive samples (the target categories to which they belong), and the similarity between the feature vectors of the in-distribution samples and the second text feature vectors of the negative samples (the target categories that they do not belong to); sum the first cross-entropy loss and the first contrastive loss with weights to obtain the first loss function value, and train the description context of each target category according to the first loss function value.
[0137] In an alternative embodiment, the training samples include in-distribution samples and out-of-distribution samples, and the description context includes perceptual context. Specifically, the loss function value can be calculated in the following manner:
[0138] Calculate the first cross-entropy loss according to the similarity between the feature vectors of the in-distribution samples in the training samples and the second text feature vectors of each target category; for the second text feature vector of any target category, take the feature vectors of the in-distribution samples corresponding to this target category as the positive samples, take the feature vectors of the out-of-distribution samples corresponding to this target category as the negative samples, and calculate the second contrastive loss according to the similarity between the feature vectors of the in-distribution samples (positive samples) and the second text feature vectors of the target categories to which they belong, and the similarity between the feature vectors of the out-of-distribution samples (negative samples) and the second text feature vectors of the corresponding target categories; train the description context of each target category according to the first cross-entropy loss and the second contrastive loss.
[0139] It should be noted that in this embodiment, the training of the description context of each target category adopts a prompt tuning framework. The parameters of the pre-trained model are fixed, and the description context (serving as a prompt word) of the input text encoder is trained. The learned description context of each category can be reused. When the category of interest in the task scenario changes, for example, trains are no longer needed to be detected in the object detection scenario, or trucks need to be newly added for detection, etc., only the corresponding description context needs to be trained on the newly added category, and the description context of the currently required category can be used during the inference process, without retraining on all categories, which greatly reduces the training cost, can efficiently handle the task scenario with category scalability, reduces the training and deployment costs, and improves the training efficiency.
[0140] The method of this embodiment obtains the training samples corresponding to each target category, takes the description context of each target category as the parameter to be trained, initializes the description context of each target category, inputs the category name and the current description context of each target category into the text encoder of the pre-trained model for encoding to obtain the second text feature vector of each target category, and inputs the training samples into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the training samples; according to the similarity between the second text feature vector of each target category and the feature vector of the training samples, the description context of each target category is trained, and the parameters of the pre-trained model are fixed during the training process, making full use of the highly generalized representation ability accumulated in the large-scale corpus during the pre-training stage of the model, having good robustness to the offset of the scenario task. The trained description context of each target category can accurately describe the boundary feature information of each target category, thereby improving the accuracy of OOD detection. Moreover, in this embodiment, the prior knowledge of the pre-trained model is used to automatically generate the fallacy OOD samples of each target category as training samples, which can improve the training quality of the perception context and fallacy context of each target category, and further improve the accuracy of OOD detection.
[0141] In an optional embodiment, the description context includes a perception context and a fallacy context. The perception context describes the in-distribution sample feature information of the target category, and the fallacy context describes the boundary feature information of the target category. Figure 5 The flowchart of the method for training the perception context and fallacy context of each target category provided by an exemplary embodiment of the present application is as Figure 5 shown. The specific process of training the perception context and fallacy context of each target category is as follows:
[0142] Step S501, obtain the training samples corresponding to each target category.
[0143] In this embodiment, the training samples corresponding to each target category obtained include in-distribution samples and out-of-distribution samples (which can be spurious OOD samples). For the specific implementation of this step, refer to the aforementioned step S401.
[0144] To improve the training effect, for the out-of-distribution samples corresponding to each target category in the training data, spurious OOD samples that are relatively similar to each target category can be selected. For example, for the target category of "apple", the in-distribution samples can be apple images in various scenarios (such as images under different lighting conditions, perspectives, backgrounds, or different apples, etc.). Compared with selecting images of some objects that are not similar to apples (such as flowers, etc.) as out-of-distribution samples, selecting images of some objects that are relatively similar to apples (such as peaches, etc.) as spurious OOD samples for the "apple" category, the descriptive context of the "apple" category obtained through training can more accurately describe the "apple" category, which can improve the training quality and effect. The spurious OOD samples corresponding to each target category can be determined by manual screening or generated using a large model by configuring appropriate prompts.
[0145] Step S502: Use the descriptive context of each target category as a parameter to be trained, and initialize the descriptive context of each target category.
[0146] In this embodiment, the initialized descriptive context of each target category includes an initialized perceptual context and a spurious context. Among them, both the perceptual context and the spurious context are composed of multiple learnable word vectors. By prompting and fine-tuning to train these learnable word vectors, the perceptual context and the spurious context of each target category are obtained.
[0147] In this step, when initializing the perceptual context of each target category, each word vector in the perceptual context of each target category can be randomly initialized. Similarly, when initializing the spurious context of each target category, each word vector in the spurious context of each target category can be randomly initialized.
[0148] Since the parameter to be trained includes the perceptual context and the spurious context, the processing process of the aforementioned step S403 is implemented by steps S503 - 504 in this embodiment.
[0149] Step S503: Concatenate the category name of each target category with the current perceptual context and input it into a text encoder for encoding to obtain the second perceptual feature vector of each target category.
[0150] In this step, the class names of each target category are concatenated with the current perceptual context and then input into the text encoder of the pre-trained model for encoding. The obtained text feature vectors of each target category are denoted as the second perceptual feature vectors. The second perceptual feature vectors of the target categories contain rich feature information of the in-distribution samples of the target categories. In this embodiment, in order to distinguish from the first perceptual feature vectors during OOD detection, the text feature vectors generated based on the perceptual context during the training process are denoted as the second perceptual feature vectors.
[0151] Step S504: Concatenate the class names of each target category with the current fallacy context and then input them into the text encoder for encoding to obtain the second fallacy feature vectors of each target category.
[0152] In this step, the class names of each target category are concatenated with the current fallacy context and then input into the text encoder of the pre-trained model for encoding. The obtained text feature vectors of each target category are denoted as the second fallacy feature vectors. The second fallacy feature vectors of the target categories contain rich features of the out-of-distribution samples (fallacy OOD samples) of the target categories. In this embodiment, in order to distinguish from the first fallacy feature vectors during OOD detection, the text feature vectors generated based on the fallacy context during the training process are denoted as the second fallacy feature vectors.
[0153] Step S505: Input the training samples into the sample encoder of the pre-trained model for encoding to obtain the feature vectors of the training samples.
[0154] In this step, the training samples are input into the sample encoder of the pre-trained model, and the feature vectors of the samples to be detected are obtained through encoding by the sample encoder.
[0155] In this embodiment, the process of sample encoding in step S505 and the process of text encoding in steps S503 - S504 can be performed in parallel or executed in any order successively. Here, no specific limitation is imposed on the execution order between these two steps.
[0156] In this embodiment, since the description context includes two types of context information, namely the perceptual context and the fallacy context, two second text feature vectors (the second perceptual feature vectors and the second fallacy feature vectors) of each target category are obtained through steps S503 and S504. In the aforementioned step S405, the description context of each target category is trained according to the similarity between the second text feature vectors of each target category and the feature vectors of the training samples. In this embodiment, it can be specifically implemented through steps S506 - S508.
[0157] Step S506: Calculate the first loss based on the similarity between the feature vector of the in-distribution sample and the second perceptual feature vectors of each target category, the target category corresponding to the in-distribution sample, and the similarity between the feature vector of the in-distribution sample and the second fallacy feature vectors of each target category.
[0158] Among them, the first loss is used to maximize the similarity between the feature vector of the in-distribution sample and the second perceptual feature vector of the target category it belongs to, minimize the similarity between the feature vector of the in-distribution sample and the second perceptual feature vector of the target category it does not belong to, and minimize the similarity between the feature vector of the in-distribution sample and the second fallacy feature vectors of each target category.
[0159] Exemplarily, in this step, according to the target category corresponding to the in-distribution sample, the similarity between the feature vector of the in-distribution sample and the second perceptual feature vector of the target category it belongs to, the similarity between the feature vector of the in-distribution sample and the second perceptual feature vector of the target category it does not belong to, and the similarity between the feature vector of the in-distribution sample and the second fallacy feature vector of the target category it does not belong to can be calculated. Further, the cross-entropy loss is calculated using the following formula (3) to obtain the first loss:
[0160]
[0161] Among them, L1 represents the first loss, n represents the number of in-distribution samples in the training samples, C represents the number of target categories. x i represents the feature vector of the i-th in-distribution sample, y i represents the target category to which x i belongs, represents the second perceptual feature vector of the target category y i . represents the similarity between x i and the second perceptual feature vector of y i . represents the similarity between x i and the second perceptual feature vector of the target category k (k ≠ y i ). represents the similarity between x i and the second fallacy feature vector of the target category k (k ≠ y i ). e is the natural constant, and log is the logarithmic operation.
[0162] Optionally, the first loss can also be calculated using the following formula (4):
[0163]
[0164] The meanings of the symbols in formula (4) are the same as those in formula (3), and will not be elaborated here.
[0165] In this embodiment, the description context of the target category includes the perception context and the fallacy context. Among them, the perception context is mainly used to distinguish different target categories and implement the classification tasks of different target categories. An out-of-distribution category is added for each target category (which can be regarded as a category determined by out-of-distribution samples or fallacy OOD samples of the target category, and can also be called the fallacy OOD category). In order to obtain a more strict classification boundary for the target category, we expand the classification task of the original C target categories to 2C categories, and use the Cross-Entropy (CE) loss function to distinguish the in-distribution samples of the k-th category from the target categories and out-of-distribution categories that it does not belong to, so that the trained perception context and fallacy context can accurately and hierarchically describe the category boundary of the target category, thereby improving the accuracy of OOD detection.
[0166] Step S507: Calculate the second loss according to the similarity between the feature vectors of the in-distribution samples and the second perception feature vectors and second fallacy feature vectors of each target category, and the similarity between the feature vectors of the out-of-distribution samples and the second perception feature vectors and second fallacy feature vectors of the target categories corresponding to the out-of-distribution samples.
[0167] Specifically, calculate the third loss according to the third similarity between the feature vector of the in-distribution sample and the second perception feature vector of the target category to which it belongs, and the fourth similarity between the feature vector of the in-distribution sample and the second fallacy feature vector of the target category to which it belongs. The third loss is used to maximize the third similarity and minimize the fourth similarity.
[0168] Exemplarily, the third loss can be calculated using the following formula (5):
[0169]
[0170] Among them, L3 represents the third loss, and n represents the number of in-distribution samples in the training samples. x i represents the feature vector of the i-th in-distribution sample, y i represents the target category to which x i belongs, represents the second perception feature vector of the target category y i . represents the similarity between x i and the second perception feature vector of y i . represents the similarity between x i and the second fallacy feature vector of y i . e is the natural constant, and log is the logarithmic operation.
[0171] Optionally, the third loss can also be calculated using the following formula (6):
[0172]
[0173] The meanings represented by each symbol in formula (6) are the same as those in formula (5), and will not be elaborated here.
[0174] Further, according to the fifth similarity between the feature vector of the out-of-distribution sample and the second spurious feature vector of the corresponding target category, and the sixth similarity between the feature vector of the out-of-distribution sample and the second perceptual feature vector of the corresponding target category, the fourth loss is calculated. The fourth loss is used to maximize the fifth similarity and minimize the sixth similarity.
[0175] Exemplarily, the fourth loss can be calculated using the following formula (7):
[0176]
[0177] where L4 represents the fourth loss, represents the number of out-of-distribution samples in the training samples. represents the feature vector of the i-th out-of-distribution sample, represents the corresponding target category, represents the target category 's second perceptual feature vector. represents and 's similarity between the second perceptual feature vectors. represents the target category 's second spurious feature vector. represents and 's similarity between the second spurious feature vectors. e is the natural constant, and log is the logarithmic operation.
[0178] Optionally, the fourth loss can also be calculated using the following formula (8):
[0179]
[0180] The meanings represented by each symbol in formula (8) are the same as those in formula (7), and will not be elaborated here.
[0181] Further, the sum of the third loss and the fourth loss is used as the second loss.
[0182] Exemplarily, the second loss is obtained by calculating with the following formula (9):
[0183]
[0184] Among them, L2 represents the second loss. For the meanings represented by other symbols, refer to the relevant descriptions in Formulas (5) and (7), which will not be elaborated here.
[0185] Optionally, the second loss can also be obtained by weighted summing the third loss and the fourth loss, where the weight coefficients of the third loss and the fourth loss can be configured and adjusted according to the requirements of the actual task scenario and empirical values, and no specific limitation is made here.
[0186] Since it is more difficult to distinguish the target category and the corresponding out-of-distribution category (fallacy OOD category), in the training process, the dual binary cross-entropy (Binary Cross-Entropy, abbreviated as BCE), that is, the logarithmic loss, is adopted in this step, which makes the in-distribution samples move away and makes the generated out-of-distribution samples (fallacy OOD samples) move away from the target category, so as to achieve a better effect of distinguishing the target category and the out-of-distribution category (fallacy OOD category), enabling the trained perceptual context and fallacy context to accurately and hierarchically describe the category boundaries of the target category, thereby improving the accuracy of OOD detection.
[0187] Steps S506 and S507 can be executed in parallel or in any order, and no specific limitation is made here.
[0188] Step S508, train the perceptual context and fallacy context of each target category according to the first loss and the second loss.
[0189] In this step, according to the first loss and the second loss, calculate the comprehensive loss, and update the perceptual context and fallacy context of each target category by means of backpropagation of the reverse gradient according to the comprehensive loss.
[0190] Specifically, the sum of the first loss and the second loss can be used as the comprehensive loss; alternatively, the comprehensive loss can be obtained by weighted summing the first loss and the second loss, where the weight coefficients of the first loss and the second loss can be configured and adjusted according to the requirements of the actual task scenario and empirical values, and no specific limitation is made here.
[0191] In this embodiment, the description context of each target category includes the perceptual context and the fallacy context. By obtaining the in-distribution samples and out-of-distribution samples (fallacy OOD samples) corresponding to each target category, taking the perceptual context and fallacy context of each target category as the parameters to be trained, fixing the parameters of the pre-trained model, and only training the perceptual context and fallacy context of each target category, the highly generalized representation ability accumulated in the large-scale corpus in the pre-training stage of the model is fully utilized, which has good robustness to the offset of the scenario task. The trained perceptual context and fallacy context of each target category can accurately and hierarchically describe the boundary feature information of each target category, thereby improving the accuracy of OOD detection.
[0192] Based on the foregoing Figure 4 or Figure 5 On the basis of the corresponding embodiments, in an alternative embodiment, the server first obtains the in-distribution sample sets of each target category, and then automatically generates the spurious OOD samples of each target category in the following manner: For any target category, determine the clustering center of the in-distribution sample set of the target category, and sample multiple feature vectors in the feature vector space determined by the in-distribution sample set of the target category as the feature vectors of multiple out-of-distribution samples corresponding to the target category. Among them, the distances between the sampled multiple feature vectors and the clustering center are greater than or equal to a distance threshold.
[0193] Specifically, for any target category, encode the in-distribution samples of the target category into feature vectors, and calculate and determine the central vector of the feature vector space determined by these feature vectors as the clustering center of the in-distribution sample set of the target category. Sample feature vectors in this feature vector space whose distances from the clustering center are greater than or equal to the distance threshold as the spurious OOD samples of the target category. The spurious OOD samples obtained in this way are sampling points far from the clustering center in this feature vector space. Among them, the distance threshold can be configured and adjusted according to the requirements and empirical values of the actual application scenario, and no specific limitation is made here. The out-of-distribution samples of each target category can be automatically generated in the same way.
[0194] Optionally, after sampling the feature vectors of multiple out-of-distribution samples (which can be spurious OOD samples) corresponding to the target category, the sampled out-of-distribution samples can also be filtered to ensure that the remaining out-of-distribution samples after filtering are indeed out-of-distribution samples rather than in-distribution samples belonging to the target category, thereby ensuring the training effect. Specifically, noise data (such as randomly masking one or more word vectors in the perceptual context) can be added to the current perceptual context of the target category as the perceptual context of the out-of-distribution category of the target category to simulate the perceptual context of an unknown out-of-distribution category different from the target category. Further, calculate the seventh similarity between the feature vectors of the multiple out-of-distribution samples corresponding to the target category obtained by sampling and the feature vectors of the perceptual context of the out-of-distribution category of the target category, and filter out the feature vectors of the out-of-distribution samples whose seventh similarity is less than or equal to the similarity threshold.
[0195] Among them, the seventh similarity reflects the similarity between the out-of-distribution samples sampled and the out-of-distribution categories simulated. The larger the seventh similarity, the closer the out-of-distribution samples are to the simulated out-of-distribution categories, and thus the farther away from the target category; while the smaller the seventh similarity, the farther the out-of-distribution samples are from the simulated out-of-distribution categories and the closer to the target category. By filtering out the out-of-distribution samples with a smaller seventh similarity, the remaining out-of-distribution samples can be ensured to be truly out-of-distribution samples, which can avoid the incorrect use of in-distribution samples as out-of-distribution samples in the training data, improve the quality of out-of-distribution samples in the training data, and thus improve the training quality and effect. Among them, the similarity threshold can be configured and adjusted according to the needs of the actual application scenario and empirical values, and no specific limitation is made here.
[0196] Figure 6 It is a framework diagram for training perceptual context and fallacy context provided by an exemplary embodiment of the present application. Figure 6 Among them, CLS_1…CLS_k respectively represent the category names of the 1-kth target categories. It represents the m word vectors included in the perceptual context of the first target category. It represents the m word vectors included in the perceptual context of the kth target category. It represents the m word vectors included in the fallacy context of the first target category. It represents the m word vectors included in the fallacy context of the kth target category. u represents the masked word vector. It represents the perceptual context of the out-of-distribution category of the first target category simulated by masking the second word vector in the perceptual context of the first target category. It represents the perceptual context of the out-of-distribution category of the kth target category simulated by masking the first word vector in the perceptual context of the kth target category. Among them, m represents the number of word vectors included in the perceptual context, which can be specifically configured and adjusted according to the needs of the actual task scenario. For example, m = 16, and no specific limitation is made here.
[0197] Such as Figure 6As shown in the figure, the training samples are input into the sample encoder of the pre-trained model for encoding to obtain the feature vectors of the samples. The class names of each target class are concatenated with the perceptual context and then input into the text encoder of the pre-trained model for encoding to obtain the second perceptual feature vectors of the target classes; the class names of each target class are concatenated with the fallacy context and then input into the text encoder of the pre-trained model for encoding to obtain the second fallacy feature vectors of the target classes. The second perceptual feature vectors of different target classes can determine the classification boundaries between different target classes. And the second perceptual feature vectors and the second fallacy feature vectors of the same target class can determine the class boundary of that target class. Through the perceptual context and the fallacy context, the boundaries of each target class can be accurately described hierarchically, realizing more accurate OOD detection.
[0198] It should be noted that the dashed arrow represents the perceptual context of the out-of-distribution classes of the target class simulated by adding noise data, which is only used in the training process of the perceptual context and the fallacy context of the target class and will not be used during OOD detection. The perceptual context of the out-of-distribution classes of the target class is determined by adding perturbations (noise) to the perceptual context of the target class, which is not a learnable parameter and does not require the loss function value to be updated through backpropagation of the reverse gradient.
[0199] Next, the OOD detection method will be exemplarily described in combination with a specific task scenario. In this embodiment, the target detection scenario is taken as an example. The edge device uses the target detection algorithm to perform target detection on the original image based on the given set of classes of interest, and the target detection result includes the positions and classes of one or more target objects belonging to the classes of interest that appear in the original image. Based on the positions of the target objects in the target detection result, the edge device can segment the image of the region where the target objects are located from the original image to obtain the segmented images of the target objects. By performing OOD detection on these segmented images of the target objects, it is determined whether the target objects are OOD samples, thereby determining the misdetection situation of the target detection algorithm and realizing the automatic evaluation of the performance of the target detection algorithm.
[0200] Figure 7 This is a flowchart of the out-of-distribution detection method in the target detection scenario provided by an exemplary embodiment of the present application. The execution subject of this embodiment is the server for performing OOD detection. As Figure 7 shown, the process of out-of-distribution detection in the target detection scenario is as follows:
[0201] Step S701: Receive the detection request of the edge device for the segmented images of the target objects in the target detection result, and obtain at least one class of interest of the target detection task and the description context of each class of interest. The description context contains the sample feature information of the class of interest, and the description contexts of different classes of interest are different.
[0202] In this embodiment, after the edge device obtains the segmentation image of the target object, it sends a detection request for the segmentation image of the target object in the target detection result to the server. The detection request includes the segmentation image to be detected and the categories of interest in the current target detection task.
[0203] After receiving the detection request for the segmentation image of the target object in the target detection result, the server can extract the segmentation image to be detected and the categories of interest in the current target detection task from the detection request.
[0204] Based on the categories of interest in the current target detection task, the server obtains the description context of each category of interest. The specific implementation method is similar to that of obtaining the description context of each target category in step S201 and step S301. Just take the category of interest as the target category and use the method of step S201 or step S301 to obtain the description context of the target category. For specific details, refer to the relevant content of the foregoing embodiments and will not be elaborated here.
[0205] Step S702: Input the category names and description contexts of each category of interest into the text encoder of the pre-trained model for encoding to obtain the third text feature vectors of each category of interest.
[0206] In this embodiment, the pre-trained model used includes a text encoder and an image encoder (as a sample encoder). Any existing pre-trained vision-language model (VLM) can be used as the pre-trained model, such as the Contrastive Language-Image Pretraining (CLIP) model, or other cross-text and image multi-modal pre-trained models. There is no specific limitation here.
[0207] In this step, after concatenating the category names and description contexts of each category of interest, input them into the text encoder of the pre-trained model for encoding to obtain the text feature vectors of each category of interest (denoted as the third text feature vectors). The third text feature vectors of the categories of interest contain rich sample feature information of the categories of interest.
[0208] Step S703: Input the segmentation image into the image encoder of the pre-trained model for encoding to obtain the feature vector of the segmentation image.
[0209] In this step, input the segmentation image into the image encoder of the pre-trained model, and obtain the feature vector of the segmentation image through the encoding of the image encoder.
[0210] Step S704: Determine whether the segmentation image is an out-of-distribution sample according to the similarity between the feature vector of the segmentation image and the third text feature vectors of each category of interest, and obtain the out-of-distribution detection result.
[0211] In this embodiment, the implementation manners of steps S702 - S704 are similar to those of the foregoing steps S202 - S204. The category of interest is used as the target category, and the segmented image is used as the sample to be detected. For the specific implementation process, refer to the relevant content of the foregoing embodiment, which will not be elaborated here.
[0212] Step S705: Return the out - of - distribution detection result to the edge device.
[0213] After obtaining the out - of - distribution detection result of the segmented image of the target object, the server returns the out - of - distribution detection result to the edge device. Further, the edge device can calculate the false detection rate of the target detection result according to the number of segmented images and the number of out - of - distribution samples included in the segmented images, and use this as a performance index of the target detection algorithm, as reference data for further diagnosing or optimizing the target detection algorithm.
[0214] In an optional embodiment, after obtaining the out - of - distribution detection result of the segmented image of the target object, the server can also calculate the false detection rate of the target detection result according to the number of segmented images and the number of out - of - distribution samples included in the segmented images. Further, the server returns the false detection rate of the target detection result to the edge device. The edge device can optimize the target detection algorithm used for target detection according to the false detection rate of the target detection result returned by the server.
[0215] The method of this embodiment is applied to the target detection scenario. By pre - obtaining the description context of each category of interest in the current target detection task, the description context of the category of interest can accurately describe the sample feature information of the category of interest, and the description contexts of different categories of interest are different. The third text feature vector encoded by the category name and description context of the category of interest contains rich sample feature information of the category of interest. Based on the similarity between the third text feature vectors of each category of interest and the feature vector of the segmented image, it is determined whether the image to be segmented is an out - of - distribution sample, realizing the OOD detection of the segmented sample, which can greatly improve the accuracy of OOD detection.
[0216] In an optional embodiment, the description context of each category of interest includes a perception context and a fallacy context. Among them, the perception context contains the feature information of in - distribution samples of the category of interest, and the fallacy context contains the feature information of fallacy samples outside the distribution of the category of interest (denoted as fallacy OOD samples). The perception context and the fallacy context are used to hierarchically describe the boundaries of each category, so as to accurately detect out - of - distribution samples that fall outside the boundaries of the category of interest.
[0217] Figure 8This is a flowchart of an out-of-distribution detection method based on perceptual context and fallacy context in the target detection scenario provided by the embodiments of the present application. As Figure 8 shown, the specific steps of this method are as follows:
[0218] Step S800: The edge device performs target detection on the original image to obtain the positions and categories of the target objects of the interested categories that appear in the original image.
[0219] Step S801: The edge device segments the segmentation image of the target object from the original image.
[0220] Step S802: The edge device sends a detection request for the segmentation image of the target object in the target detection result to the server.
[0221] Step S803: The server receives the detection request for the segmentation image of the target object in the target detection result and obtains at least one interested category of the target detection task.
[0222] Step S804: The server obtains the perceptual context and fallacy context of each interested category.
[0223] For obtaining the perceptual context and fallacy context of each interested category in this step, refer to the specific content of obtaining the perceptual context and fallacy context of each target category in the foregoing embodiments, which will not be elaborated here.
[0224] Step S805: The server splices the category name of each interested category with the perceptual context and inputs it into the text encoder of the pre-trained model for encoding to obtain the first perceptual feature vector of each interested category; splices the category name of each interested category with the fallacy context and inputs it into the text encoder for encoding to obtain the first fallacy feature vector of each interested category; inputs the segmentation image into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the segmentation image.
[0225] The specific implementation of this step is similar to the implementation manners of steps S302 - S304 in the foregoing embodiments. Regarding the interested category as the target category and the segmentation image as the sample to be detected, for the specific implementation manner, refer to the relevant content of the foregoing embodiments, which will not be elaborated here.
[0226] Step S806: For any interested category, the server calculates the first similarity between the feature vector of the segmentation image and the first perceptual feature vector of the interested category, and the second similarity between the feature vector of the segmentation image and the first fallacy feature vector of the interested category, and calculates the classification confidence that the segmentation image belongs to the interested category according to the first similarity and the second similarity.
[0227] In this step, the classification confidence of the segmented image belonging to the category of interest is calculated, which is similar to the implementation method of calculating the classification confidence of the sample to be detected belonging to the target category in the previous step S305. The category of interest is regarded as the target category, and the segmented image is regarded as the sample to be detected. For the specific implementation method, refer to the relevant content of the previous embodiment and will not be elaborated here.
[0228] Step S807: The server determines the confidence that the segmented image is an out-of-distribution sample based on the classification confidence of the segmented image belonging to each category of interest.
[0229] This step is similar to the implementation method in the previous step S306. The category of interest is regarded as the target category, and the segmented image is regarded as the sample to be detected. For the specific implementation method, refer to the relevant content of the previous embodiment and will not be elaborated here.
[0230] Step S808: The server determines whether the segmented image is an out-of-distribution sample based on the confidence that the segmented image is an out-of-distribution sample, and obtains the OOD detection result of the segmented image.
[0231] This step is similar to the implementation method in the previous step S307. The category of interest is regarded as the target category, and the segmented image is regarded as the sample to be detected. For the specific implementation method, refer to the relevant content of the previous embodiment and will not be elaborated here.
[0232] Step S809: The server returns the OOD detection result of the segmented image to the edge device.
[0233] Step S810: The edge device calculates the false detection rate of the segmented image according to the OOD detection result of the segmented image, and optimizes the object detection algorithm used for object detection.
[0234] This embodiment provides a complete example of training the perceptual context and fallacy context of the category of interest and implementing OOD detection based on the perceptual context and fallacy context of the category of interest when applied to the object detection scenario. For the specific implementation methods of each step in this embodiment, refer to the relevant descriptions of the previous embodiments and will not be elaborated here.
[0235] Figure 9 It is a schematic structural diagram of a server provided by an embodiment of the present application. As Figure 9 shown, the server includes: a memory 901 and a processor 902. The memory 901 is used to store computer execution instructions and can be configured to store various other data to support operations on the server. The processor 902 is communicatively connected to the memory 901 and is used to execute the computer execution instructions stored in the memory 901 to implement the technical solutions provided by any of the above method embodiments. Its specific functions and achievable technical effects are similar and will not be elaborated here.
[0236] Optionally, asFigure 9 As shown, the server further includes other components such as a firewall 903, a load balancer 904, a communication component 905, and a power supply component 906. Figure 9 Only some components are schematically shown in [the figure], and it does not mean that the server only includes Figure 9 the components shown.
[0237] The embodiments of the present application further provide a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the methods of any of the foregoing embodiments are implemented. The specific functions and achievable technical effects are not described herein again.
[0238] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the methods of any of the foregoing embodiments are implemented. The computer program is stored in a readable storage medium. At least one processor of the server can read the computer program from the readable storage medium, and at least one processor executing the computer program causes the server to execute the technical solutions provided in any of the foregoing method embodiments. The specific functions and achievable technical effects are not described herein again.
[0239] The embodiments of the present application provide a chip, including a processing module and a communication interface. The processing module can execute the technical solutions of the server in the foregoing method embodiments. Optionally, the chip further includes a storage module (such as a memory). The storage module is used to store instructions, and the processing module is used to execute the instructions stored in the storage module, and the execution of the instructions stored in the storage module causes the processing module to execute the technical solutions provided in any of the foregoing method embodiments.
[0240] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods of the embodiments of the present application.
[0241] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed and completed by a hardware processor, or by a combination of hardware and software modules in the processor. The memory may include high-speed random access memory (RAM), and may also include non-volatile storage, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0242] The above-mentioned memory may be an object storage (Object Storage Service, abbreviated as OSS). The above-mentioned memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Read Only Memory, abbreviated as PROM), read-only memory (Read Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk. The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile hotspot (WiFi), a second-generation mobile communication system (2G), a third-generation mobile communication system (3G), a fourth-generation mobile communication system (4G) / Long Term Evolution (abbreviated as LTE), a fifth-generation mobile communication system (5G), etc. mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (Near Field Communication, abbreviated as NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (Radio Frequency Identification, abbreviated as RFID) technology, infrared technology, ultra-wideband (Ultra Wide Band, abbreviated as UWB) technology, Bluetooth technology and other technologies. The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located. The above-mentioned storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or dedicated computer.
[0243] An exemplary storage medium is coupled to a processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit. Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master device.
[0244] It should be noted that in this document, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including such element.
[0245] The order of the above embodiments of the present application is for description only and does not represent the superiority or inferiority of the embodiments. Additionally, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. They are merely used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. Additionally, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types. The meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0246] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0247] Other embodiments of the present application will be readily contemplated by those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0248] The above are only the preferred embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of the present application.
Claims
1. An out-of-distribution detection method, characterized in that, Including: Obtaining a sample to be detected, at least one given target category, and the description context of each of the target categories, where the description context contains sample feature information of the target category, and the description contexts of different target categories are different; Inputting the category names and description contexts of each of the target categories into the text encoder of the pre-trained model for encoding to obtain the first text feature vectors of each of the target categories, and inputting the sample to be detected into the sample encoder of the pre-trained model for encoding to obtain the feature vector of the sample to be detected; Determining whether the sample to be detected is an out-of-distribution sample according to the similarity between the feature vector of the sample to be detected and the first text feature vectors of each of the target categories.
2. The method according to claim 1, wherein The description context includes a perception context and a fallacy context, the perception context contains the feature information of the in-distribution samples of the target category, and the fallacy context contains the feature information of the fallacy samples outside the distribution of the target category. The step of inputting the category names and description contexts of each of the target categories into the text encoder of the pre-trained model for encoding to obtain the first text feature vectors of each of the target categories includes: Inputting the concatenation of the category name of each of the target categories and the perception context into the text encoder of the pre-trained model for encoding to obtain the first perception feature vectors of each of the target categories; Inputting the concatenation of the category name of each of the target categories and the fallacy context into the text encoder for encoding to obtain the first fallacy feature vectors of each of the target categories.
3. The method according to claim 2, characterized in that The step of determining whether the sample to be detected is an out-of-distribution sample according to the similarity between the feature vector of the sample to be detected and the first text feature vectors of each of the target categories includes: Determining the confidence that the sample to be detected is an out-of-distribution sample according to the first similarity between the feature vector of the sample to be detected and the first perception feature vector of the target category, and the second similarity between the feature vector of the sample to be detected and the first fallacy feature vector of the target category; If the confidence that the sample to be detected is an out-of-distribution sample is less than or equal to the confidence threshold, determining that the sample to be detected is an out-of-distribution sample.
4. The method according to claim 3, characterized in that, The step of determining the confidence that the sample to be detected is an out-of-distribution sample according to the first similarity between the feature vector of the sample to be detected and the first perception feature vector of the target category, and the second similarity between the feature vector of the sample to be detected and the first fallacy feature vector of the target category includes: For any one of the target categories, calculating the first similarity between the feature vector of the sample to be detected and the first perception feature vector of the target category, and the second similarity between the feature vector of the sample to be detected and the first fallacy feature vector of the target category, and calculating the classification confidence that the sample to be detected belongs to the target category according to the first similarity and the second similarity; Determining the confidence that the sample to be detected is an out-of-distribution sample according to the classification confidence that the sample to be detected belongs to each of the target categories.
5. The method according to claim 4, characterized in that Calculating the classification confidence that the sample to be detected belongs to the target category according to the first similarity and the second similarity includes: According to the first similarity and the second similarity, taking the second similarity as an inhibitory term of the first similarity, and calculating a first confidence that the sample to be detected belongs to the target category; Calculating the product of the first similarity and the first confidence to obtain the classification confidence that the sample to be detected belongs to the target category.
6. The method according to any one of claims 1-5, characterized in that, Obtaining the description context of each target category includes: Obtaining the training samples corresponding to each target category; Initializing the description context of each target category by taking the description context of each target category as a parameter to be trained; Inputting the category name and the current description context of each target category into the text encoder of the pre-trained model for encoding to obtain a second text feature vector of each target category, and inputting the training sample into the sample encoder of the pre-trained model for encoding to obtain a feature vector of the training sample; Training the description context of each target category according to the similarity between the second text feature vector of each target category and the feature vector of the training sample.
7. The method according to claim 6, wherein The initialized description context of each target category includes an initialized perception context and a fallacy context. The step of inputting the category name and the current description context of each target category into the text encoder for encoding to obtain a second text feature vector of each target category includes: Concatenating the category name of each target category with the current perception context and inputting the concatenated result into the text encoder for encoding to obtain a second perception feature vector of each target category; Concatenating the category name of each target category with the current fallacy context and inputting the concatenated result into the text encoder for encoding to obtain a second fallacy feature vector of each target category.
8. The method according to claim 7, wherein The training sample includes in-distribution samples and out-of-distribution samples. The step of training the description context of each target category according to the similarity between the second text feature vector of each target category and the feature vector of the training sample includes: Calculating the cross-entropy loss according to the similarity between the feature vector of the in-distribution sample and the second text feature vector of each target category; Calculating the contrastive loss according to the similarity between the feature vector of the in-distribution sample and the second text feature vector of each target category, and the similarity between the feature vector of the out-of-distribution sample and the second text feature vector of the target category corresponding to the out-of-distribution sample; Training the description context of each target category according to the cross-entropy loss and the contrastive loss.
9. The method according to claim 6, characterized in that, The step of obtaining the training samples corresponding to each target category includes: Obtaining the in-distribution sample set of each target category; For any target category, determining the clustering center of the in-distribution sample set of the target category, and sampling a plurality of feature vectors in the feature vector space determined by the in-distribution sample set of the target category as the feature vectors of a plurality of out-of-distribution samples corresponding to the target category, where the distance between the sampled plurality of feature vectors and the clustering center is greater than or equal to a distance threshold.
10. The method according to claim 9, wherein After sampling the feature vectors of multiple out-of-distribution samples corresponding to the target category, it further includes: Adding noise data to the current perceptual context of the target category as the perceptual context of the out-of-distribution category of the target category; Calculating the seventh similarity between the feature vectors of the multiple out-of-distribution samples and the feature vectors of the perceptual context of the out-of-distribution category of the target category, and filtering out the feature vectors of the out-of-distribution samples whose seventh similarity is less than or equal to the similarity threshold.
11. An out-of-distribution detection method, characterized in that, It includes: The receiving end-side device receives a detection request for the segmentation image of the target object in the target detection result, obtains at least one interested category of the target detection task and the description context of each of the interested categories, and the description context contains the sample feature information of the interested category, and the description contexts of different interested categories are different; Inputting the category names and description contexts of each of the interested categories into the text encoder of the pre-trained model for encoding to obtain the third text feature vectors of each of the interested categories, and inputting the segmentation image into the image encoder of the pre-trained model for encoding to obtain the feature vector of the segmentation image; Determining whether the segmentation image is an out-of-distribution sample according to the similarity between the feature vector of the segmentation image and the third text feature vectors of each of the interested categories, and obtaining an out-of-distribution detection result; Returning the out-of-distribution detection result to the end-side device.
12. The method according to claim 11, characterized in that After determining whether the segmentation image is an out-of-distribution sample, it further includes: Calculating the false detection rate of the target detection result according to the number of the segmentation images and the number of out-of-distribution samples included in the segmentation images; Returning the false detection rate of the target detection result to the end-side device, and the false detection rate of the target detection result is used to optimize the target detection algorithm used for target detection.
13. A server, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the server to execute the method according to any one of claims 1-12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1-12 is implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-12 is implemented.
Citation Information
Cited By
Visual language model zero sample distribution external detection method, medium and computer equipment
CN121861453A
Visual language model zero-sample distribution out-of-distribution detection methods, media and computer equipment
CN121861453B