Image out-of-distribution detection method and device, electronic equipment and storage medium
By extracting the embedding vectors of the image and combining the classification lists inside and outside the distribution, the problem of low reliability of the out-of-distribution detection method is solved, the reliability and accuracy of the model are improved, and overconfident prediction of out-of-distribution data is avoided.
Patent Information
- Application Number
- CN202510119986.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-06-10
AI Technical Summary
At this stage, the reliability of the off-distribution detection method is low, which leads to excessive confidence in prediction of off-distribution data in deep neural networks, affecting the reliability and security of the model in practical applications.
By extracting the embedding vector of the image to be detected, and determining the detection result of the image to be detected based on the embedding vector, the classification list corresponding to the data in the text distribution, and the classification list corresponding to the data outside the text distribution. This method constructs a classification list inside and outside the distribution, and judges through similarity thresholds and scoring thresholds, avoiding the shortcomings of the traditional method.
It improves the reliability and accuracy of the model in complex environments, avoids excessive confidence prediction of out-of-distribution data, and enhances the robustness of the model.
Smart Images

Figure CN120125876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular, to a method, device, electronic device and storage medium for out-of-distribution detection of images. Background Art
[0002] Object recognition based on image classification is one of the basic tasks in the fields of computer vision and artificial intelligence. However, in the real world, when using neural network models based on deep learning, these models often face distribution shift problems such as covariate shift and semantic shift, which pose challenges to the robustness of the models.
[0003] Generally, the distribution of the training data of the model is called in-distribution (ID), and these data are obtained from one or more data distributions through a fixed sampling method; while the input data that does not match or is unexpected from the training data distribution is regarded as out-of-distribution (OOD). For example, if a model is trained only with pictures of fish in the Atlantic Ocean, when pictures of fish unique to the Pacific Ocean are input, this input belongs to out-of-distribution data because of semantic shift. Ideally, the model should be able to detect these samples and give a low confidence, but traditional deep neural networks often make overly confident predictions for OOD data. Common performance evaluation metrics only consider the closed-set accuracy, assuming that the model will only encounter known data. Even if the OOD samples do not belong to any known categories, the model will choose the "closest" one from these categories.
[0004] To solve this problem, researchers have proposed the concept of out-of-distribution detection. However, at present, out-of-distribution detection has the disadvantages of overconfident prediction of out-of-distribution data and imperfect evaluation metrics, which seriously affect the reliability and safety of the model in practical applications. Summary of the Invention
[0005] Embodiments of the present invention provide a method, device, electronic device and storage medium for out-of-distribution detection of images to solve the problem of low reliability of the current out-of-distribution detection method.
[0006] In a first aspect, an embodiment of the present invention provides a method for out-of-distribution detection of images, including:
[0007] Obtain an image to be detected;
[0008] Extract the embedding vector of the image to be detected;
[0009] Determine the detection result of the image to be detected according to the embedding vector, the classification list corresponding to the in-distribution data in text, and the classification list corresponding to the out-of-distribution data in text; wherein, the classification list corresponding to the in-distribution data in text and the classification list corresponding to the out-of-distribution data in text are respectively constructed based on the in-distribution data set and the in-distribution data set.
[0010] In a possible implementation, according to the embedding vector and the classification list corresponding to the data within the text distribution, determine the detection result of each image to be detected in the classification task to be processed, including:
[0011] Based on the embedding vector and the classification list corresponding to the data within the text distribution, calculate the first critical similarity corresponding to the image to be detected under each classification list corresponding to the data within the text distribution;
[0012] If the first critical similarity corresponding to the image to be detected under each classification list corresponding to the data within the text distribution is less than the first preset similarity threshold; then it is determined that the image to be detected belongs to a class of samples and belongs to out-of-distribution data.
[0013] In a possible implementation, according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution, determine the detection result of each image to be detected in the classification task to be processed, including:
[0014] If among the first critical similarities corresponding to the image to be detected under each classification list corresponding to the data within the text distribution, there is exactly one first critical similarity of the classification list corresponding to the data within the text distribution that is not less than the first preset similarity threshold, then according to the embedding vector and the classification list corresponding to the data outside the text distribution, calculate the second critical similarity corresponding to the image to be detected under the classification list corresponding to the data outside the text distribution;
[0015] If the second critical similarity corresponding to the image to be detected under the classification list corresponding to the data outside the text distribution is less than the second preset similarity threshold, then it is determined that the image to be detected belongs to a class of samples and belongs to in-distribution data.
[0016] In a possible implementation, according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution, determine the detection result of each image to be detected in the classification task to be processed, including:
[0017] If among the second critical similarities corresponding to the image to be detected under the classification list corresponding to the data outside the text distribution, there is a second critical similarity that is not less than the second preset similarity threshold, then determine that the image to be detected is a class of samples;
[0018] Calculate the similarities between the image to be detected and the classification list corresponding to the data within the text distribution and the perturbation variable group respectively, to obtain the in-distribution similarity and the out-of-distribution similarity; where the classification list corresponding to the data outside the text distribution;
[0019] Compare the in-distribution similarity and the out-of-distribution similarity;
[0020] If the in-distribution similarity is greater than the out-of-distribution similarity, and the total score of the in-distribution similarity is less than the preset score threshold, then it is determined that the image to be detected is out-of-distribution data.
[0021] In a possible implementation, according to the embedding vector, the classification list corresponding to the in-distribution data of the text, and the classification list corresponding to the out-of-distribution data of the text, determine the detection result of each image to be detected in the classification task to be processed, including:
[0022] If there are more than two first critical similarities not less than the first preset similarity threshold among the first critical similarities corresponding to the classification list of each in-distribution data of the text for the image to be detected, then it is determined that the image to be detected is a four-category sample; and based on the preset determination strategy, determine the detection result of the image to be detected.
[0023] In a possible implementation, based on the preset determination strategy, determine the detection result of the image to be detected, including:
[0024] If the preset determination strategy is a simple determination strategy, then it is determined that the image to be detected is in-distribution data;
[0025] If the preset determination strategy is a complete determination strategy, then determine whether there is a second critical similarity not less than the second preset similarity threshold among the in-distribution data of the text corresponding to the first critical similarity not less than the first preset similarity threshold in the image to be detected;
[0026] If there is, then calculate the similarity between the image to be detected and the in-distribution data of the text corresponding to the second critical similarity not less than the second preset similarity threshold, and the similarity of the perturbation variable group respectively, to obtain the in-distribution similarity and the out-of-distribution similarity;
[0027] Compare the in-distribution similarity and the out-of-distribution similarity;
[0028] If the in-distribution similarity is greater than the out-of-distribution similarity, and the total score of the in-distribution similarity is less than the preset score threshold, then it is determined that the image to be detected is out-of-distribution data; if the total score of the in-distribution similarity is not less than the preset score threshold, then it is determined that the image to be detected is in-distribution data.
[0029] In a possible implementation, the perturbation variable group is determined by the following method:
[0030] For each variable in the classification list corresponding to the out-of-distribution data of the text, obtain the perturbation variable group corresponding to the variable through the text encoder.
[0031] In a second aspect, an out-of-distribution detection device for images provided by an embodiment of the present invention includes:
[0032] An acquisition module, configured to acquire an image to be detected;
[0033] An extraction module for extracting an embedding vector of an image to be detected;
[0034] A detection module for determining a detection result of the image to be detected according to the embedding vector, a classification list corresponding to the in-distribution data of the text, and a classification list corresponding to the out-of-distribution data of the text; wherein, the classification list corresponding to the in-distribution data of the text and the classification list corresponding to the out-of-distribution data of the text are respectively constructed based on the in-distribution data set and the in-distribution data set.
[0035] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.
[0036] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.
[0037] An embodiment of the present invention provides an out-of-distribution detection method, device, electronic device, and storage medium for images. When classifying an image to be detected, the embodiment of the present invention extracts an embedding vector for characterizing the features of the image to be detected, and then respectively analyzes and calculates the extracted embedding vector with a classification list corresponding to the in-distribution data of the text and a classification list corresponding to the out-of-distribution data of the text determined in advance to obtain corresponding calculation results, and can accurately judge the category to which the image to be detected belongs according to the set judgment rules, avoiding the overconfident prediction of the out-of-distribution data by the traditional deep neural network, thereby improving the reliability and accuracy of the model in a complex environment. In addition, compared with the traditional single-category text, the embodiment of the present invention comprehensively considers the in-distribution data and out-of-distribution data of the text to obtain the corresponding classification list, better captures the data distribution characteristics in the analysis process, provides a richer discrimination basis for out-of-distribution detection, and further improves the detection effect. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a flowchart of the implementation of the out-of-distribution detection method for images provided by the embodiment of the present invention;
[0040] Figure 2 It is a flowchart of the implementation of the out-of-distribution detection method for images provided by another embodiment of the present invention;
[0041] Figure 3 It is a schematic structural diagram of the out-of-distribution detection device for images provided by an embodiment of the present invention;
[0042] Figure 4 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0043] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.
[0045] Figure 1 It is a flowchart of the implementation of the out-of-distribution detection method for images provided by an embodiment of the present invention. As Figure 1 shown, the method may include:
[0046] Step 110: Obtain the image to be detected.
[0047] In this embodiment, the image to be detected may be a single image or multiple images.
[0048] In this embodiment, each image to be detected corresponds to a specific classification task. Based on the classification task, the corresponding ID classification list and the preset conditional variables corresponding to the ID classification list can be found in the pre-constructed dataset; wherein, the preset conditional variables may include the number of out-of-distribution classifications to be maintained for each ID classification, the number of in-distribution texts to be maintained for each ID classification, the number of absolute perturbation variables, the number of relative perturbation variables, the critical similarity of positive samples, and the critical similarity of negative samples; the pre-constructed dataset includes an in-distribution dataset and an in-distribution dataset.
[0049] In this embodiment, the pre-constructed dataset further includes the classification list corresponding to the in-distribution data in the text distribution corresponding to each ID classification list, the classification list corresponding to the out-of-distribution data, and the perturbation variable group. Among them, the classification list corresponding to the in-distribution data and the classification list corresponding to the out-of-distribution data are pre-constructed based on the in-distribution dataset and the in-distribution dataset respectively, and the perturbation variable group is determined by the classification list corresponding to the out-of-distribution data corresponding to it.
[0050] The determination process of the classification list corresponding to the in-distribution data in the text distribution is described below:
[0051] In this embodiment, the classification list corresponding to the in-distribution data in each ID classification list is jointly determined by the number of in-distribution data that each ID classification needs to maintain and the in-distribution text obtained after processing the in-distribution image data.
[0052] The following uses an optional embodiment to illustrate how to determine the classification list corresponding to the in-distribution data in the text distribution.
[0053] Step 1: For the in-distribution images in each ID classification list, send them into the visual feature extraction module to obtain their visual embedding vectors x e .
[0054] Exemplarily, the visual encoder in the visual feature extraction module can be a visual encoder based on Vision-Transformer.
[0055] In the visual feature extraction module, the Python third-party library PIL library can be used to read the picture file and perform corresponding preprocessing, and then input the preprocessed image into the Vision-Transformer-based visual encoder to obtain its feature encoding x'. e . Then, through post-processing steps to perform further feature extraction on the obtained feature encoding x' e to obtain the final embedding vector x e . In the embodiment of the present invention, a fully connected layer can be used to implement this function.
[0056] Step 2: Based on the obtained visual embedding vector x e , use a vision-language model with text generation ability to construct its description text. Exemplarily, the vision-language model can be a pre-trained model (Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, BLIP2) that freezes the existing vision model and the large language model.
[0057] Optionally, through the Q-Former encoder in the BLIP2 model, based on the pre-trained Query Token query vector and the visual embedding vector x obtained in step one e , feature extraction and text decoding are performed, and a specific text description t is generated d . The in-distribution image is a transport plane. After processing, a text description with the content of "a large military plane flying through a cloudy sky" is generated
[0058] Step three: Based on the obtained text description t d , construct the in-distribution text t ID to obtain the classification list corresponding to the in-distribution data of the text
[0059] Exemplarily, this step can be implemented in this embodiment by a large language model with conversation ability (Virtual Large Language Model, VLLM), such as pre-trained language models (Large Language Model Meta AI, LLaMa), large-scale vision language models (Qwen Large Vision Language Model, Qwen-VL), etc. This embodiment is illustrated by taking Qwen-VL as an example
[0060] In this embodiment, since the obtained text description is the overall description of the picture, and the image data inevitably contains noise such as the image background in addition to the image foreground, the obtained text description t d usually inevitably contains text irrelevant to the image subject. Exemplarily, if the in-distribution image is a transport plane, then "flying through a cloudy sky" can be considered as text irrelevant to the nature of the image subject
[0061] To obtain in-distribution data of the text with higher quality and less noise, it is necessary to post-process this text and screen out the text "a large military plane" related to the subject as the final text description t ID , that is, the in-distribution text data
[0062] The in-distribution data t of the text obtained by this method ID has stronger distribution description ability than the category text composed of a single category text of "a photo of......", because it retains the partial covariate shift of the original distribution compared to the category text. Therefore, the in-distribution data t of the text obtained through this link IDIt has a more powerful detection ability when facing difficult out-of-distribution detection. By collecting t constructed through this embodiment ID , an embedded vector feature library T can be constructed ID and embedded as an optional feature into the out-of-distribution detection module
[0063] The determination process of the classification list corresponding to the out-of-distribution data of the text will be described first below
[0064] In this embodiment, the classification list corresponding to the out-of-distribution data of the text in each ID classification list is jointly determined by the number of out-of-distribution classifications that each ID classification needs to maintain and the ID classification task described above. Exemplarily, when the classification task is ImageNet-10 vs ImageNet-20 and τ = 1, it means selecting the class with the highest similarity to the ID classification list from the ImageNet20 classification task as a member of the classification list corresponding to the out-of-distribution data of the text (i.e., Out-of-Distribution Detection, OOD); where "the highest similarity" means the highest similarity of the embedded vectors obtained by encoding the text concept descriptions of the two categories via the Text Encoder; τ is the number of out-of-distribution classifications that each ID classification needs to maintain
[0065] Exemplarily, for each ID classification list, the category information in the ID classification list is encoded through the text encoding module to obtain text embedded vectors
[0066] Retain 1 category with the most similar category text to its own among the category texts of all OOD classifications as its out-of-distribution data of the text[n ij .
[0067] The determination process of the perturbation variable group will be described below
[0068] In this embodiment, the perturbation variable group is composed of α absolute perturbation variables and β relative perturbation variables. Among them, the absolute perturbation variable is obtained by adding a random variable with a mean of 0 to the embedded vector obtained by encoding the text concept description of the OOD class n ij , and the relative perturbation variable is obtained by encoding the text concept description of the OOD class n ij after being described and expanded by the description expansion module through the text encoder
[0069] Exemplarily, two relative perturbation variables can be constructed, calculate the mean of the offsets of these two relative perturbation variables compared to the category text, and randomly generate 1 absolute perturbation variable within the hypersphere region centered on the embedded vector of the OOD category text with this mean as the radius, and mark the above three perturbation variables as the perturbation variable group[n ijk .
[0070] Step 120: Extract the embedding vector of the image to be detected.
[0071] In this embodiment, the image to be detected can be subjected to feature extraction through a visual feature extraction module to obtain its corresponding embedding vector.
[0072] Exemplarily, a Vision-Transformer-based visual encoder can be used as the visual encoder in the visual feature extraction module. After reading the picture file through the PIL library and performing corresponding preprocessing, it is input into the Vision-Transformer-based visual encoder to obtain its feature encoding x'. e 。
[0073] Then, through post-processing steps, the obtained feature encoding is further subjected to feature extraction to obtain the final embedding vector x e (Visual embedding vector).
[0074] Step 130: Determine the detection result of the picture to be detected according to the embedding vector, the classification list corresponding to the in-distribution data of the text, and the classification list corresponding to the out-of-distribution data of the text; wherein, the classification list corresponding to the in-distribution data of the text and the classification list corresponding to the out-of-distribution data of the text are respectively constructed based on the in-distribution data set and the in-distribution data set.
[0075] In this embodiment, the extracted embedding vector is respectively analyzed and calculated with the pre-determined classification list corresponding to the in-distribution data of the text and the classification list corresponding to the out-of-distribution data of the text to obtain corresponding calculation results, and the category to which the image to be detected belongs can be accurately judged according to the set judgment rules.
[0076] In summary, in the embodiment of the present invention, when classifying the image to be detected, by extracting the embedding vector for characterizing the features of the image to be detected, and then respectively analyzing and calculating the extracted embedding vector with the pre-determined classification list corresponding to the in-distribution data of the text and the classification list corresponding to the out-of-distribution data of the text to obtain corresponding calculation results, the category to which the image to be detected belongs can be accurately judged according to the set judgment rules, avoiding the overconfident prediction of the out-of-distribution data by the traditional deep neural network, thereby improving the reliability and accuracy of the model in a complex environment. In addition, compared with the traditional single-category text, the embodiment of the present invention comprehensively considers the in-distribution data and out-of-distribution data of the text to obtain the corresponding classification list, better captures the data distribution characteristics in the analysis process, provides a richer discriminant basis for out-of-distribution detection, and further improves the detection effect.
[0077] In an alternative embodiment, in step 130, determining the detection result of each picture to be detected in the classification task to be processed according to the embedding vector and the classification list corresponding to the in-distribution data of the text may include:
[0078] Based on the classification list corresponding to the data within the text distribution of the embedding vector, calculate the first critical similarity of the image to be detected under the classification list corresponding to the data within each text distribution.
[0079] If the first critical similarity of the image to be detected under the classification list corresponding to the data within each text distribution is less than the first preset similarity threshold; then it is determined that the image to be detected belongs to a certain type of sample and belongs to out-of-distribution data.
[0080] In this embodiment, in the target classification task corresponding to the image to be detected, determine the ID classification list under this target classification task. For the classification list corresponding to the data within the text distribution corresponding to the determined ID classification list, calculate the first critical similarity of the detected image under the classification list corresponding to the data within each text distribution.
[0081] If the first critical similarity of the image to be detected under the classification list corresponding to the data within each text distribution is less than the first preset similarity threshold; then it is determined that the image to be detected belongs to a certain type of sample and belongs to out-of-distribution data.
[0082] That is, if all the first critical similarities satisfy the following conditions, it is considered that the image to be detected belongs to a certain type of sample and belongs to out-of-distribution data:
[0083]
[0084] Among them, is used to represent the first critical similarity; y i is the data within the i-th text distribution; x e is the embedding vector; θ p is the positive sample critical similarity.
[0085] In this embodiment, a certain type of sample is a sample determined to have semantic deviation.
[0086] In an alternative embodiment, in step 130, according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution, determining the detection result of each image to be detected in the classification task to be processed may include:
[0087] If among the first critical similarities of the image to be detected under the classification list corresponding to the data within each text distribution, there is exactly one first critical similarity of the classification list corresponding to the data within a text distribution that is not less than the first preset similarity threshold, then according to the embedding vector and the classification list corresponding to the data outside the text distribution, calculate the second critical similarity of the image to be detected under the classification list corresponding to the data outside the text distribution;
[0088] If all the second critical similarities corresponding to the test image under the classification list of out-of-distribution data are less than the second preset similarity threshold, it is determined that the test image belongs to the second-class samples and belongs to the in-distribution data.
[0089] In this embodiment, if there is exactly one first critical similarity that is not less than the first preset similarity threshold under the classification list of all in-distribution data of the text, that is:
[0090]
[0091] Then it is considered that the test image does not belong to the first-class samples; then further judgment needs to be made on this test image.
[0092] According to the embedding vector and the classification list corresponding to the out-of-distribution data of the text, calculate the second critical similarity corresponding to the test image under the classification list corresponding to the out-of-distribution data of the text. If all the calculated second critical similarities are less than the second preset similarity threshold, it is determined that the test image belongs to the second-class samples and belongs to the in-distribution data. That is:
[0093]
[0094] Among them, used to represent the second critical similarity; n ij is the i-th data in the out-of-distribution classification list; θ n is the negative sample critical similarity.
[0095] In this embodiment, the second-class samples are samples that have been judged by the model and do not have semantic drift.
[0096] In an alternative embodiment, step 130 may include determining the detection result of each test image in the to-be-processed classification task according to the embedding vector, the classification list corresponding to the in-distribution data of the text, and the classification list corresponding to the out-of-distribution data of the text:
[0097] If there is a second critical similarity that is not less than the second preset similarity threshold among the second critical similarities corresponding to the test image under the classification list corresponding to the out-of-distribution data of the text, determine that the test image is a third-class sample.
[0098] Calculate the similarities between the test image and the classification list corresponding to the in-distribution data of the text and the perturbation variable group respectively to obtain the in-distribution similarity and the out-of-distribution similarity; among them, the classification list corresponding to the out-of-distribution data of the text.
[0099] Compare the in-distribution similarity and the out-of-distribution similarity.
[0100] If the in-distribution similarity is greater than the out-of-distribution similarity, and the total score of the in-distribution similarity is less than the preset score threshold, then it is determined that the image to be detected is out-of-distribution data.
[0101] In this embodiment, if there is a second critical similarity that is not less than the second preset similarity threshold, then it is considered that the image to be detected is not a type-two sample, but a type-three sample. That is:
[0102]
[0103] In this case, for each quantity in the embedding vector obtained for the image to be detected, calculate its similarity with each data in the classification list corresponding to the in-distribution data of the text and the perturbation variable group respectively. For any one quantity, compare its in-distribution similarity and out-of-distribution similarity respectively. If the in-distribution similarity is greater than the out-of-distribution similarity, then set the out-of-distribution similarity to zero and do not score it. Statistically calculate the total score of the in-distribution similarity corresponding to the image to be detected. If the calculated total score of the in-distribution similarity is less than the preset score threshold, then it is determined that the image to be detected is out-of-distribution data; otherwise, it is considered that the image to be detected is in-distribution data.
[0104] In this embodiment, a type-three sample means that there is an in-distribution data classification list for the image to be detected that satisfies semantic similarity and there is also an out-of-distribution data classification list that satisfies semantic similarity. We consider this to be a Hard-OOD problem. For example, it can be understood as the problem of distinguishing a coyote (out-of-distribution) from a husky (in-distribution).
[0105] Exemplarily, assume that the embedding vector includes 1 -a n quantities. For each quantity in this embedding vector, if the in-distribution similarity corresponding to 1 a is greater than its corresponding out-of-distribution similarity, then the in-distribution similarity corresponding to 1 a can be recorded as 1, and its corresponding out-of-distribution similarity can be recorded as 0; if the in-distribution similarity corresponding to 2 a is less than its corresponding out-of-distribution similarity, then the in-distribution similarity corresponding to 2 a can be recorded as 0, and its corresponding out-of-distribution similarity can be recorded as 1. Statistically calculate the total score corresponding to the in-distribution similarity of each quantity in the embedding vector to determine whether the total score of the in-distribution similarity is less than the preset score threshold, so as to determine whether the image to be detected is in-distribution data or out-of-distribution data.
[0106] In an alternative embodiment, in step 130, according to the embedding vector, the classification list corresponding to the in-distribution data of the text, and the classification list corresponding to the out-of-distribution data of the text, determining the detection result of each image to be detected in the classification task to be processed may include:
[0107] If there are two or more first critical similarities that are not less than the first preset similarity threshold value in the first critical similarities corresponding to the classification list corresponding to the data in each text distribution of the image to be detected, then the image to be detected is determined to be a four-category sample; and based on the preset judgment strategy, the detection result of the image to be detected is determined.
[0108] In this embodiment, if there are two or more first critical similarities that are not less than the first preset similarity threshold, it is determined that the image to be detected is a sample of four categories. That is:
[0109]
[0110] For this embodiment, a more popular explanation is that there are two images similar to the image to be detected in the data within the distribution.
[0111] In this case, the following optional embodiment can be used to determine whether it is out-of-distribution data.
[0112] In an optional embodiment, determining the detection result of the image to be detected based on a preset determination strategy may include:
[0113] If the preset determination strategy is a simple determination strategy, the image to be detected is determined to be in-distribution data.
[0114] If the preset determination strategy is a complete determination strategy, it is determined whether there is a second critical similarity not less than a second preset similarity threshold in the text distribution data corresponding to the first critical similarity not less than the first preset similarity threshold in the image to be detected.
[0115] If so, the similarity between the image to be detected and the text distribution data corresponding to the second critical similarity not less than the second preset similarity threshold and the disturbance variable group is calculated respectively to obtain the in-distribution similarity and the out-of-distribution similarity.
[0116] Compare in-distribution similarity with out-of-distribution similarity.
[0117] If the similarity within the distribution is greater than the similarity outside the distribution, and the similarity within the distribution is less than the preset scoring threshold, the image to be detected is determined to be out-of-distribution data; if the similarity within the distribution is not less than the preset scoring threshold, the image to be detected is determined to be in-distribution data.
[0118] In this embodiment, there are two preset determination strategies. The first one is a simple determination strategy. If this strategy is adopted, it can be directly considered that the image to be detected belongs to the in-distribution data.
[0119] The second is the complete judgment strategy, under which the image to be detected needs to be further judged and processed. Specifically:
[0120] Determine the classification list corresponding to the in-distribution data with the first critical similarity not less than the first preset similarity threshold according to the first critical similarity not less than the first preset similarity threshold, and determine the corresponding classification list of the out-of-distribution data according to the determined classification list of the in-distribution data.
[0121] Calculate the second critical similarity between the embedding vector corresponding to the image to be detected and the classification list corresponding to the determined out-of-distribution data. If there is at least one second critical similarity not less than the second preset similarity threshold, it is considered that the property of the image to be detected satisfies the properties of the three types of samples. For example, in this case, there are two images with relatively high similarity in the in-distribution data for the image to be detected, and there is one image with relatively high similarity in the out-of-distribution data.
[0122] In this case, for each quantity in the embedding vector obtained for the image to be detected, calculate its similarity with the classification list corresponding to the in-distribution data of the text and each data in the perturbation variable group respectively. For any one quantity, compare the in-distribution similarity and the out-of-distribution similarity of this quantity respectively. If the in-distribution similarity is greater than the out-of-distribution similarity, set the out-of-distribution similarity to zero and do not score. Statistically calculate the total in-distribution similarity score corresponding to the image to be detected. If the calculated total in-distribution similarity score is less than the preset score threshold, determine that the image to be detected is out-of-distribution data, otherwise it is considered that the image to be detected is in-distribution data, and its classification label can be: among the classification lists corresponding to the in-distribution data that meet the critical similarity condition, the one with the highest similarity, that is:
[0123]
[0124] Wherein, in this embodiment, the critical similarity condition characterizes whether the corresponding critical similarity meets its corresponding preset critical similarity value.
[0125] Figure 2 is the implementation flowchart of the out-of-distribution detection method for images provided by another embodiment of the present invention. As Figure 2 shown, this method can be divided into a pre-set data preparation process and a process of detecting the image to be detected according to the pre-set data. Correspondingly, the method is specifically as follows:
[0126] The data preparation process may include:
[0127] Step 1: According to the specific classification task, formulate the corresponding classification strategy and detection strategy. Specifically, in this embodiment, it is agreed that the detection task is to use the CIFAR-10 dataset as the in-distribution data and the SVHN dataset as the out-of-distribution data, and the preset conditional variables are as follows:
[0128] τ = 1, ε = 1, α = 1, β = 2, θ p= arccos(0.95), θ n = arccos(0.90).
[0129] where τ is the number within the text distribution that needs to be maintained for each ID classification, ε is the number outside the text distribution that needs to be maintained for each ID classification; α is the number of absolute perturbation variables, β is the number of relative perturbation variables, and θ p is the critical similarity for positive samples, and θ n is the critical similarity for negative samples.
[0130] Step 2. For each ID classification list, maintain a set of OOD classification lists [n ij and a set of in-distribution data [t ij .
[0131] Under the preset condition variables set in this embodiment, first construct the class texts of CIFAR-10 and SVHN. Specifically, before each class text in these two datasets, add the prompt word prefix "a photo of"; and based on this, send it into the text encoding module in the foregoing technical solution to obtain its text embedding vector and construct the benchmark embedding vector feature library T ID and T OOD . Then, according to the conditions given in Step 1, for each ID classification, retain the 1 class with the most similar class text to its own among all the class texts in the OOD classification list as its [n ij .
[0132] Step 3. For each n ij in each OOD classification list [n ij , construct a perturbation variable group [n ijk .
[0133] In this embodiment, 2 relative perturbation variables can be constructed under the conditions determined in Step 1, calculate the mean of the offsets of these 2 relative perturbation variables compared to the class text, and randomly generate 1 absolute perturbation variable within the hypersphere region centered on the embedding vector of the OOD class text with this mean as the radius, and mark the above three perturbation variables as [n ijk .
[0134] In this embodiment, the detection process can include:
[0135] Step 4. For each image x to be detected input , construct its embedding vector x e .
[0136] Step 5. Based on the embedding vector x e obtained in Step 4, determine whether it satisfies the following conditions to determine the class of the image to be detected:
[0137] Based on the embedding vector, calculate the cosine similarity with the classification list corresponding to the distributed interior trim to obtain the first critical similarity. If the first critical similarity corresponding to the image to be detected under the classification list corresponding to the data in each text distribution is less than the first preset similarity threshold; then it is determined that the image to be detected belongs to a class of samples and belongs to out-of-distribution data.
[0138] That is, if all the first critical similarities satisfy the following conditions, it is considered that the image to be detected belongs to a class of samples and belongs to out-of-distribution data and a report is generated, and then it is judged whether all the images to be detected have been completed. If so, the process ends; if not, continue the detection:
[0139]
[0140] Step Six. The embedding vector x obtained in Step Four e , and determine whether it satisfies the following conditions to determine the category of the image to be detected:
[0141] If there is exactly one first critical similarity that is not less than the first preset similarity threshold under the classification list corresponding to the data in all text distributions, that is:
[0142]
[0143] Then it is considered that the image to be detected does not belong to a class of samples; then further judgment needs to be made on the image to be detected.
[0144] According to the embedding vector and the classification list corresponding to the out-of-text-distribution data, calculate the second critical similarity corresponding to the image to be detected under the classification list corresponding to the out-of-text-distribution data. If all the calculated second critical similarities are less than the second preset similarity threshold, it is determined that the image to be detected belongs to a class of samples and belongs to in-distribution data. That is:
[0145]
[0146] Otherwise, it is considered that the sample does not belong to a class of samples, and subsequent steps are carried out.
[0147] Step Seven. Based on the embedding vector x obtained in Step Four e , and determine whether it satisfies the following conditions to determine the category of the image to be detected:
[0148] If there is exactly one first critical similarity that is not less than the first preset similarity threshold under the classification list corresponding to the data in all text distributions, that is:
[0149]
[0150] Then it is considered that the image to be detected does not belong to the first type of samples; then further judgment needs to be made on this image to be detected.
[0151] According to the embedding vector and the classification list corresponding to the out-of-distribution data of the text, calculate the second critical similarity corresponding to the image to be detected under the classification list corresponding to the out-of-distribution data of the text. If there is a second critical similarity not less than the second preset similarity threshold, then it is considered that the image to be detected is not a second type of sample, but a third type of sample. That is:
[0152]
[0153] Then it is considered that the sample belongs to the third type of samples. For each quantity in the embedding vector obtained for the image to be detected, calculate its similarity with each data in the classification list corresponding to the in-distribution data of the text and the perturbation variable group respectively. For any one quantity, compare its in-distribution similarity and out-of-distribution similarity respectively. If the in-distribution similarity is greater than the out-of-distribution similarity, then set the out-of-distribution similarity to zero and do not score. Statistically calculate the total in-distribution similarity score corresponding to the image to be detected. If the calculated total in-distribution similarity score is less than the preset score threshold, then determine that the image to be detected is out-of-distribution data, otherwise it is considered that the image to be detected is in-distribution data.
[0154] Step eight: Based on the embedding vector x obtained in step four e , determine whether it meets the following conditions to determine the category of the image to be detected:
[0155] If there are two or more first critical similarities not less than the first preset similarity threshold under the classification list corresponding to all in-distribution data of the text, that is:
[0156]
[0157] Then it is considered that the sample belongs to the fourth type of samples.
[0158] For the processing of the fourth type of samples, it can be selected based on different preset determination strategies:
[0159] Simple determination strategy. If this strategy is adopted, it can be directly considered that the image to be detected belongs to in-distribution data.
[0160] Full determination strategy. Under this strategy, for each y that meets the critical similarity condition i , according to the first critical similarity not less than the first preset similarity threshold, determine the classification list corresponding to the in-distribution data where the first critical similarity is not less than the first preset similarity threshold, and according to the determined classification list corresponding to the in-distribution data, determine the corresponding classification list of the out-of-distribution data.
[0161] Calculate a second critical similarity between the embedding vector corresponding to the image to be detected and the classification list corresponding to the determined out-of-distribution data. If there is at least one second critical similarity not less than the second preset similarity threshold, it is considered that the property of the image to be detected satisfies the properties of the three types of samples.
[0162] For each quantity in the embedding vector obtained for the image to be detected, calculate its similarity with the classification list corresponding to the in-distribution data of the text and each data in the perturbation variable group respectively. For any one quantity, compare the in-distribution similarity and the out-of-distribution similarity of this quantity respectively. If the in-distribution similarity is greater than the out-of-distribution similarity, set the out-of-distribution similarity to zero and do not score. Statistically calculate the total in-distribution similarity score corresponding to the image to be detected. If the calculated total in-distribution similarity score is less than the preset score threshold, determine that the image to be detected is out-of-distribution data, otherwise, consider that the image to be detected is in-distribution data. Its classification label is the one with the highest similarity in [y i , that is:
[0163]
[0164] In summary, the method provided by the embodiments of the present invention has the advantages of being able to detect input pictures with out-of-distribution characteristics, avoiding to a certain extent the phenomenon of the model being overconfident about out-of-distribution data, and improving the robustness of the model in the face of out-of-distribution data.
[0165] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0166] The following is an apparatus embodiment of the present invention. For details not described in detail, reference may be made to the corresponding method embodiments above.
[0167] Figure 3 The structural schematic diagram of the out-of-distribution detection apparatus for images provided by the embodiments of the present invention is shown. For the sake of convenience of description, only parts related to the embodiments of the present invention are shown and are described in detail as follows:
[0168] As Figure 3 shown, the out-of-distribution detection apparatus 3 for images includes:
[0169] An acquisition module 31, configured to acquire an image to be detected;
[0170] An extraction module 32, configured to extract the embedding vector of the image to be detected;
[0171] The detection module 33 is used to determine the detection result of the image to be detected according to the embedding vector, the classification list corresponding to the in-distribution data of the text distribution, and the classification list corresponding to the out-of-distribution data of the text distribution; wherein, the classification list corresponding to the in-distribution data of the text distribution and the classification list corresponding to the out-of-distribution data of the text distribution are respectively constructed based on the in-distribution data set and the in-distribution data set.
[0172] In a possible implementation manner, the detection module 33 is specifically used for:
[0173] Based on the embedding vector and the classification list corresponding to the in-distribution data of the text distribution, calculate the first critical similarity corresponding to the image to be detected under each classification list corresponding to the in-distribution data of the text distribution;
[0174] If the first critical similarity corresponding to the image to be detected under each classification list corresponding to the in-distribution data of the text distribution is less than the first preset similarity threshold; then it is determined that the image to be detected belongs to a class of samples and belongs to out-of-distribution data.
[0175] In a possible implementation manner, the detection module 33 is specifically used for:
[0176] If among the first critical similarities corresponding to the image to be detected under each classification list corresponding to the in-distribution data of the text distribution, there is exactly one first critical similarity of the classification list corresponding to the in-distribution data of the text distribution that is not less than the first preset similarity threshold, then according to the embedding vector and the classification list corresponding to the out-of-distribution data of the text distribution, calculate the second critical similarity corresponding to the image to be detected under the classification list corresponding to the out-of-distribution data of the text distribution;
[0177] If the second critical similarity corresponding to the image to be detected under the classification list corresponding to the out-of-distribution data of the text distribution is less than the second preset similarity threshold, then it is determined that the image to be detected belongs to a class of samples and belongs to in-distribution data.
[0178] In a possible implementation manner, the detection module 33 is specifically used for:
[0179] If there is a second critical similarity that is not less than the second preset similarity threshold among the second critical similarities corresponding to the image to be detected under the classification list corresponding to the out-of-distribution data of the text distribution, then determine that the image to be detected is a class of samples;
[0180] Calculate the similarities between the image to be detected and the classification list corresponding to the in-distribution data of the text distribution and the group of perturbation variables respectively, to obtain the in-distribution similarity and the out-of-distribution similarity; wherein, the classification list corresponding to the out-of-distribution data of the text distribution;
[0181] Compare the in-distribution similarity and the out-of-distribution similarity;
[0182] If the in-distribution similarity is greater than the out-of-distribution similarity, and the total score of the in-distribution similarity is less than the preset score threshold, then it is determined that the image to be detected is out-of-distribution data.
[0183] In a possible implementation manner, the detection module 33 is specifically configured to:
[0184] If there are more than two first critical similarities that are not less than the first preset similarity threshold among the first critical similarities corresponding to the classification list of each in-distribution data of the image to be detected, then it is determined that the image to be detected is a four-category sample; and based on the preset determination strategy, the detection result of the image to be detected is determined.
[0185] In a possible implementation manner, the detection module 33 is specifically configured to:
[0186] If the preset determination strategy is a simple determination strategy, then it is determined that the image to be detected is in-distribution data;
[0187] If the preset determination strategy is a complete determination strategy, then it is determined whether there is a second critical similarity that is not less than the second preset similarity threshold among the in-distribution data corresponding to the first critical similarities that are not less than the first preset similarity threshold in the image to be detected;
[0188] If there is, then the similarities between the image to be detected and the in-distribution data corresponding to the second critical similarities that are not less than the second preset similarity threshold, as well as the perturbation variable group, are calculated respectively to obtain the in-distribution similarity and the out-of-distribution similarity;
[0189] Compare the in-distribution similarity and the out-of-distribution similarity;
[0190] If the in-distribution similarity is greater than the out-of-distribution similarity, and the total score of the in-distribution similarity is less than the preset score threshold, then it is determined that the image to be detected is out-of-distribution data; if the total score of the in-distribution similarity is not less than the preset score threshold, then it is determined that the image to be detected is in-distribution data.
[0191] In a possible implementation manner, the perturbation variable group is determined by the following method:
[0192] For each variable in the classification list corresponding to the out-of-distribution data of the text, the corresponding perturbation variable group is obtained through the text encoder.
[0193] Figure 4 It is a schematic diagram of the electronic device provided by the embodiments of the present invention. As Figure 4As shown, the electronic device 4 of this embodiment includes: a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. When the processor 40 executes the computer program 42, it implements the steps in the embodiments of the out-of-distribution detection method for each of the above images. For example Figure 1 the steps 110 to 130 shown. Alternatively, when the processor 40 executes the computer program 42, it implements the functions of each module / unit in the above device embodiments. For example Figure 3 the functions of the modules shown.
[0194] Exemplarily, the computer program 42 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 41 and executed by the processor 40 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 42 in the electronic device 4. For example, the computer program 42 may be divided into Figure 3 the modules shown.
[0195] The electronic device 4 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device 4 may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 these are merely examples of the electronic device 4 and do not constitute a limitation on the electronic device 4. It may include more or fewer components than shown, or combine certain components, or have different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0196] The so-called processor 40 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0197] The memory 41 may be an internal storage unit of the electronic device 4, such as a hard disk or memory of the electronic device 4. The memory 41 may also be an external storage device of the electronic device 4, such as a plug-in hard disk equipped on the electronic device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 41 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 41 is used to store the computer program and other programs and data required by the electronic device. The memory 41 may also be used to temporarily store data that has been output or will be output.
[0198] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0199] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0200] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0201] In the embodiments provided by the present invention, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0202] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0204] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments for out-of-distribution detection of each image can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0205] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for out-of-distribution detection of an image, characterized in that: include: Acquire the image to be detected; Extracting an embedding vector of the image to be detected; The detection result of the image to be detected is determined according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution; wherein the classification list corresponding to the data within the text distribution and the classification list corresponding to the data outside the text distribution are constructed based on the in-distribution data set and the in-distribution data set, respectively.
2. The out-of-distribution detection method of an image according to claim 1, characterized in that: Determining the detection result of each to-be-detected picture in the classification task to be processed according to the embedding vector and the classification list corresponding to the data in the text distribution includes: Based on the embedding vector and the classification list corresponding to the data in the text distribution, calculate the first critical similarity of the image to be detected under the classification list corresponding to the data in each text distribution; If the first critical similarity of the image to be detected in the classification list corresponding to each text distribution data is less than the first preset similarity threshold; then it is determined that the image to be detected belongs to a type of sample and is out-of-distribution data.
3. The out-of-distribution detection method of an image according to claim 2, characterized in that: Determining the detection result of each to-be-detected picture in the classification task to be processed according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution includes: If, among the first critical similarities corresponding to the image to be detected under the classification list corresponding to each data within the text distribution, there is only one first critical similarity of the classification list corresponding to the data within the text distribution that is not less than the first preset similarity threshold, then according to the embedding vector and the classification list corresponding to the data outside the text distribution, calculate the second critical similarity corresponding to the image to be detected under the classification list corresponding to the data outside the text distribution; If the second critical similarity corresponding to the image to be detected in the classification list corresponding to the data outside the text distribution is less than the second preset similarity threshold, it is determined that the image to be detected belongs to the second category sample and belongs to the data within the distribution.
4. The out-of-distribution detection method of an image according to claim 3, characterized in that: Determining the detection result of each to-be-detected picture in the classification task to be processed according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution includes: If the image to be detected has a second critical similarity that is not less than the second preset similarity threshold value in the second critical similarities corresponding to the classification list corresponding to the data outside the text distribution, then the image to be detected is determined to be a three-category sample; Calculate the similarity between the classification list and the disturbance variable group corresponding to the image to be detected and the data within the text distribution respectively, and obtain the similarity within the distribution and the similarity outside the distribution; wherein the classification list corresponding to the data outside the text distribution; comparing the in-distribution similarity and the out-of-distribution similarity; If the in-distribution similarity is greater than the out-of-distribution similarity, and the total score of the in-distribution similarity is less than a preset score threshold, it is determined that the image to be detected is out-of-distribution data.
5. The out-of-distribution detection method of an image according to claim 2, characterized in that: Determining the detection result of each to-be-detected picture in the classification task to be processed according to the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution includes: If there are two or more first critical similarities that are not less than the first preset similarity threshold value in the first critical similarities corresponding to the classification list corresponding to the data in each text distribution of the image to be detected, then the image to be detected is determined to be a four-category sample; and based on the preset judgment strategy, the detection result of the image to be detected is determined.
6. The out-of-distribution detection method of an image according to claim 5, characterized in that: The step of determining the detection result of the image to be detected based on a preset determination strategy includes: If the preset determination strategy is a simple determination strategy, determining that the image to be detected is in-distribution data; If the preset determination strategy is a complete determination strategy, determining whether there is a second critical similarity not less than a second preset similarity threshold in the text distribution data corresponding to the first critical similarity not less than the first preset similarity threshold in the image to be detected; If so, respectively calculating the similarity between the image to be detected and the text distribution data corresponding to the second critical similarity not less than the second preset similarity threshold, and the disturbance variable group, to obtain the similarity within the distribution and the similarity outside the distribution; comparing the in-distribution similarity and the out-of-distribution similarity; If the in-distribution similarity is greater than the out-distribution similarity, and the total score of the in-distribution similarity is less than the preset score threshold, then the image to be detected is determined to be out-of-distribution data; if the total score of the in-distribution similarity is not less than the preset score threshold, then the image to be detected is determined to be in-distribution data.
7. The out-of-distribution detection method of an image according to claim 4, characterized in that: The disturbance variable group is determined in the following way: For each variable in the text out-of-distribution data, a perturbation variable group corresponding to the variable is obtained through a text encoder.
8. An out-of-distribution detection device for an image, characterized in that: include: An acquisition module, used for acquiring an image to be detected; An extraction module, used for extracting the embedding vector of the image to be detected; A detection module is used to determine the detection result of the image to be detected based on the embedding vector, the classification list corresponding to the data within the text distribution, and the classification list corresponding to the data outside the text distribution; wherein the classification list corresponding to the data within the text distribution and the classification list corresponding to the data outside the text distribution are constructed based on the in-distribution data set and the in-distribution data set, respectively.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.