Sample generation method and device, image detection method and model training method

By obtaining the attribute labels and text description features of the sample image set, high-quality training samples are generated, which solves the problems of scarce number of outlier samples and feature differences, improves the robustness and security of the model in out-of-distribution detection, and is suitable for fields such as autonomous driving and smart healthcare.

CN120673189APending Publication Date: 2025-09-19TIANJIN UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510567188.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In existing technologies, the number of outlier samples used for training is scarce, and feature differences lead to poor performance of the model in out-of-distribution detection. The lack of effective semantic information supervision makes it difficult to improve the robustness and security of the model in real scenarios.

Method used

By obtaining the attribute labels of the sample image set, determining the target attribute labels that are similar to the text description features, and using the similarity between visual features and text features to generate training samples, the semantic representation dimension is expanded, high-quality semantic supervision and adversarial training samples are provided, and the performance of the model in out-of-distribution detection is enhanced.

Benefits of technology

The model's performance in out-of-distribution detection is improved, the model's robustness and security are enhanced, and it can better identify samples that are different from the training distribution. It is suitable for safety-critical applications such as autonomous driving and smart healthcare.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673189A_ABST
    Figure CN120673189A_ABST
Patent Text Reader

Abstract

The invention provides a sample generation method, which is applied to the field of machine learning, and comprises the following steps: obtaining a sample image set which comprises a plurality of first sample images and an attribute label of a first object in each first sample image; determining a first text description feature corresponding to each attribute tag; based on each first text description feature, determining a target attribute tag from the plurality of initial attribute tags to obtain a plurality of target attribute tags; based on the similarity between the visual feature of the second sample image and the first text description features corresponding to the multiple attribute tags, and the similarity between the visual feature of the second sample image and the target text description features corresponding to the multiple target attribute tags; determining a detection result representing whether an attribute of a second object of the second sample image exists in a plurality of attribute tags in the sample image set; and generating a training sample based on the detection result corresponding to the second sample image and the second sample image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human machine learning, and in particular to a sample generation method, device, image detection method, and model training method. Background Art

[0002] Out-of-distribution detection is a key technology in machine learning. It aims to identify outlier samples in test data that differ significantly from the distribution of the model's training data. Its core goal is to enhance the robustness and security of the model in real-world scenarios, preventing the model from making high-confidence erroneous predictions about unknown data. Related technologies utilize auxiliary outlier samples, or construct outlier samples based on in-distribution data mining, to aid model training.

[0003] In the process of implementing the concept of the present disclosure, it was found that there are at least the following problems in the relevant technology: the number of outlier samples used in the training process is sparse, and the characteristics of the outlier samples used for training are different from the characteristics of the samples used for out-of-distribution detection. Each outlier sample used for training cannot provide accurate and effective semantic information supervision, resulting in poor performance of the model in out-of-distribution detection. Summary of the Invention

[0004] In view of this, the present disclosure provides a sample generation method, including: obtaining a sample image set, the sample image set including multiple first sample images and attribute labels of the first object in each first sample image; determining a first text description feature corresponding to each attribute label; based on the first text description feature corresponding to each attribute label, determining a target attribute label from multiple initial attribute labels in a knowledge base, respectively, to obtain multiple target attribute labels, the multiple attribute labels and the multiple target attribute labels are all different; based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels, determining a detection result of whether the attribute of the second object used to characterize the second sample image exists in the multiple attribute labels in the sample image set; based on the detection result corresponding to the second sample image and the second sample image, generating a training sample.

[0005] According to an embodiment of the present disclosure, the attribute labels in the sample image set belong to the same attribute category; based on the first text description features corresponding to each attribute label, target attribute labels are respectively determined from multiple initial attribute labels in the knowledge base, and multiple target attribute labels are obtained, including: respectively performing feature extraction on each initial attribute label in the multiple initial attribute labels in the knowledge base to obtain multiple word embedding features; clustering the multiple word embedding features to obtain multiple target clusters; for each first text description feature, based on the distance between the cluster center word embedding features of each of the multiple target clusters and the first text description feature, screening the first target word embedding features from the target clusters to obtain multiple first target word embedding features; and determining the initial attribute labels corresponding to the multiple first target word embedding features as the target attribute labels to obtain multiple target attribute labels.

[0006] According to an embodiment of the present disclosure, based on the first text description features corresponding to each attribute label, target attribute labels are respectively determined from multiple initial attribute labels in the knowledge base, and multiple target attribute labels are obtained, including: for each first sample image: cropping the first sample image to respectively obtain M image segments, where M is an integer greater than or equal to 2; based on the similarity between the sub-visual features corresponding to each of the M image segments and the first text description feature, screening L first target sub-visual features from the M sub-visual features as the first predicted text description features, and screening L second target sub-visual features as the second predicted text description features, where 0 < L < M / 2 and L is an integer, and the first target sub-visual features are different from the second target sub-visual features; according to the similarity between each word embedding feature in the multiple word embedding features and the first predicted text description feature, the similarity between each word embedding feature and the second predicted text description feature, and the similarity between each word embedding feature and the first text description feature, respectively obtaining the first similarity value, the second similarity value, and the third similarity value corresponding to each word embedding feature, where the multiple word embedding features are obtained by respectively performing feature extraction on each initial attribute label in the multiple initial attribute labels in the knowledge base; based on the first similarity value, the second similarity value, and the third similarity value corresponding to each word embedding feature, screening the second target word embedding features from the multiple word embedding features; and determining the initial attribute labels corresponding to the second target word embedding features corresponding to each first sample image as the target attribute labels to obtain multiple target attribute labels.

[0007] According to an embodiment of the present disclosure, based on the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature, a second target word embedding feature is screened from multiple word embedding features, including: based on the weight coefficients corresponding to the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature, the first similarity value, the second similarity value and the third similarity value are weighted and summed to obtain a joint similarity value corresponding to each word embedding feature; based on the joint similarity value corresponding to each word embedding feature, the second target word embedding feature is screened from multiple word embedding features; wherein the weight coefficient corresponding to the first similarity value is greater than the weight coefficient corresponding to the second similarity value, and the weight coefficient corresponding to the third similarity value.

[0008] According to an embodiment of the present disclosure, based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels, a detection result is determined as to whether the attribute of the second object used to characterize the second sample image exists in the multiple attribute labels in the sample image set, including: grouping the multiple target attribute labels to obtain N groups of target attribute label sets, where N is an integer greater than or equal to 2; for each group of target attribute label sets, based on the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels in the target attribute label set, Based on the similarity between the visual features of the second sample image and the first text description features corresponding to the multiple attribute labels in the target attribute label set, a plurality of fourth similarity values ​​are determined; for each set of target attribute label sets, an initial detection probability value is determined based on the similarity values ​​corresponding to the multiple target attribute labels in the target attribute label set and the plurality of fourth similarity values; a target detection probability value is determined based on the initial detection probability value corresponding to each set of target attribute label sets; and based on the target detection probability value, a detection result is determined as to whether the attribute of the second object characterizing the second sample image exists in the multiple attribute labels in the sample image set.

[0009] According to another aspect of the present disclosure, a method for training an image detection model is provided, comprising: inputting a training sample including at least a second sample image into an image detection model, and outputting a training detection result for characterizing whether an attribute of a second object in the second sample image exists in a plurality of attribute labels in a sample image set; training the image detection model based on the training detection result corresponding to the second sample image and the sample label of the second sample image; wherein the sample label of the second sample image is determined based on the detection result obtained by any of the above-mentioned sample generation methods.

[0010] According to another aspect of the present disclosure, an image detection method is provided, comprising: inputting an image to be detected into a pre-trained image detection model, and outputting a target detection result for characterizing whether an attribute of an object in the image to be detected exists in a plurality of attribute labels in a sample image set; wherein the pre-trained image detection model is obtained based on an image detection model training method.

[0011] According to an embodiment of the present disclosure, the image detection method also includes: when it is determined that the attribute of the object in the image to be detected represented by the target detection result does not exist in the multiple attribute labels in the sample image set, based on the similarity between the visual features of the image to be detected and each word embedding feature in the multiple word embedding features, a third target word embedding feature is screened from the multiple word embedding features, wherein the multiple word embedding features are obtained by performing feature extraction on each of the multiple initial attribute labels in the knowledge base respectively; and the attribute label corresponding to the third target word embedding feature is added to the multiple target attribute labels.

[0012] According to an embodiment of the present disclosure, the image detection method also includes: when it is determined that the detection result represents that the attribute of the image to be detected exists in multiple attribute labels in the sample image set, determining the attribute of the object in the image to be detected based on the similarity between the visual features of the image to be detected and each first text description feature.

[0013] According to another aspect of the present disclosure, a sample generation device is provided, characterized in that the device includes: an acquisition module for acquiring a sample image set, the sample image set including multiple first sample images and attribute labels of the first object in each first sample image; a first determination module for determining a first text description feature corresponding to each attribute label; a second determination module for determining a target attribute label from multiple initial attribute labels in a knowledge base based on the first text description feature corresponding to each attribute label, to obtain multiple target attribute labels, wherein the multiple attribute labels and the multiple target attribute labels are different; a third determination module for determining a detection result of whether the attribute of the second object used to characterize the second sample image exists in the multiple attribute labels in the sample image set based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels; a fourth determination module for generating a training sample based on the detection result corresponding to the second sample image and the second sample image.

[0014] According to an embodiment of the present disclosure, by obtaining first text description features corresponding to multiple attribute labels in a sample image set, the dimension of the image semantic representation of the first sample image can be expanded. According to the first text description features corresponding to each attribute label after the dimension of the semantic representation is expanded, target attribute labels different from the multiple attribute labels are determined from the multiple initial attribute labels in the knowledge base, and multiple target attribute labels are obtained to enhance the semantic difference between the multiple target attribute labels and the multiple attribute labels. Based on the similarity between the visual features of the second sample image and the hint features within the distribution, and the similarity between the visual features of the second sample image and the hint features outside the distribution, the detection result of whether the attribute of the second object in the second sample image belongs to the multiple attribute labels is determined, and a training sample is generated according to the detection result and the second sample image, thereby providing a plurality of training samples with higher quality semantic supervision and more adversarial for model training, which are used to improve the performance of the image detection model training in the out-of-distribution detection process. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0016] Figure 1 The application scenario diagram of the sample generation method and device according to the embodiments of the present disclosure is schematically shown.

[0017] Figure 2 A flow chart of a sample generation method according to an embodiment of the present disclosure is shown.

[0018] Figure 3 A schematic diagram of determining multiple target attribute labels according to an embodiment of the present disclosure is shown.

[0019] Figure 4 A schematic diagram of image segments corresponding to the first predicted text description feature and the second predicted text description feature according to an embodiment of the present disclosure is shown.

[0020] Figure 5 A schematic diagram of determining multiple target attribute labels according to another embodiment of the present disclosure is shown.

[0021] Figure 6 A schematic diagram of an image detection method according to an embodiment of the present disclosure is shown.

[0022] Figure 7 The structural block diagram of the sample generating device according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0023] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0024] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0027] In implementing the concepts of this disclosure, out-of-distribution detection enables models to accurately identify samples that differ from the training distribution during deployment, thereby improving the model's performance in identifying out-of-distribution samples. This is crucial for safety-critical applications such as autonomous driving and smart healthcare, as it minimizes unreliable decision outcomes.

[0028] Traditional out-of-distribution detection algorithms are typically trained or fine-tuned based on the entire training data, including those based on post-hoc reasoning and those based on regularization during training. These traditional algorithms are often driven by a single modality and fail to leverage information from multimodal data representations, resulting in poor generalization capabilities.

[0029] As visual language models pre-trained on large-scale datasets demonstrate strong zero-shot generalization capabilities in multiple downstream tasks, some researchers have explored leveraging the powerful multimodal representation capabilities of visual language models for out-of-distribution detection tasks. These methods, in the context of few-shot or zero-shot training, seek supervisory information for model learning by mining high-value outlier representations from within-distribution training images, leveraging pre-defined within-distribution cue representations, or retrieving relevant text from external corpora. Their performance has even surpassed traditional algorithms that use the entire training set.

[0030] Because the supervisory information utilized by these methods mostly comes from a small amount of in-distribution data (few-shot setting), or even only from known in-distribution text labels (zero-shot setting), the overall distribution of this supervisory information lacks diversity at the macro level, and the learned prompts have limitations in the semantic space. Specifically, the supervisory information distribution constructed from a small number of in-distribution samples cannot fully cover the out-of-distribution samples with complex and diverse semantic characteristics that may appear during the testing phase. Furthermore, the prompts learned based on in-distribution text labels, due to their over-reliance on known semantics, may deviate significantly from the semantics of actual out-of-distribution samples, making it difficult to capture the diverse out-of-distribution representations in real test datasets. Such limitations affect the model's ability to detect out-of-distribution samples, ultimately making it difficult for the learned prompts to have strong generalization capabilities.

[0031] In addition, conventional supervision information is constructed based only on image or text representations in the training phase, without efficiently utilizing the information of potential out-of-distribution samples in the testing phase. The difference between supervision information and actual out-of-distribution test samples leads to insufficient accuracy of the knowledge distribution contained in the supervision information. The above two problems make it difficult to further improve the performance of detecting out-of-distribution samples when the model is deployed in open scenarios.

[0032] Based on this, an embodiment of the present disclosure provides a sample generation method. The method includes: obtaining a sample image set, the sample image set including multiple first sample images and an attribute label of a first object in each first sample image; determining a first text description feature corresponding to each attribute label; based on the first text description feature corresponding to each attribute label, determining a target attribute label from multiple initial attribute labels in a knowledge base, respectively, to obtain multiple target attribute labels, wherein the multiple attribute labels and the multiple target attribute labels are all different; based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels, determining a detection result of whether the attribute of the second object of the second sample image is present in the multiple attribute labels in the sample image set; and generating a training sample based on the detection result corresponding to the second sample image and the second sample image.

[0033] Figure 1 The application scenario diagram of the sample generation method and device according to the embodiments of the present disclosure is schematically shown.

[0034] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0035] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0036] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0037] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0038] It should be noted that the sample generation method provided in the embodiments of the present disclosure can generally be executed by the server 105. The sample generation device can generally be configured on the server 105. The sample generation method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the sample generation device can generally be configured in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0040] The following will be passed Figures 2 to 5 The sample generation method of the embodiment of the present disclosure is described in detail.

[0041] Figure 2 A flow chart of a sample generation method according to an embodiment of the present disclosure is shown.

[0042] like Figure 2 As shown, the large model deployment method 200 includes: operations S210 to S250.

[0043] In operation S210 , a sample image set is acquired, where the sample image set includes a plurality of first sample images and an attribute label of a first object in each first sample image.

[0044] In operation S220 , a first text description feature corresponding to each attribute tag is determined.

[0045] In operation S230 , based on the first text description feature corresponding to each attribute tag, target attribute tags are determined from the multiple initial attribute tags in the knowledge base to obtain multiple target attribute tags, where the multiple attribute tags and the multiple target attribute tags are different.

[0046] In operation S240, based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute tags, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute tags, a detection result is determined as to whether the attribute of the second object used to characterize the second sample image exists in the multiple attribute tags in the sample image set.

[0047] In operation S250 , a training sample is generated based on the detection result corresponding to the second sample image and the second sample image.

[0048] According to an embodiment of the present disclosure, a sample image set can be obtained from a plurality of open-source training data. The sample image set is used as training samples for training an image detection model to complete target tasks such as classification, regression, and detection. A first object in a first sample image represents an object to be identified.

[0049] For example, for a picture in which a water cup is the focus, the first sample image is the picture, and the first object is the water cup in the picture. The attribute label of the first object can be the text "water cup" for identifying the category of the first object.

[0050] According to an embodiment of the present disclosure, the first text description feature represents an in-distribution hint feature related to the attribute label. The first text description feature may be a feature generated based on text describing the first sample image.

[0051] For example, the description text for the attribute tag "vehicle" can be "This is a red vehicle, taken in an outdoor environment"; the text description for the attribute tag "dog" can be "This is a photo of an old yellow dog."

[0052] According to an embodiment of the present disclosure, the first text description feature may determine the first text description feature corresponding to each attribute tag from a mapping table including multiple attribute tags and multiple first text description features according to the attribute tag.

[0053] According to an embodiment of the present disclosure, the first text description feature can also be determined by the following method: a first sample image and a prompt word that prompts the large model to describe the first sample image are input into the large model, the large model recognizes the first sample image, and outputs text describing the first sample image. The text output by the large model is input into a text encoder to obtain the first text description feature.

[0054] According to an embodiment of the present disclosure, the knowledge base can be an external corpus containing a large amount of text. For example, "red," "blue," "steel," "plastic," "triangle," "circle," "ear," etc. The target attribute label represents out-of-distribution hint information that has a low correlation with multiple attribute labels.

[0055] According to an embodiment of the present disclosure, multiple initial attribute labels of a knowledge base are encoded using a preset text encoder to obtain features corresponding to the multiple initial attribute labels, thereby determining the similarity of each initial attribute label with the first text description feature, and selecting initial attribute labels with lower similarity from the multiple initial attribute features as target attribute labels to obtain multiple target attribute labels. The multiple attribute labels and the multiple target attribute labels are different.

[0056] For example, when the multiple initial attribute labels of the knowledge base are "vehicle", "triangle", and "ear", for the first text description feature corresponding to "This is a red vehicle, photographed in an outdoor environment", the similarity between "vehicle", "triangle", and "ear" and the first text description feature can be calculated respectively, thereby determining the target attribute labels "triangle" and "ear" with lower similarity to the first text description feature.

[0057] According to an embodiment of the present disclosure, the second sample image is an image with a significantly different data distribution from that used when training the image detection model. For example, when the image detection model is used to classify cats and dogs, the samples used in the training process are images of cats or dogs. The visual features of the second sample image can be obtained by inputting the second sample image into a visual encoder.

[0058] According to an embodiment of the present disclosure, the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels is used to characterize the probability that the second sample image belongs to an in-distribution sample. An in-distribution sample characterizes that the attributes of the second object corresponding to the second sample image belong to the multiple attribute labels. For example, if the multiple attribute labels are "cat" and "dog", and the attribute of the second object is "cat" or "dog", then the second sample image belongs to the in-distribution sample.

[0059] According to an embodiment of the present disclosure, the target text description feature represents a feature encoding matrix of the target attribute label, and the target text description feature can be obtained by inputting the target attribute label into a text encoder. The similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels is used to represent the probability that the second sample image belongs to an out-of-distribution sample. The in-distribution sample represents that the attribute of the second object corresponding to the second sample image does not belong to multiple attribute labels. For example, the multiple attribute labels are "cat" and "dog", the attribute of the second object is "bird", and the second sample image belongs to an out-of-distribution sample.

[0060] According to an embodiment of the present disclosure, the similarity between the visual features of the second sample image and the first text description features of each attribute label, as well as the similarity between the visual features of the second sample image and the first text description features of each attribute label, can be determined by calculating the cosine similarity between the features.

[0061] According to an embodiment of the present disclosure, based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels, the label with the highest similarity to the visual features of the second sample image is determined to determine the detection result according to whether the label is an attribute label.

[0062] According to an embodiment of the present disclosure, when the detection result characterizes that the attributes of the second object of the second sample image exist in multiple attribute labels, the attributes of the second sample image are used as sample labels to generate training samples for training the image detection model to complete target tasks such as classification, regression, and detection, that is, training samples within the distribution.

[0063] According to an embodiment of the present disclosure, when the detection result characterizes that the non-attribute of the second object of the second sample image exists in multiple attribute labels, the attribute of the second sample image is used as a sample label to generate training samples for improving the robustness, security and generalization ability of the image detection model, that is, out-of-distribution training samples or outlier samples.

[0064] According to an embodiment of the present disclosure, when the detection result characterizes that the non-attribute of the second object of the second sample image exists in multiple attribute labels, "unknown", "other", etc. can also be set as the sample label of the second sample image to generate out-of-distribution training samples.

[0065] According to an embodiment of the present disclosure, in order to improve the quality of multiple training samples, multiple out-of-distribution training samples may be determined based on the second sample image.

[0066] According to an embodiment of the present disclosure, by obtaining first text description features corresponding to multiple attribute labels in a sample image set, the dimension of the image semantic representation of the first sample image can be expanded. According to the first text description features corresponding to each attribute label after the dimension of the semantic representation is expanded, target attribute labels different from the multiple attribute labels are determined from the multiple initial attribute labels in the knowledge base, and multiple target attribute labels are obtained to enhance the semantic difference between the multiple target attribute labels and the multiple attribute labels. Based on the similarity between the visual features of the second sample image and the hint features within the distribution, and the similarity between the visual features of the second sample image and the hint features outside the distribution, the detection result of whether the attribute of the second object in the second sample image belongs to the multiple attribute labels is determined, and a training sample is generated according to the detection result and the second sample image, thereby providing a plurality of training samples with higher quality semantic supervision and more adversarial for model training, which are used to improve the performance of the image detection model training in the out-of-distribution detection process.

[0067] According to an embodiment of the present disclosure, the attribute labels in the sample image set belong to the same attribute category; based on the first text description feature corresponding to each attribute label, the target attribute label is determined from the multiple initial attribute labels in the knowledge base respectively, and multiple target attribute labels are obtained, including: feature extraction is performed on each of the multiple initial attribute labels in the knowledge base respectively to obtain multiple word embedding features; the multiple word embedding features are clustered to obtain multiple target clusters; for each of the first text description features, based on the distance between the cluster center word embedding feature of each of the multiple target clusters and the first text description feature, the first target word embedding feature is screened from the target cluster to obtain multiple first target word embedding features; the initial attribute label corresponding to each of the multiple first target word embedding features is determined as the target attribute label to obtain multiple target attribute labels.

[0068] According to an embodiment of the present disclosure, the attribute labels in the sample image set belong to the same attribute category. For example, the sample image set is used to train the color classification capability of the image detection model, then multiple attribute labels in the sample image set belong to the color category, and the multiple attribute labels of the color category may include: "red", "blue", etc.

[0069] For another example, the sample image set is used to train the image detection model's ability to recognize objects in images. Then, multiple attribute labels in the sample image set may belong to the object recognition category, and the multiple attribute labels of the object recognition category may include: "horse", "car", "person" and other labels.

[0070] According to an embodiment of the present disclosure, a CLIP (Contrastive Language-Image Pre-training) text encoder is used to extract a text representation of each of multiple initial attribute labels in a knowledge base, and normalize the text to obtain multiple word embedding features; the word embedding features represent a matrix representation of the initial attribute labels.

[0071] According to the embodiments of the present disclosure, since the external corpus contains a massive amount of text, directly clustering it may affect the efficiency of knowledge distribution construction, and thus affect the efficiency of prompt enhancement in the training stage. A small batch K-means algorithm can be used to cluster multiple word embedding features in batches to obtain multiple target clusters to improve clustering efficiency and thus improve the efficiency of generating training samples.

[0072] According to an embodiment of the present disclosure, an initial attribute label with lower similarity is determined from multiple initial attribute labels in the knowledge base including external knowledge based only on the similarity between the initial attribute label and the first text description feature. This only considers the accuracy of determining multiple target attribute labels, but lacks semantic diversity to a certain extent. It is impossible for multiple target attribute labels to cover the various semantic information in the representation space as comprehensively as possible. This defect is more obvious when the number of attribute labels is small. There will be too many texts with accuracy but similar semantics in the multiple target attribute labels determined, and multiple target attribute labels that meet the diversity of semantic types cannot be obtained, which makes it impossible for the model to generate a comprehensive and tight decision boundary, resulting in poor performance of the model in out-of-distribution detection.

[0073] Based on this, for each first text description feature, the similarity between the cluster center word embedding feature of each of the multiple target clusters and the first text description feature is calculated to determine a target cluster with lower similarity to the first text description feature from the multiple target clusters, thereby obtaining multiple target clusters with lower similarity. P word embedding features with lower similarity to the first text description feature can be screened from the multiple target clusters with lower similarity to obtain multiple first target word embedding features, thereby reducing semantically similar texts, where the value of P can be preset based on expert experience.

[0074] According to an embodiment of the present disclosure, multiple first target word embedding features are used to characterize diverse out-of-distribution hint features. Determine the P embedding features with lower similarity as the i-th target cluster with lower similarity The first target word embedding feature , the method can be expressed by formula (1):

[0075] (1);

[0076] in, represents the i-th target cluster The similarity between the word embedding feature and the first text description feature.

[0077] According to an embodiment of the present disclosure, the initial attribute labels corresponding to the respective first target word embedding features are determined as target attribute labels, thereby obtaining multiple target attribute labels.

[0078] According to the embodiments of the present disclosure, if the retrieval and construction of knowledge distribution is performed only based on the similarity between the initial attribute label and the attribute label, the knowledge distribution generated by this construction strategy that only considers accuracy as the principle lacks semantic diversity and cannot fully cover the various semantic information in the representation space. This defect is more obvious when the number of attribute labels is small.

[0079] According to an embodiment of the present disclosure, feature extraction is performed on each of the multiple initial attribute labels of the knowledge base to obtain multiple word embedding features. The multiple word embedding features are clustered to obtain multiple target clusters. For each first text description feature, based on the distance between the cluster center word embedding feature of each of the multiple target clusters and the first text description feature, the first target word embedding feature is screened from the target cluster to obtain multiple first target word embedding features, so as to reduce the amount of computation used for distance calculation and increase the diversity of the semantic types of the multiple first target word embedding features. The initial attribute labels corresponding to each of the multiple first target word embedding features are determined as target attribute labels to obtain multiple target attribute labels. The multiple target attribute labels are determined based on the principle of accuracy and diversity to increase the diversity of the semantic types of the multiple target attribute labels.

[0080] Figure 3 A schematic diagram of determining multiple target attribute labels according to an embodiment of the present disclosure is shown.

[0081] like Figure 3 As shown, multiple initial attribute labels in the knowledge base are clustered to obtain target clusters 1, 2, 3, and 4. For each of the first text description features of multiple attribute labels in the sample image set, for example, the first text description feature of "cow," the distance between the cluster center word embedding feature and the first text description feature of each target cluster is determined, and the first target word embedding features are filtered from the target clusters to obtain multiple first target word embedding features. Among them, the initial attribute label corresponding to the cluster center word embedding feature of target cluster 1 is "airplane," the initial attribute label corresponding to the cluster center word embedding feature of target cluster 2 is "jeans," the initial attribute label corresponding to the cluster center word embedding feature of target cluster 3 is "sophora tree," and the initial attribute label corresponding to the cluster center word embedding feature of target cluster 4 is "wood."

[0082] For target cluster 4, the cosine similarity between the first text description feature of "cat" in the target attribute label and the cluster center embedded word feature corresponding to the initial attribute label "wood" is determined. If the cosine similarity is greater than the preset threshold, the cluster center embedded word feature of target cluster 4 is used as the first target word embedded feature.

[0083] Based on the same method, the first target word embedding features are filtered from target cluster 1, target cluster 2, and target cluster 3 to obtain multiple first target word embedding features.

[0084] The initial attribute labels corresponding to the plurality of first target word embedding features are determined as target attribute labels to obtain a plurality of target attribute labels.

[0085] According to an embodiment of the present disclosure, based on the first text description features corresponding to each attribute label, target attribute labels are determined from multiple initial attribute labels in the knowledge base respectively, obtaining multiple target attribute labels, including: for each first sample image: cropping the first sample image to obtain M image segments respectively, where M is an integer greater than or equal to 2; based on the similarity between the sub-visual features corresponding to each of the M image segments and the first text description features, screening L first target sub-visual features from the M sub-visual features as the first predicted text description features, and screening L second target sub-visual features as the second predicted text description features, 0 < L < M / 2, and L is an integer, and the first target sub-visual features are different from the second target sub-visual features; according to the similarity between each word embedding feature in the multiple word embedding features and the first predicted text description features, and the similarity between each word embedding feature and the second predicted text description features, and the similarity between each word embedding feature and the first text description features, obtaining the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature respectively, where the multiple word embedding features are obtained by performing feature extraction on each initial attribute label in the multiple initial attribute labels of the knowledge base; screening second target word embedding features from the multiple word embedding features based on the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature; determining the initial attribute label corresponding to the second target word embedding feature corresponding to each first sample image as the target attribute label, obtaining multiple target attribute labels.

[0086] According to an embodiment of the present disclosure, the first sample image is cropped to obtain M image segments respectively, where M is an integer greater than or equal to 2. The sub-visual features corresponding to each of the M image segments can be determined using the pre-trained visual encoder of the CLIP model. Cropping the image into M segments enables the visual encoder to capture features from both global and local perspectives: details that may be overlooked by the global features, such as the shape, texture, and color of objects, can be explicitly retained through the sub-visual features.

[0087] According to an embodiment of the present disclosure, the cosine similarity between the sub-visual features corresponding to each of the M image segments and the first text description features is calculated, and L sub-visual features with lower cosine similarities are screened from the M sub-visual features to obtain L first target sub-visual features, and the L first target sub-visual features are used as the first predicted text description features.

[0088] According to an embodiment of the present disclosure, L sub-visual features with higher cosine similarities are screened from the M sub-visual features to obtain L second target sub-visual features, and the L second target sub-visual features are used as the second predicted text description features.

[0089] According to an embodiment of the present disclosure, the first predicted text description feature representation is similar to the first text description feature and frequently appears in the feature neighborhood of the first object, but is not a feature of the first object. For example, the first object in the first sample image is a dog. After cropping the first sample image, M image segments are obtained. The sub-visual features of some of these image segments represent background-related content, but are not characteristics of the dog. Background-related content, such as lawns and frisbees, are image content that frequently appears with the first object, the dog.

[0090] According to an embodiment of the present disclosure, the second target sub-visual feature represents a feature that is highly correlated with the feature representing the first object. For example, if the first object in the first sample image is a goldfish, then the second target sub-visual feature in the M cropped image segments may be features representing the goldfish's fins, scales, eyes, and other parts.

[0091] Figure 4 A schematic diagram of image segments corresponding to the first predicted text description feature and the second predicted text description feature according to an embodiment of the present disclosure is shown.

[0092] like Figure 4 As shown, the first sample image 400 is cropped to obtain multiple image segments, among which the sub-visual features corresponding to the image segment 401 have fewer features related to the first object "cat", and are mainly used to characterize the features on the sofa where the cat is. The similarity between the sub-visual features corresponding to the image segment 401 and the first text description features will be low, and the sub-visual features corresponding to the image segment 401 can be the first predicted text description features.

[0093] Image segment 402 includes features that are relatively relevant to the first object "cat". The similarity between the sub-visual feature corresponding to image segment 402 and the first text description feature is relatively high, and the sub-visual feature corresponding to image segment 402 can be used as the second predicted text description feature.

[0094] According to an embodiment of the present disclosure, a text encoder of CLIP is used to extract a text representation of each of multiple initial attribute labels in a knowledge base, and normalize the text to obtain multiple word embedding features; the word embedding feature is a matrix representation of the initial attribute label.

[0095] According to an embodiment of the present disclosure, the first similarity value, the second similarity value, and the third similarity value may be determined by calculating the cosine similarity between the features.

[0096] According to an embodiment of the present disclosure, the first similarity value can be used to represent the degree of association between the initial attribute label and the image segment used to represent the non-first object feature. For example, "sofa" and Figure 4The correlation between the image segment 401, "dog" and Figure 4 The correlation of the image segment 401 in FIG.

[0097] According to an embodiment of the present disclosure, the second similarity can be used to represent: the degree of association between the initial attribute label and the image segment used to represent the first object feature. For example, "cat" and Figure 4 The correlation between the image segment 402, "fish" and Figure 4 The correlation of the image segment 402 in FIG.

[0098] According to an embodiment of the present disclosure, the third similarity can be used to represent the degree of association between the initial attribute label and the text content used to represent the first object feature, for example, the relevance between "cat" and "a yellow cat", or the relevance between "fish" and "a yellow cat".

[0099] According to an embodiment of the present disclosure, based on the first similarity value, the second similarity value and the third similarity value corresponding to the word embedding feature, a word embedding feature with a higher first similarity value and lower second and third similarity values ​​is selected from multiple word embedding features as the second target word embedding feature.

[0100] According to an embodiment of the present disclosure, the initial attribute label corresponding to the second target word embedding feature obtained corresponding to each first sample image is determined as the target attribute label, and a plurality of target attribute labels are obtained.

[0101] According to an embodiment of the present disclosure, a first sample image is cropped to obtain M image segments and sub-visual features corresponding to the M image segments, so as to retain the fine-grained image information in the first sample image and improve the feature expression ability of the sub-visual features in tiny details. Based on the similarity between the sub-visual features corresponding to each image segment and the first text description feature, a first predicted text description feature with a higher similarity to the first text description feature and a second predicted text description feature with a lower similarity to the first text description feature are obtained. Based on the similarity values ​​of the word embedding feature and the first predicted text description feature, the second predicted text description feature and the first text description feature, i.e., the first similarity value, the second similarity value and the third similarity value, a second target word embedding feature is screened from the multiple word embedding features to reduce the overlap of multiple initial attribute labels with the first description, improve the accuracy and effectiveness of the target attribute label, and thus provide training samples carrying high-value supervisory signals for the training process.

[0102] Figure 5 A schematic diagram of determining multiple target attribute labels according to another embodiment of the present disclosure is shown.

[0103] like Figure 5As shown, the first sample image includes outlier representations, which are generally background features unrelated to the first object. The first sample image also includes in-distribution representations, which are generally features that represent the first object. The first sample image is cropped to obtain M image segments to separate the in-distribution representations from the outlier representations.

[0104] The multiple initial attribute labels in the knowledge base are divided into groups, resulting in Group 1: triangle, square, circle, and bar. Group 2: silk, steel, plastic, and wood. Group 3: arms, ears, legs, and tail. Group 4: red, blue, and gray. The cluster center feature of Group 1 can be a bar; the cluster center feature of Group 2 can be a tail; the cluster center feature of Group 3 can be gray; and the cluster center feature of Group 4 can be wood.

[0105] Based on the similarity between the sub-visual features corresponding to each image segment in the M image segments and the first text description features, L first target sub-visual features are screened from the M sub-visual features as the first predicted text description features, i.e., outlier features, and L second target sub-visual features are screened as the second predicted text description features, i.e., in-distribution features.

[0106] According to the similarity between each word embedding feature in the plurality of word embedding features and the first predicted text description feature, the similarity between each word embedding feature and the second predicted text description feature, and the similarity between each word embedding feature and the first text description feature, respectively obtaining a first similarity value, a second similarity value, and a third similarity value corresponding to each word embedding feature, wherein the plurality of word embedding features are obtained by performing feature extraction on each of the plurality of initial attribute labels in the knowledge base;

[0107] Based on the first, second, and third similarity values ​​corresponding to each word embedding feature, a second target word embedding feature is selected from the multiple word embedding features to obtain a second target word embedding feature that maximizes similarity with the outlier feature and minimizes similarity with the in-distribution representation and the multiple attribute labels. Thus, the initial attribute label corresponding to the second target word embedding feature obtained for each first sample image is determined as the target attribute label, resulting in multiple target attribute labels.

[0108] for Figure 5 For the first sample image "cat", based on the first similarity value, the second similarity value and the third similarity value, a second target word embedding feature is determined that maximizes the similarity with the background part and minimizes the similarity with the first object and the attribute label "cat".

[0109] According to an embodiment of the present disclosure, based on the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature, a second target word embedding feature is screened from multiple word embedding features, including: based on the weight coefficients corresponding to the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature, the first similarity value, the second similarity value and the third similarity value are weighted and summed to obtain a joint similarity value corresponding to each word embedding feature; based on the joint similarity value corresponding to each word embedding feature, the second target word embedding feature is screened from multiple word embedding features; wherein the weight coefficient corresponding to the first similarity value is greater than the weight coefficient corresponding to the second similarity value, and the weight coefficient corresponding to the third similarity value.

[0110] According to an embodiment of the present disclosure, a weighted sum is performed on the first similarity value, the second similarity value, and the third similarity value corresponding to each word embedding feature to obtain a total similarity value of each word embedding feature, and the top X word embedding features with higher joint similarity values ​​are selected from multiple word embedding features as the second target word embedding features.

[0111] According to an embodiment of the present disclosure, the weight coefficient corresponding to the second similarity value is greater than 0, and the weight coefficient corresponding to the first similarity value and the weight coefficient corresponding to the third similarity value are less than 0.

[0112] According to an embodiment of the present disclosure, the weight coefficient corresponding to the first similarity value may be set to be much larger than the weight coefficients corresponding to the second similarity value and the weight coefficients corresponding to the third similarity value. For example, the weight coefficient corresponding to the first similarity value is 0.99, the weight coefficient corresponding to the second similarity value is 0.005, and the weight coefficient corresponding to the third similarity value is 0.005.

[0113] According to an embodiment of the present disclosure, the weight coefficient corresponding to the first similarity value for determining the joint similarity is set to be greater than the weight coefficient corresponding to the second similarity value and the weight coefficient corresponding to the third similarity value. By increasing the similarity ratio between each word embedding feature in the multiple word embedding features in the joint similarity value and the first predicted text description feature, the overlap of multiple initial attribute labels with the first description is reduced, and the accuracy and effectiveness of the target attribute label are improved.

[0114] According to an embodiment of the present disclosure, multiple target attribute labels are determined The formula can be expressed by formula (2):

[0115] (2);

[0116] in, Represents multiple word embedding features, The joint similarity value representing the word embedding features, Indicates the number of multiple target attribute labels.

[0117] According to an embodiment of the present disclosure, based on the similarity between the visual features of the second sample image and the first text description features corresponding to each of the multiple attribute labels, and the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels, a detection result is determined as to whether the attribute of the second object used to characterize the second sample image exists in the multiple attribute labels in the sample image set, including: grouping the multiple target attribute labels to obtain N groups of target attribute label sets, where N is an integer greater than or equal to 2; for each group of target attribute label sets, based on the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels in the target attribute label set, According to the similarity between the visual features of the second sample image and the first text description features corresponding to the multiple attribute labels, a plurality of similarity values ​​are determined; according to the similarity between the visual features of the second sample image and the first text description features corresponding to the multiple attribute labels, a plurality of fourth similarity values ​​are obtained; for each group of target attribute label sets, an initial detection probability value is determined according to the similarity values ​​corresponding to the multiple target attribute labels in the target attribute label set and the plurality of fourth similarity values; according to the initial detection probability value corresponding to each group of target attribute label sets, a target detection probability value is determined; based on the target detection probability value, a detection result is determined as to whether the attribute of the second object used to characterize the second sample image exists in the multiple attribute labels in the sample image set.

[0118] According to an embodiment of the present disclosure, multiple target attribute labels may be grouped according to the prompt word grouping integration strategy used in NegLabel to obtain N sets of target attribute label sets, where N is an integer greater than or equal to 2.

[0119] For example, multiple predefined labels, such as "red" and "stripes", are divided into N groups according to the correlation between the semantic multiple target attribute labels and the predefined labels, such as group 1 color: {red, blue, green}; group 2 texture: {stripes, plaid, solid color}.

[0120] According to an embodiment of the present disclosure, for each group of target attribute label sets, the cosine similarity between the visual features of the second sample image and the target text description features corresponding to each target attribute label is calculated, and the similarity values ​​corresponding to each of the multiple target attribute labels in the target attribute label set are determined.

[0121] According to an embodiment of the present disclosure, the mean similarity value is used to characterize the correlation between the attributes of the second sample image and the target attribute label set. For example, the correlation between "yellow" and the color group 1: {red, blue, green}. For example, the correlation between "yellow" and the texture group 2: {stripes, plaid, solid color}.

[0122] According to an embodiment of the present disclosure, for each attribute label, the cosine similarity between the visual feature of the second sample image and the first text description feature corresponding to the attribute label is calculated to determine a fourth similarity value corresponding to each attribute label.

[0123] According to an embodiment of the present disclosure, the fourth similarity value represents the correlation between the attribute of the second sample image and the attribute label, for example, the correlation between "dog" and "yellow cat"; for example, the correlation between "cat" and "yellow cat".

[0124] According to an embodiment of the present disclosure, for each set of target attribute label sets, the similarity values ​​corresponding to each of the multiple target attribute labels in the target attribute label set are summed to obtain a sum of similarities corresponding to the target attribute label set. The sum of the similarities corresponding to the target attribute label set and the sum of the multiple fourth similarity values ​​is added as the denominator, and the sum of the multiple fourth similarity values ​​is used as the numerator to calculate the initial detection probability value of the second sample image for the target attribute label set, thereby determining the initial detection probability values ​​corresponding to the second sample image and the multiple target attribute label sets respectively.

[0125] According to an embodiment of the present disclosure, the average of the initial detection probability values ​​corresponding to the second sample image and multiple target attribute label sets is calculated, and the average of the N initial detection probability values ​​is determined as the target detection probability value of the second sample image, wherein the target detection probability value represents the probability that the attributes of the second object in the second sample image belong to multiple attribute labels.

[0126] According to an embodiment of the present disclosure, when the target detection probability value is greater than a preset value, it is determined that the attribute used to characterize the second object in the second sample image is present in the plurality of attribute labels in the sample image set; and when the target detection probability value is less than the preset value, it is determined that the attribute used to characterize the second object in the second sample image is not present in the plurality of attribute labels in the sample image set. The preset value may be, for example, 0.5 or 0.6.

[0127] According to an embodiment of the present disclosure, multiple target attribute labels are grouped to obtain N groups of target attribute label sets. According to the initial detection probability values ​​corresponding to the second sample image and the multiple target attribute label sets, the probability that the attributes of the second object in the second sample image belong to the multiple attribute labels, that is, the target detection probability value, is calculated, so as to determine the detection result of whether the attributes of the second object used to characterize the second sample image exist in the multiple attribute labels in the sample image set. This method of calculating the target detection probability value in units of the target attribute label set to determine the detection result reduces the influence of a single target attribute label on determining the detection result, thereby reducing the probability of inaccurate detection results due to semantic overlap between the target attribute label and the attribute label, and improving the accuracy of the detection result.

[0128] According to another aspect of the present disclosure, a method for training an image detection model is provided, comprising: inputting a training sample comprising at least a second sample image into an image detection model, and outputting a training detection result for characterizing whether an attribute of a second object in the second sample image exists in a plurality of attribute labels in a sample image set; training the image detection model based on the training detection result corresponding to the second sample image and the sample label of the second sample image; wherein the sample label of the second sample image is determined by the detection result obtained by the sample generation method of any of the above-mentioned embodiments.

[0129] According to an embodiment of the present disclosure, the image detection model may be trained based on a sample image set.

[0130] For example, the sample image set includes multiple first sample images with attribute labels of "cat" and "dog." A second sample image with the attribute of the second object being "cat" is input into the image detection model, and the image detection model outputs a training detection result of "yes."

[0131] For example, the sample image set includes multiple first sample images with attribute labels of "cat" and "dog." A second sample image with the attribute of the second object being "car" is input into the image detection model, and the image detection model outputs a training detection result of "no."

[0132] According to an embodiment of the present disclosure, in a case where the attribute of the second object exists in a plurality of attribute labels in the sample image set, the sample label of the second sample image may be the attribute of the second object.

[0133] According to an embodiment of the present disclosure, when the attribute of the second object does not exist in the plurality of attribute labels in the sample image set, the sample label of the second sample image may be the attribute of the second object.

[0134] According to an embodiment of the present disclosure, in a case where the attribute of the second object does not exist in the multiple attribute labels in the sample image set, the sample label of the second sample image may also be “other”.

[0135] According to an embodiment of the present disclosure, a second sample image is input into an image detection model. The model then performs operations such as feature extraction and object recognition on the image based on its own network structure and parameters, outputting a preliminary recognition result. This preliminary recognition result is then carefully compared with the sample label. By calculating the error between the two, such as classification error and positioning error, the error signal is propagated back to each network layer of the model using a backpropagation algorithm, enabling targeted adjustments to the parameters in the image detection model.

[0136] According to an embodiment of the present disclosure, an image detection model is trained using the second sample image and sample label obtained by the sample generation method of any of the above embodiments, providing the image detection model training with multiple training samples with higher-quality semantic supervision and greater adversarial nature, so as to improve the performance of the image detection model training in the out-of-distribution detection process.

[0137] According to another aspect of the present disclosure, an image detection method is provided, comprising: inputting an image to be detected into a pre-trained image detection model, and outputting a target detection result for characterizing whether an attribute of an object in the image to be detected exists in a plurality of attribute labels in a sample image set; wherein the pre-trained image detection model is obtained by the above-mentioned image detection model training method.

[0138] According to an embodiment of the present disclosure, the image to be detected may be an image used for performing out-of-distribution detection on an image detection model, and a target detection result is determined using a pre-trained image detection model.

[0139] For example, multiple attribute labels include: "cat" and "dog". The image to be detected with the attribute of the object being "cat" is input into the image detection model, and the pre-trained image detection model outputs the target detection result representing that the attributes of the object in the image to be detected exist in multiple attribute labels in the sample image set.

[0140] For example, multiple attribute labels include: "cat" and "dog", and the image to be detected whose attribute is "tiger" is input into the image detection model. The pre-trained image detection model outputs a target detection result representing that the attribute of the object in the image to be detected does not exist in the multiple attribute labels in the sample image set.

[0141] According to an embodiment of the present disclosure, out-of-distribution detection is performed on a pre-trained image detection model using an image to be detected to determine a target detection result, thereby determining the robustness of the image detection model.

[0142] According to an embodiment of the present disclosure, the image detection method also includes: when it is determined that the attribute of the object in the image to be detected represented by the target detection result does not exist in the multiple attribute labels in the sample image set, based on the similarity between the visual features of the image to be detected and each word embedding feature in the multiple word embedding features, a third target word embedding feature is screened from the multiple word embedding features, wherein the multiple word embedding features are obtained by performing feature extraction on each of the multiple initial attribute labels in the knowledge base respectively; and the attribute label corresponding to the third target word embedding feature is added to the multiple target attribute labels.

[0143] According to an embodiment of the present disclosure, the visual features of the image to be detected may be obtained by inputting the image to be detected into a pre-trained visual encoder.

[0144] According to an embodiment of the present disclosure, there is a large amount of redundancy in the multiple initial attribute labels of the knowledge base. Based on the cosine similarity between the visual features of the image to be detected and each word embedding feature in the multiple word embedding features, the word embedding feature with higher cosine similarity is selected from the multiple word embedding features as the third target word embedding feature.

[0145] According to an embodiment of the present disclosure, the attribute label corresponding to the third target word embedding feature is added to multiple target attribute labels to expand the number of target attribute labels, and the image to be detected corresponding to the attribute label can be used as a training sample to improve the robustness of the model.

[0146] Figure 6 A schematic diagram of an image detection method according to an embodiment of the present disclosure is shown.

[0147] like Figure 6 As shown, multiple word embedding features are obtained by respectively extracting features from each of the multiple initial attribute labels in the knowledge base. When it is determined that the attribute of the object in the target detection result representing the object in the image to be detected 1 does not exist in the multiple attribute labels in the sample image set, based on the similarity between the visual features of the image to be detected 1 and each word embedding feature in the multiple word embedding features, a third target word embedding feature is screened from the multiple word embedding features; the attribute labels "drooping tail", "wolf" and "pointed ears" corresponding to the third target word embedding feature are added to the multiple target attribute labels.

[0148] For image n to be detected, if it is determined that the attribute of the object in the target detection result representing the target image n does not exist in the multiple attribute labels in the sample image set, a third target word embedding feature is selected from the multiple word embedding features based on the same method; and the attribute labels "car" and "steel" corresponding to the third target word embedding feature are added to the multiple target attribute labels to expand the multiple target attribute labels.

[0149] In the case that the target detection results of other images to be detected, including the image to be detected 2, represent that the attributes of the object in the image to be detected do not exist in the multiple attribute labels in the sample image set, the multiple target attribute labels can be expanded by the above method.

[0150] According to an embodiment of the present disclosure, the image detection method also includes: when it is determined that the detection result represents that the attribute of the image to be detected exists in multiple attribute labels in the sample image set, determining the attribute of the object in the image to be detected based on the similarity between the visual features of the image to be detected and each first text description feature.

[0151] According to an embodiment of the present disclosure, when it is determined that the detection result represents that the attribute of the image to be detected exists in multiple attribute labels in the sample image set, the similarity between the visual features of the image to be detected and each first text description feature is calculated, and the first text description feature with the greatest similarity to the visual features of the image to be detected is determined, and the attribute label corresponding to the first text description feature is determined as the attribute of the object in the image to be detected.

[0152] According to the embodiments of the present disclosure, by directly comparing visual features with text description features, image content and semantic information can be aligned in a unified high-dimensional space. This alignment fully utilizes the generalization capabilities of pre-trained image detection models to capture the deep semantic connections between images and text, thereby improving the accuracy of fine-grained attribute recognition.

[0153] Based on the above sample generation method, the present disclosure also provides a sample generation device. Figure 7 The device is described in detail.

[0154] Figure 7 The structural block diagram of the sample generating device according to an embodiment of the present disclosure is schematically shown.

[0155] like Figure 7 As shown, the sample generating apparatus 700 of this embodiment includes an acquisition module 710 , a first determination module 720 , a second determination module 730 , a third determination module 740 and a fourth determination module 750 .

[0156] an acquisition module, configured to acquire a sample image set, the sample image set comprising a plurality of first sample images and an attribute label of a first object in each first sample image;

[0157] A first determining module, configured to determine a first text description feature corresponding to each attribute tag;

[0158] A second determining module is configured to determine target attribute labels from a plurality of initial attribute labels in a knowledge base based on a first text description feature corresponding to each attribute label, to obtain a plurality of target attribute labels, wherein the plurality of attribute labels and the plurality of target attribute labels are different;

[0159] a third determining module, configured to determine a detection result of whether an attribute of a second object representing the second sample image exists in the plurality of attribute labels in the sample image set based on a similarity between the visual feature of the second sample image and the first text description features corresponding to each of the plurality of attribute labels, and a similarity between the visual feature of the second sample image and the target text description features corresponding to each of the plurality of target attribute labels;

[0160] The fourth determination module generates training samples based on the detection results corresponding to the second sample image and the second sample image.

[0161] According to an embodiment of the present disclosure, the attribute labels in the sample image set belong to the same attribute category; the second determination module includes: a feature determination unit, a cluster determination unit, a first screening unit, and a first determination unit.

[0162] The feature determination unit is configured to perform feature extraction on each of the multiple initial attribute labels in the knowledge base to obtain multiple word embedding features.

[0163] The cluster determination unit is configured to cluster the multiple word embedding features to obtain multiple target clusters.

[0164] The first screening unit is configured to, for each first text description feature, based on the distances between the cluster center word embedding features of the multiple target clusters and the first text description feature, screen the first target word embedding features from the target clusters to obtain multiple first target word embedding features.

[0165] The first determination unit is configured to determine the initial attribute labels corresponding to the multiple first target word embedding features as target attribute labels to obtain multiple target attribute labels.

[0166] According to an embodiment of the present disclosure, the second determination module includes: for each first sample image: an image cropping unit, a second screening unit, a similarity determination unit, a third screening unit, and a second determination unit.

[0167] The image cropping unit is configured to crop the first sample image to obtain M image segments respectively, where M is an integer greater than or equal to 2;

[0168] The second screening unit is configured to, based on the similarities between the sub-visual features corresponding to each of the M image segments and the first text description feature, screen L first target sub-visual features from the M sub-visual features as the first predicted text description features, and screen L second target sub-visual features as the second predicted text description features, where 0 < L < M / 2 and L is an integer, and the first target sub-visual features are different from the second target sub-visual features;

[0169] The similarity determination unit is configured to respectively obtain a first similarity value, a second similarity value, and a third similarity value corresponding to each word embedding feature according to the similarities between each word embedding feature in the multiple word embedding features and the first predicted text description feature, the similarities between each word embedding feature and the second predicted text description feature, and the similarities between each word embedding feature and the first text description feature, where the multiple word embedding features are obtained by performing feature extraction on each of the multiple initial attribute labels in the knowledge base.

[0170] The third screening unit is used to screen a second target word embedding feature from the multiple word embedding features based on the first similarity value, the second similarity value and the third similarity value corresponding to each word embedding feature.

[0171] The second determining unit is configured to determine an initial attribute label corresponding to the second target word embedding feature obtained corresponding to each first sample image as a target attribute label, to obtain a plurality of target attribute labels.

[0172] According to an embodiment of the present disclosure, the third screening unit includes: a first determining subunit and a second determining subunit.

[0173] The first determination subunit is used to weightedly sum the first similarity value, the second similarity value, and the third similarity value based on the weight coefficients corresponding to the first similarity value, the second similarity value, and the third similarity value corresponding to each word embedding feature to obtain a joint similarity value corresponding to each word embedding feature.

[0174] The second determining subunit is configured to select a second target word embedding feature from the plurality of word embedding features based on a joint similarity value corresponding to each word embedding feature, wherein the weight coefficient corresponding to the first similarity value is greater than the weight coefficient corresponding to the second similarity value and the weight coefficient corresponding to the third similarity value.

[0175] According to an embodiment of the present disclosure, the third determining module includes: a label grouping unit, a third determining unit, a fourth determining unit, a fifth determining unit, a sixth determining unit, and a seventh determining unit.

[0176] The label grouping unit is used to group multiple target attribute labels to obtain N groups of target attribute label sets, where N is an integer greater than or equal to 2.

[0177] The third determination unit is used to determine, for each group of target attribute label sets, the similarity values ​​corresponding to each of the multiple target attribute labels in the target attribute label set based on the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels in the target attribute label set.

[0178] The fourth determining unit is configured to obtain a plurality of fourth similarity values ​​according to similarities between the visual features of the second sample image and the first text description features corresponding to each of the plurality of attribute tags.

[0179] The fifth determining unit is configured to determine, for each target attribute tag set, an initial detection probability value according to the similarity values ​​corresponding to the multiple target attribute tags in the target attribute tag set and the multiple fourth similarity values.

[0180] a sixth determining unit, configured to determine a target detection probability value according to the initial detection probability value corresponding to each set of target attribute label sets;

[0181] The seventh determining unit is configured to determine, based on the target detection probability value, a detection result of whether the attribute of the second object characterizing the second sample image exists in the plurality of attribute labels in the sample image set.

[0182] According to embodiments of the present disclosure, any multiple modules in the acquisition module 710, the first determination module 720, the second determination module 730, the third determination module 740, and the fourth determination module 750 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the acquisition module 710, the first determination module 720, the second determination module 730, the third determination module 740, and the fourth determination module 750 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the acquisition module 710, the first determination module 720, the second determination module 730, the third determination module 740, and the fourth determination module 750 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0183] It should be noted that the large model deployment device part in the embodiment of the present disclosure corresponds to the large model deployment method part in the embodiment of the present disclosure. The description of the large model deployment device part specifically refers to the large model deployment method part, which will not be repeated here.

[0184] Those skilled in the art will appreciate that various combinations and / or combinations of features described in the various embodiments and / or claims of this disclosure may be made, even if such combinations or combinations are not explicitly described in this disclosure. In particular, various combinations and / or combinations of features described in the various embodiments and / or claims of this disclosure may be made, without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0185] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A sample generation method, characterized in that: The method includes: Obtaining a sample image set, where the sample image set includes multiple first sample images and the attribute labels of the first object in each of the first sample images; Determining first text description features corresponding to each of the attribute labels; Based on the first text description features corresponding to each of the attribute labels, respectively determining target attribute labels from multiple initial attribute labels in a knowledge base, obtaining multiple target attribute labels, where the multiple attribute labels and the multiple target attribute labels are all different; Based on the similarity between the visual features of a second sample image and the first text description features corresponding to the multiple attribute labels respectively, and the similarity between the visual features of the second sample image and the target text description features corresponding to the multiple target attribute labels respectively, determining a detection result on whether the attributes of the second object representing the second sample image exist in the multiple attribute labels in the sample image set; Generating a training sample based on the detection result corresponding to the second sample image and the second sample image.

2. The method according to claim 1, characterized in that The attribute labels in the sample image set belong to the same attribute category; The determining, based on the first text description features corresponding to each of the attribute labels, of target attribute labels from multiple initial attribute labels in a knowledge base, obtaining multiple target attribute labels, includes: Respectively performing feature extraction on each of the multiple initial attribute labels in the knowledge base to obtain multiple word embedding features; Clustering the multiple word embedding features to obtain multiple target clusters; For each of the first text description features, based on the distances between the cluster center word embedding features of the multiple target clusters and the first text description feature, screening first target word embedding features from the target clusters to obtain multiple first target word embedding features; Determining the initial attribute labels corresponding to the multiple first target word embedding features as the target attribute labels to obtain the multiple target attribute labels.

3. The method according to claim 1, characterized in that The determining, based on the first text description features corresponding to each of the attribute labels, of target attribute labels from multiple initial attribute labels in a knowledge base, obtaining multiple target attribute labels, includes: For each of the first sample images: Cropping the first sample image to respectively obtain M image segments, where M is an integer greater than or equal to 2; Based on the similarity between the sub-visual features corresponding to each of the M image segments and the first text description feature, screening L first target sub-visual features as first predicted text description features and screening L second target sub-visual features as second predicted text description features from the M sub-visual features, where 0 < L < M / 2 and L is an integer, and the first target sub-visual features are different from the second target sub-visual features; According to the similarity between each word embedding feature in the plurality of word embedding features and the first predicted text description feature, the similarity between each word embedding feature and the second predicted text description feature, and the similarity between each word embedding feature and the first text description feature, respectively obtaining a first similarity value, a second similarity value, and a third similarity value corresponding to each word embedding feature, wherein the plurality of word embedding features are obtained by performing feature extraction on each of the plurality of initial attribute labels in the knowledge base; Based on the first similarity value, the second similarity value, and the third similarity value corresponding to each of the word embedding features, selecting a second target word embedding feature from the multiple word embedding features; The initial attribute label corresponding to the second target word embedding feature obtained corresponding to each of the first sample images is determined as the target attribute label, to obtain a plurality of the target attribute labels.

4. The method according to claim 3, characterized in that The selecting a second target word embedding feature from the plurality of word embedding features based on the first similarity value, the second similarity value, and the third similarity value corresponding to each word embedding feature includes: Based on the weight coefficients corresponding to the first similarity value, the second similarity value, and the third similarity value corresponding to each of the word embedding features, weightedly summing the first similarity value, the second similarity value, and the third similarity value to obtain a joint similarity value corresponding to each of the word embedding features; Based on the joint similarity value corresponding to each of the word embedding features, filtering out the second target word embedding feature from the multiple word embedding features; The weight coefficient corresponding to the first similarity value is greater than the weight coefficient corresponding to the second similarity value and the weight coefficient corresponding to the third similarity value.

5. The method according to claim 1, characterized in that The detecting result of determining whether the attribute of the second object representing the second sample image exists in the plurality of attribute labels in the sample image set based on the similarity between the visual feature of the second sample image and the first text description features corresponding to each of the plurality of attribute labels, and the similarity between the visual feature of the second sample image and the target text description features corresponding to each of the plurality of target attribute labels, includes: Grouping the plurality of target attribute labels to obtain N sets of target attribute label sets, where N is an integer greater than or equal to 2; For each target attribute label set, determining a similarity value corresponding to each of the multiple target attribute labels in the target attribute label set based on the similarity between the visual features of the second sample image and the target text description features corresponding to each of the multiple target attribute labels in the target attribute label set; Obtaining a plurality of fourth similarity values ​​according to similarities between the visual features of the second sample image and the first text description features corresponding to each of the plurality of attribute tags; For each target attribute tag set, determining an initial detection probability value according to the similarity values ​​corresponding to the plurality of target attribute tags in the target attribute tag set and the plurality of fourth similarity values; Determining a target detection probability value according to the initial detection probability value corresponding to each group of the target attribute label sets; Based on the target detection probability value, determining whether the attribute of the second object characterizing the second sample image exists in the detection result of the plurality of attribute labels in the sample image set.

6. A method for training an image detection model, characterized in that: The method comprises: Inputting a training sample including at least a second sample image into an image detection model, and outputting a training detection result for characterizing whether an attribute of a second object in the second sample image exists in a plurality of attribute labels in the sample image set; Training the image detection model based on the training detection result corresponding to the second sample image and the sample label of the second sample image; The sample label of the second sample image is determined based on the detection result obtained by the method according to any one of claims 1 to 5.

7. An image detection method, characterized in that: The method comprises: Inputting the image to be detected into a pre-trained image detection model, and outputting a target detection result for characterizing whether the attribute of the object in the image to be detected exists in multiple attribute labels in the sample image set; Wherein, the pre-trained image detection model is obtained based on the method described in claim 6.

8. The method according to claim 7, characterized in that The method further comprises: In a case where it is determined that the attribute of the object in the to-be-detected image represented by the target detection result does not exist in the plurality of attribute labels in the sample image set, selecting a third target word embedding feature from the plurality of word embedding features based on a similarity between the visual feature of the to-be-detected image and each word embedding feature in the plurality of word embedding features, wherein the plurality of word embedding features are obtained by performing feature extraction on each of the plurality of initial attribute labels in the knowledge base; The attribute label corresponding to the third target word embedding feature is added to the multiple target attribute labels.

9. The method according to claim 8, characterized in that The method further comprises: When it is determined that the detection result represents that the attribute of the image to be detected exists in multiple attribute labels in the sample image set, the attribute of the object in the image to be detected is determined based on the similarity between the visual features of the image to be detected and each of the first text description features.

10. A sample generating device, characterized in that: The device comprises: an acquisition module, configured to acquire a sample image set, wherein the sample image set includes a plurality of first sample images and an attribute label of a first object in each of the first sample images; A first determining module, configured to determine a first text description feature corresponding to each of the attribute tags; a second determining module, configured to determine a target attribute label from a plurality of initial attribute labels in a knowledge base based on the first text description feature corresponding to each of the attribute labels, to obtain a plurality of target attribute labels, wherein the plurality of attribute labels and the plurality of target attribute labels are different; a third determining module, configured to determine, based on a similarity between a visual feature of the second sample image and each of the first text description features corresponding to each of the plurality of attribute labels, and a similarity between the visual feature of the second sample image and each of the target text description features corresponding to each of the plurality of target attribute labels, a detection result of determining whether an attribute of the second object characterizing the second sample image exists in the plurality of attribute labels in the sample image set; The fourth determining module generates a training sample based on the detection result corresponding to the second sample image and the second sample image.

Citation Information

Patent Citations

  • Zero-sample image recognition method and system based on generative adversarial network

    CN111476294A

  • Zero-sample learning algorithm based on multi-network cooperation

    CN111738313A

  • Text generation method and device based on multiple modes and model training method and device based on multiple modes

    CN114298121A

  • Sample generation method and device, terminal equipment and computer readable storage medium

    CN115204267A

  • Image class incremental learning method and system based on out-of-distribution detection

    CN117079011A