Attribute recognition model training method and device, equipment, storage medium and program product
Through visual encoding and text encoding of data by graphic and text, combined with noise learning to optimize the attribute recognition model, the training difficulties of existing models in open set attribute recognition are solved, and the recognition of unseen attributes in the real world is realized, and the applicability and efficiency of the model are improved.
Patent Information
- Application Number
- CN202510277420.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-07-11
AI Technical Summary
Existing attribute recognition models are difficult to achieve effective training in open set-oriented attribute recognition tasks, especially when the number of attributes far exceeds the number of attributes in the database in the arbitrary description of objects in the real world, resulting in huge training resources.
By obtaining the graphic and text data, using the attribute recognition model to visually encode and text code the image and text description information, determine the matching relationship between candidate targets and attribute words, and train the model based on the matching results, introduce noise learning to extract the finer-grained target-attribute matching relationship, and optimize the model to adapt to open set attribute recognition.
It realizes the ability to identify attributes that do not appear in the database, improves the open set-oriented attribute recognition performance, reduces resource consumption, and improves the applicability of the model in the real world.
Smart Images

Figure CN120299052A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer application technologies, and particularly to a method, device, equipment, storage medium and program product for training an attribute recognition model. Background Art
[0002] Describing target attributes not only facilitates the accuracy of target recognition, but also enables a clearer and more accurate understanding of the target. The diversity of attributes in nature has brought great difficulties to attribute annotation. Early work mainly focused on target attribute recognition in fixed fields, such as clothes, shoes, animals, human faces, etc., which would limit the number of attribute categories. However, if attribute recognition is extended from a fixed single field to the entire field, the number of attributes in existing datasets is far from sufficient to achieve arbitrary descriptions of objects in the real world, and large-scale attribute annotation for different targets is extremely resource-consuming. Based on this, how to train an attribute recognition model to ensure that the trained attribute recognition model can achieve an open-set oriented attribute recognition task is a technical problem that urgently needs to be solved currently. Summary of the Invention
[0003] Embodiments of the present application provide a method, device, equipment, storage medium and program product for training an attribute recognition model, which can achieve an open-set oriented attribute recognition task.
[0004] On the one hand, an embodiment of the present application provides a method for training an attribute recognition model, and the method includes:
[0005] Obtain a first training sample; wherein, the first training sample includes graphic-text pair data, and the graphic-text pair data includes a first training image and text description information of the first training image;
[0006] Call an attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain visual encoding features of each candidate target;
[0007] Perform text encoding on each attribute word included in the text description information to obtain text encoding features of each attribute word;
[0008] Based on the visual encoding features of each candidate target and the text encoding features of each attribute word, determine a matching result between each candidate target and each attribute word; wherein, the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word;
[0009] Train the attribute recognition model in the direction of reducing the difference between the matching results of each determined candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on the image to be recognized.
[0010] In one embodiment, the method further includes:
[0011] Obtain a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate the target attribute word of the target included in the second training image;
[0012] Crop a local image from the second training image based on the attribute annotation label; wherein, the local image includes the target described by the target attribute word indicated by the attribute label;
[0013] Call a pre-trained attribute recognition model to perform visual encoding on the local image to obtain the visual encoding feature of the local image;
[0014] Perform text encoding on the target attribute word to obtain the text encoding feature of the target attribute word;
[0015] Based on the visual encoding feature of the local image and the text encoding feature of the target attribute word, obtain the matching probability between the local image and the target attribute word;
[0016] Take that the matching probability is higher than the first probability threshold as the optimization objective, and optimize the pre-trained attribute recognition model to obtain the attribute recognition model.
[0017] In one embodiment, the cropping the local image from the second training image based on the attribute label includes:
[0018] Perform image slicing on the second training image to obtain candidate images; wherein, the candidate images at least include the target described by the attribute word indicated by the attribute label, and the size of the candidate images is larger than the size of the detection frame of the target in the second training image;
[0019] Perform random cropping on the candidate images to obtain the local images; wherein, the size of the local images is smaller than the size of the candidate images, and the local images include part or all of the content of the target described by the target attribute word.
[0020] In one embodiment, the pre-trained attribute recognition model includes a visual encoder and a text encoder. The visual encoder is used to perform visual encoding on the local image, and the text encoder is used to perform text encoding on the target attribute word.
[0021] Optimizing the pre-trained attribute recognition model with the matching probability being higher than the first probability threshold as the optimization objective to obtain the attribute recognition model includes:
[0022] Optimizing the text encoder with the matching probability being higher than the first probability threshold as the optimization objective to obtain the attribute recognition model, where the attribute recognition model includes the visual encoder and the optimized text encoder.
[0023] In one embodiment, obtaining the matching probability between the local image and the attribute word indicated by the attribute annotation label based on the visual encoding feature of the local image and the text encoding feature of the target attribute word includes:
[0024] Obtaining the similarity between the visual encoding feature of the local image and the text encoding feature of the target attribute word;
[0025] Based on the similarity, obtaining the matching probability between the local image and the target attribute word; wherein, the matching probability between the local image and the target attribute word is positively correlated with the similarity.
[0026] In one embodiment, the method further includes:
[0027] Invoking the attribute recognition model to perform visual encoding on the first training image to obtain the visual encoding feature of the first training image;
[0028] Performing text encoding on the text description information to obtain the text encoding feature of the text description information;
[0029] Based on the visual encoding feature of the first training image and the text encoding feature of the text description information, determining the matching result between the first training image and the text description information;
[0030] Training the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation to obtain the trained attribute recognition model includes:
[0031] Train the attribute recognition model in the direction of reducing the difference between the matching result of the determined first training image and the text description information and the preset matching label of the first training image and the text description information, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word that are labeled, to obtain the trained attribute recognition model; wherein, the preset matching label of the first training image and the text description information is used to indicate that the first training image and the text description information have a matching relationship.
[0032] In one embodiment, the method further includes:
[0033] Call the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding feature of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image;
[0034] Based on the visual encoding feature of the target unit image and the text encoding features of each attribute word, determine the matching results of the target unit image and each attribute word;
[0035] The step of training the attribute recognition model in the direction of reducing the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word that are labeled to obtain the trained attribute recognition model includes:
[0036] Train the attribute recognition model in the direction of reducing the difference between the matching results of the determined target unit image and each attribute word and the preset matching label of the target unit image and each attribute word, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word that are labeled, to obtain the trained attribute recognition model; wherein, the preset matching label of the target unit image and each attribute word is used to indicate that the target unit image and each attribute word have a matching relationship.
[0037] In one embodiment, the method further includes:
[0038] Call the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding feature of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image;
[0039] Perform text encoding on each noun phrase included in the text description information to obtain the text encoding features of each noun phrase;
[0040] Based on the visual encoding features of the target unit image and the text encoding features of each noun phrase, determine the matching results between the target unit image and each noun phrase;
[0041] Training the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model, including:
[0042] Training the attribute recognition model in the direction of reducing the difference between the determined matching results of the target unit image and each noun phrase and the preset matching labels of the target unit image and each noun phrase, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the preset matching labels of the target unit image and each noun phrase are used to indicate that the target unit image and each noun phrase have a matching relationship.
[0043] In one embodiment, the method further includes:
[0044] Invoke the attribute recognition model, and based on the visual encoding features of each candidate target and the text encoding features of each attribute word, obtain the matching probabilities of each candidate target and each attribute word;
[0045] Compare the matching probabilities of each candidate target and each attribute word with a second probability threshold, and label the matching labels of each candidate target and each attribute word according to the comparison result.
[0046] On the other hand, an embodiment of the present application provides a training device for an attribute recognition model, and the training device for the attribute recognition model includes:
[0047] An acquisition unit, configured to acquire a first training sample; wherein, the first training sample includes graphic-text pair data, and the graphic-text pair data includes a first training image and the text description information of the first training image;
[0048] A visual encoding unit, configured to invoke an attribute recognition model to perform visual encoding on each candidate target included in the first training image, to obtain the visual encoding features of each candidate target;
[0049] A text encoding unit, configured to perform text encoding on each attribute word included in the text description information to obtain text encoding features of each attribute word;
[0050] A determination unit, configured to determine a matching result between each candidate target and each attribute word based on the visual encoding features of each candidate target and the text encoding features of each attribute word; wherein, the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word;
[0051] A model training unit, configured to train the attribute recognition model in a direction of reducing the difference between the determined matching result between each candidate target and each attribute word and the matching label of the corresponding candidate target and the corresponding attribute word in the annotation, to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on an image to be recognized.
[0052] On the other hand, an embodiment of the present application provides a computer device, including a processor, a storage device, and a communication interface, the processor, the storage device, and the communication interface are interconnected, wherein, the storage device is configured to store a computer program for supporting the computer device to execute the above method, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the following steps:
[0053] Obtain a first training sample; wherein, the first training sample includes text-image pair data, and the text-image pair data includes a first training image and text description information of the first training image;
[0054] Call an attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain visual encoding features of each candidate target;
[0055] Perform text encoding on each attribute word included in the text description information to obtain text encoding features of each attribute word;
[0056] Based on the visual encoding features of each candidate target and the text encoding features of each attribute word, determine a matching result between each candidate target and each attribute word; wherein, the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word;
[0057] Train the attribute recognition model in a direction of reducing the difference between the determined matching result between each candidate target and each attribute word and the matching label of the corresponding candidate target and the corresponding attribute word in the annotation, to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on an image to be recognized.
[0058] On the other hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to execute the above-mentioned training method of the attribute recognition model.
[0059] On the other hand, an embodiment of the present application provides a computer program product, the computer program product including a computer program adapted to be loaded and executed by a processor to execute the above-mentioned training method of the attribute recognition model.
[0060] In the embodiment of the present application, weak supervision training can be performed on data using image-text pairs. Specifically, a first training sample is obtained. The first training sample includes image-text pair data, and the image-text pair data includes a first training image and text description information of the first training image. Then, the attribute recognition model is called to perform visual encoding on each candidate target included in the first training image to obtain visual encoding features of each candidate target, and text encoding is performed on each attribute word included in the text description information to obtain text encoding features of each attribute word. Based on the visual encoding features of each candidate target and the text encoding features of each attribute word, a matching result between each candidate target and each attribute word is determined. The matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word. The attribute recognition model is trained in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate targets and corresponding attribute words in the annotation, and a trained attribute recognition model is obtained. It can be seen that the embodiment of the present application extracts a finer-grained target-attribute matching relationship from the coarse-grained matching relationship of the image-text pair, and then introduces noise learning to obtain a good representation of the attribute. That is to say, the trained attribute recognition model obtained through the embodiment of the present application can implement an open-set oriented attribute recognition task. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 is a schematic structural diagram of a training system for an attribute recognition model provided by an embodiment of the present application;
[0063] Figure 2 is a schematic flowchart of a training method for an attribute recognition model provided by an embodiment of the present application;
[0064] Figure 3It is a schematic structural diagram of another training system for an attribute recognition model provided by an embodiment of the present application;
[0065] Figure 4 It is a schematic flowchart of another training method for an attribute recognition model provided by an embodiment of the present application;
[0066] Figure 5 It is a schematic flowchart of another training method for an attribute recognition model provided by an embodiment of the present application;
[0067] Figure 6 It is a schematic structural diagram of a training device for an attribute recognition model provided by an embodiment of the present application;
[0068] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0070] The description of the target attribute not only helps improve the accuracy of target recognition, but also enables a clearer and more accurate understanding of the target. That is, learning the attribute representation of the target in the field of computer vision can not only supplement the information loss caused by only category description, but also help the intelligent agent deploying the artificial intelligence algorithm better understand the real world. Therefore, attribute recognition should be extended from a fixed single domain to the entire domain. For example, for a target, in practical applications, in addition to describing the category it belongs to, there should also be attributes such as color, material, texture, and shape. For example, for a tree, it can be tall or short, can be emerald green or withered yellow. However, the existing attribute recognition models can only recognize the attributes already in the database, but the number of attributes already in the database is far from enough to meet the arbitrary descriptions of objects in the real world, and large-scale attribute annotation for different targets is extremely resource-consuming.
[0071] Based on this, an embodiment of the present application provides a training method for an attribute recognition model. Obtain a first training sample, where the first training sample includes image-text pair data. The image-text pair data includes a first training image and the text description information of the first training image. Then, call the attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain the visual encoding features of each candidate target, perform text encoding on each attribute word included in the text description information to obtain the text encoding features of each attribute word, and determine the matching results between each candidate target and each attribute word based on the visual encoding features of each candidate target and the text encoding features of each attribute word. The matching results are used to indicate whether there is a matching relationship between each candidate target and each attribute word. Train the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate targets and corresponding attribute words in the annotation, so as to obtain a trained attribute recognition model. The embodiment of the present application can perform weakly supervised training using image-text pair data, extract a finer-grained object-attribute matching relationship from the coarse-grained matching relationship of the image-text pair, and then introduce noise learning to obtain a good representation of the attributes. That is to say, the trained attribute recognition model obtained through the embodiment of the present application can implement an open-set attribute recognition task. In other words, the trained attribute recognition model in the embodiment of the present application can recognize attributes that do not appear in the database, that is, the trained attribute recognition model can determine whether an object has any visual attribute.
[0072] The training method for the attribute recognition model provided by the embodiment of the present application can be applied to a computer device. The computer device may include a terminal device or a server, etc. The computer device includes but is not limited to a smart phone, a camera, a wearable device, or a computer, etc. It can be understood that the computer device for training the attribute recognition model and the computer device for calling the trained attribute recognition model may be the same computer device or different computer devices. The computer device calls the trained attribute recognition model to perform attribute recognition on the image to be recognized, and the recognized attributes can be applied to the target recognition scenario, such as recognizing whether there is a driver making a call or smoking in a traffic image, or recognizing whether there are items or people endangering public safety in a captured image, etc.
[0073] Please refer to Figure 1 , Figure 1It is a schematic structural diagram of a training system for an attribute recognition model provided by an embodiment of the present application. First, a first training sample can be obtained, where the first training sample can include text-image pair data, and the text-image pair data can include a first training image 101 and text description information 102 of the first training image. Further, the first training image 101 can be subjected to object recognition by a target detector to obtain at least one candidate target 103 included in the first training image 101, and the text description information 102 can be segmented to obtain at least one attribute word 104 included in the text description information 102. Then, at least one candidate target 103 is input into a Visual Encoder, and the Visual Encoder performs visual encoding on each candidate target to obtain visual encoding features of each candidate target. At least one attribute word 104 is input into a Text Encoder, and the Text Encoder performs text encoding on each attribute word to obtain text encoding features of each attribute word. Then, based on the visual encoding features of each candidate target and the text encoding features of each attribute word, a matching result between each candidate target and each attribute word is determined, and the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word. The attribute recognition model is trained in a direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate targets and corresponding attribute words in the annotation, so as to obtain a trained attribute recognition model.
[0074] Based on Figure 1 For the training system of the attribute recognition model shown, please refer to Figure 2 , Figure 2 It is a schematic flowchart of a training method for an attribute recognition model provided by an embodiment of the present application. The training method for the attribute recognition model can be executed by a computer device; as Figure 2 shown, the training solution for the attribute recognition model includes but is not limited to steps S201 to S205, where:
[0075] S201, Obtain a first training sample, where the first training sample includes text-image pair data, and the text-image pair data includes a first training image and text description information of the first training image.
[0076] In one implementation, at least one existing image-text pair data can be obtained from a database, and the obtained at least one image-text pair data is used as the first training sample. For example, the image-text pair data in the database can be sourced from COCO (Common Objects in COntext), which is a dataset provided by the Microsoft team for image recognition. The COCO dataset can include at least one image-text pair data, and each image-text pair data includes object instances (i.e., the first training images in the embodiments of the present application) and image captions (i.e., the text description information in the embodiments of the present application).
[0077] In another implementation, the first training sample can be obtained by manual annotation. For example, at least one first training image is first obtained from a memory or downloaded from the Internet, and then each first training image is manually annotated to obtain the text description information of each first training image.
[0078] S202, call the attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain the visual encoding features of each candidate target.
[0079] In one example, the attribute recognition model in the embodiments of the present application can include a CLIP model or an ALIGN model, etc. Among them, the full name of the CLIP model is "Contrastive Language-Image Pretraining", which means using contrastive language-image pre-training. The basic idea of the CLIP model is to let the machine learn the latent semantic structure from text and images. It establishes a model that contrasts and relates the concepts in the image and text corpus, so that the machine gradually learns and masters the "correct" relationship between images and text. After the model is trained, the entire system can process new text and image instances with relatively high accuracy, such as problems like image-text correspondence and semantic analysis. The Large-scale Image and Noisy-text embedding (ALIGN) model, the ALIGN model is a model that uses large-scale noisy image-text data to expand visual and visual language representation learning. The ALIGN model can perform cross-modal search from image to text, text to image, and even jointly search for queries of image + text.
[0080] In the embodiments of the present application, the CLIP model or the ALIGN model first uses a visual encoder and a text encoder to extract the features of the scene image and the text description respectively, and then learns based on the noise contrastive estimation loss (InfoNCE loss). Thanks to the support of up to hundreds of millions of image-text pairs of data, this pre-trained model (i.e., the CLIP model or the ALIGN model) has strong zero-shot classification ability for both categories and attributes. However, directly applying this model to attribute recognition does not yield satisfactory results. Based on this, the embodiments of the present application utilize the zero-shot classification ability of the CLIP model or the ALIGN model to extract a finer-grained object-attribute matching relationship from the coarse-grained matching relationship of image-text pairs, and then introduce noise learning to obtain a good representation of the attributes, thereby enabling the open-set attribute recognition task to be achieved.
[0081] In another example, the attribute recognition model in the embodiments of the present application may refer to an attribute recognition model trained by the training method of the attribute recognition model Figure 4 shown. Specifically, reference may be made to the description in the subsequent embodiments, and the embodiments of the present application will not elaborate. Figure 5
[0082] In one implementation manner, after obtaining the first training sample, the first training image may be subjected to object recognition by a target detector to obtain at least one candidate object.
[0083] Exemplarily, a category-agnostic target detector may be trained using target detection data such as COCO. This detector follows the design of Faster RCNN, but its classification head is changed to a binary classification, that is, it only classifies the background and the object, and does not give the specific category name. This category-agnostic target detector has a certain recall ability for all objects in the first training image. To obtain the objects in each first training image, we perform forward calculation on the images in COCO Caption using this target detector to obtain the objects included in each first training image. Optionally, to ensure the accuracy of the objects, the top 20 objects with a probability greater than 0.8 may be taken as the candidate objects that can be extracted from this first training image.
[0084] In one implementation manner, the attribute recognition model may include a visual encoder. The unit images corresponding to at least one candidate object included in the first training image may be input into the visual encoder, and then the visual encoder extracts the features of the unit images corresponding to each candidate object to obtain the visual encoding features of each candidate object.
[0085] S203, perform text encoding on each attribute word included in the text description information to obtain the text encoding features of each attribute word.
[0086] In one implementation, after obtaining the first training sample, the text description information can be subjected to part-of-speech tagging and word segmentation to obtain at least one attribute word, and then each attribute word is subjected to text encoding to obtain the text encoding features of each attribute word.
[0087] Exemplarily, the open-source Part-of-Speech Tagging tool TextBlob can be used to split the text description information. For example, for a text description information, its sentence itself (c), noun phrase (p), and attribute word (a) can be obtained. For example, for the text description information "A striped zebra is eating green grass", it can be processed into c: A striped zebra is eating green grass; p: striped zebra, green grass; a: striped, green.
[0088] In one implementation, the attribute recognition model may further include a text encoder. Each attribute word can be input into the text encoder, and then each attribute word is subjected to text encoding by the text encoder to obtain the text encoding features of each attribute word.
[0089] For example, in order to obtain the text encoding of the attribute word for attribute classification, each attribute word can be filled into a pre-designed template to adapt to the input of the text encoder. For example, 16 learnable prompt vectors can be designed, and then the attribute word is inserted in the middle of these 16 prompt vectors, and then this compound sentence is sent into the text encoder to extract the text encoding features of each attribute word.
[0090] S204. Based on the visual encoding features of each candidate target and the text encoding features of each attribute word, determine the matching results of each candidate target and each attribute word. The matching results are used to indicate whether each candidate target and each attribute word have a matching relationship.
[0091] Specifically, for any candidate target among at least one candidate target included in the first training image, any attribute word among at least one attribute word included in the text description information of the first training image can determine the matching result between any candidate target and any attribute word through an attribute recognition model based on the visual coding feature of any candidate target and the text coding feature of any attribute word. The matching result is used to indicate whether there is a matching relationship between any candidate target and any attribute word. For example, assume that the first training image includes M candidate targets and the text description information of the first training image includes N attribute words. Then, through the attribute recognition model, based on the visual coding feature of the m-th candidate target and the text coding feature of the n-th attribute word, the matching result between the m-th candidate target and the n-th attribute word can be determined, where both m and n are positive integers, 1 ≤ m ≤ M, 1 ≤ n ≤ N, and finally M * N matching results can be obtained.
[0092] S205. Train the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation to obtain the trained attribute recognition model.
[0093] Exemplarily, pseudo-label supervision can be utilized to obtain the BCE loss value, and the attribute recognition model can be trained in the direction of reducing the BCE loss value to obtain the trained attribute recognition model. Specifically, the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation can be obtained, and the BCE loss value can be determined based on this difference. Among them, the BCE loss value is positively correlated with the total difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation.
[0094] Among them, the trained attribute recognition model is used for attribute recognition of the image to be recognized. For example, through the trained attribute recognition model, it can not only recognize that the image to be recognized contains a target such as a tree, but also recognize that this tree is tall and green, etc.
[0095] In one implementation, the attribute recognition model can be called to obtain the matching probabilities of each candidate target and each attribute word based on the visual coding features of each candidate target and the text coding features of each attribute word, compare the matching probabilities of each candidate target and each attribute word with a second probability threshold, and label the matching labels of each candidate target and each attribute word according to the comparison results.
[0096] Specifically, any candidate target among at least one candidate target and any attribute word among at least one attribute word may or may not have a matching relationship. For example, there is no matching relationship between the candidate target "dog" and the attribute word "withered and yellow", while there is a matching relationship between the candidate target "tree" and the attribute word "withered and yellow". Then, the attribute recognition model can be called to obtain the matching probability between each candidate target and each attribute word based on the visual coding features of each candidate target and the text coding features of each attribute word. If the matching probability between a certain candidate target and a certain attribute word is greater than the second probability threshold, it is considered that the candidate target and the attribute word are matched, and the obtained matching label for the candidate target and the attribute word can be a positive match, and this matching label is used to indicate that the candidate target and the attribute word have a matching relationship; if the matching probability between a certain candidate target and a certain attribute word is less than or equal to the second probability threshold, it is considered that the candidate target and the attribute word are not matched, and the obtained matching label for the candidate target and the attribute word can be a negative match, and this matching label is used to indicate that the candidate target and the attribute word do not have a matching relationship.
[0097] For example, the attribute recognition model can be called for pseudo-label annotation. For example, the probability of each candidate target containing each attribute word is obtained, and the attribute words with a probability greater than 0.8, or the attribute words with a probability greater than 0.55 and existing in the text description information, are regarded as positive attributes, and other attribute words are regarded as negative attributes. The matching label between the candidate target and the positive attribute is a positive match, and the matching label between the candidate target and the negative attribute is a negative match.
[0098] Optionally, the similarity between the visual coding feature of any candidate target and the text coding feature of any attribute word can be obtained, and the matching probability between any candidate target and any attribute word is obtained based on this similarity. Among them, the matching probability between any candidate target and any attribute word is positively correlated with the similarity between the visual coding feature of any candidate target and the text coding feature of any attribute word, that is, the higher the similarity between the visual coding feature of any candidate target and the text coding feature of any attribute word, the greater the matching probability between any candidate target and any attribute word.
[0099] Exemplarily, the similarity between the visual coding feature of the candidate target and the text coding feature of the attribute word can be calculated by the following formula (1).
[0100]
[0101] Among them, s can represent the similarity between the visual encoding features of the candidate target and the text encoding features of the attribute word, σ refers to the Sigmoid function, v can represent the visual encoding features of the candidate target, for example, it can be the normalized visual encoding features, a can represent the text encoding features of the attribute word, for example, it can be the normalized text encoding features, and τ can be used to adjust the similarity, for example, it can be set to 100.
[0102] In one implementation, an attribute recognition model can also be called to perform visual encoding on the first training image to obtain the visual encoding features of the first training image, perform text encoding on the text description information to obtain the text encoding features of the text description information, and determine the matching result between the first training image and the text description information based on the visual encoding features of the first training image and the text encoding features of the text description information. Then, train the attribute recognition model in the direction of reducing the difference between the determined matching result of the first training image and the text description information and the preset matching label of the first training image and the text description information, and the difference between the matching results of each candidate target and each attribute word and the corresponding matching labels of the corresponding candidate target and the corresponding attribute word, to obtain the trained attribute recognition model. Among them, the preset matching label of the first training image and the text description information is used to indicate that the first training image and the text description information have a matching relationship.
[0103] In the embodiments of the present application, for the matching of the first training image and the text description information, the noise contrast learning method of CLIP can be followed, and the infoNCE loss is used for supervision. Among them, since the first training image and the text description information are image-text pair data, the first training image and the text description information are matched. Therefore, the preset matching label of the first training image and the text description information is used to indicate that the first training image and the text description information have a matching relationship. Based on this, an attribute recognition model can be called to determine the matching result between the first training image and the text description information based on the visual encoding features of the first training image and the text encoding features of the text description information. This matching result is used to indicate whether the first training image and the text description information have a matching relationship, and then the infoNCE loss is obtained based on this matching result and the preset matching label of the first training image and the text description information. For example, the smaller the difference between this matching result and the preset matching label of the first training image and the text description information, that is, the smaller the infoNCE loss; the larger the difference between this matching result and the preset matching label of the first training image and the text description information, that is, the larger the infoNCE loss. The smaller the infoNCE loss, the more accurate the determined matching result of the first training image and the text description information; the larger the infoNCE loss, the less accurate the determined matching result of the first training image and the text description information.
[0104] For example, in order to obtain the text encoding of the text description information, the text description information can be directly input into the text encoder to extract the text encoding features of the text description information.
[0105] In one implementation, the attribute recognition model can also be called to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image. The target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image. Then, based on the visual encoding features of the target unit image and the text encoding features of each attribute word, the matching result between the target unit image and each attribute word is determined. Furthermore, the attribute recognition model can be trained in the direction of reducing the difference between the determined matching result between the target unit image and each attribute word and the preset matching label of the target unit image and each attribute word, as well as the difference between the matching result between each candidate target and each attribute word and the matching label of the corresponding candidate target and the corresponding attribute word, to obtain the trained attribute recognition model. Among them, the preset matching label of the target unit image and each attribute word is used to indicate that the target unit image and each attribute word have a matching relationship.
[0106] In the embodiments of the present application, for the matching between the target unit image and each attribute word, since the target unit image and each attribute word have a matching relationship, the MIL-NCE loss can be used for supervision. For example, the smaller the difference between the determined matching result between the target unit image and each attribute word and the preset matching label of the target unit image and the corresponding attribute word, that is, the smaller the MIL-NCE loss; the larger the difference between the determined matching result between the target unit image and each attribute word and the preset matching label of the target unit image and the corresponding attribute word, that is, the larger the MIL-NCE loss. The smaller the MIL-NCE loss, the more accurate the determined matching result between the target unit image and each attribute word; the larger the MIL-NCE loss, the less accurate the determined matching result between the target unit image and each attribute word.
[0107] Exemplarily, the MIL-NCE loss can be calculated by the following formula (2).
[0108]
[0109] Among them, L MIL-NCEIt can represent the MIL-NCE loss, that is, the sum of the differences between the matching results of the target unit image and each attribute word, and the preset matching labels of the target unit image and the corresponding attribute words. v can represent the visual encoding features of the target unit image, t can represent the text encoding features of the attribute word, τ is a regulation factor, P is a positive match, which means that the target unit image and the attribute word have a matching relationship, N is a negative match, which means that the target unit image and the attribute word do not have an explicitly given matching relationship.
[0110] In one implementation, an attribute recognition model can also be called to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image. The target unit image refers to the unit image corresponding to the target detection box in the first training image. The target detection box refers to the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image. Then, text encoding can be performed on each noun phrase included in the text description information to obtain the text encoding features of each noun phrase. Based on the visual encoding features of the target unit image and the text encoding features of each noun phrase, the matching results between the target unit image and each noun phrase are determined. Furthermore, the attribute recognition model is trained in the direction of reducing the differences between the determined matching results of the target unit image and each noun phrase and the preset matching labels of the target unit image and each noun phrase, as well as the differences between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word, to obtain the trained attribute recognition model. Among them, the preset matching labels of the target unit image and each noun phrase are used to indicate that the target unit image and each noun phrase have a matching relationship.
[0111] For example, to obtain the text encoding of a noun phrase, each noun phrase can be filled into a pre-designed template to adapt to the input of the text encoder. For example, 8 learnable prompt vectors can be designed, and then the noun phrase is inserted in the middle of these 8 prompt vectors, and then this compound sentence is sent into the text encoder to extract the text encoding features of each noun phrase.
[0112] In the embodiments of the present application, for the matching between the target unit image and each noun phrase, since there is a matching relationship between the target unit image and each noun phrase, the MIL-NCE loss can be used for supervision. For example, the smaller the difference between the determined matching result of the target unit image and each noun phrase and the preset matching label of the target unit image and the corresponding noun phrase, that is, the smaller the MIL-NCE loss; the larger the difference between the determined matching result of the target unit image and each noun phrase and the preset matching label of the target unit image and the corresponding noun phrase, that is, the larger the MIL-NCE loss. Among them, the smaller the MIL-NCE loss, the more accurate the determined matching result of the target unit image and each noun phrase; the larger the MIL-NCE loss, the less accurate the determined matching result of the target unit image and each noun phrase.
[0113] It can be understood that: during the training process of the attribute recognition model, one or more of the above training methods can be combined. For example, the attribute recognition model can be trained in the direction of reducing the difference between the determined matching result of the first training image and the text description information and the preset matching label of the first training image and the text description information, the difference between the matching results of each candidate target and each attribute word and the marked matching labels of the corresponding candidate target and the corresponding attribute word, and the difference between the determined matching result of the target unit image and each noun phrase and the preset matching label of the target unit image and each noun phrase, to obtain the trained attribute recognition model. Another example is that the attribute recognition model can be trained in the direction of reducing the difference between the determined matching result of the first training image and the text description information and the preset matching label of the first training image and the text description information, the difference between the matching results of each candidate target and each attribute word and the marked matching labels of the corresponding candidate target and the corresponding attribute word, the difference between the determined matching result of the target unit image and each noun phrase and the preset matching label of the target unit image and each noun phrase, and the difference between the determined matching result of the target unit image and each noun phrase and the preset matching label of the target unit image and the corresponding noun phrase, to obtain the trained attribute recognition model.
[0114] Optionally, the text encoder for text-encoding the attribute words, the text encoder for text-encoding the text description information, and the text encoder for text-encoding the noun phrases can be the same text encoder or different text encoders, which is not specifically limited in the embodiments of the present application.
[0115] In the case where the attribute recognition model includes a text encoder and a visual encoder, the specific process of training the attribute recognition model can be: adjusting the parameters of the text encoder and the parameters of the visual encoder respectively.
[0116] Through the above training method of the attribute recognition model, beneficial attribute representations can be effectively obtained from these noisy data, improving the open-set attribute classification performance for the real world.
[0117] In the embodiment of this application, a first training sample is obtained. The first training sample includes image-text pair data. The image-text pair data includes a first training image and text description information of the first training image. Then, the attribute recognition model is called to perform visual encoding on each candidate object included in the first training image to obtain visual encoding features of each candidate object, and text encoding is performed on each attribute word included in the text description information to obtain text encoding features of each attribute word. Based on the visual encoding features of each candidate object and the text encoding features of each attribute word, the matching result between each candidate object and each attribute word is determined. The matching result is used to indicate whether there is a matching relationship between each candidate object and each attribute word. The attribute recognition model is trained in the direction of reducing the difference between the determined matching results of each candidate object and each attribute word and the matching labels of the corresponding candidate object and the corresponding attribute word in the annotation, and the trained attribute recognition model is obtained, which can implement the open-set attribute recognition task.
[0118] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of another training system of the attribute recognition model provided by the embodiment of this application. First, a second training sample can be obtained. The second training sample can include a second training image 301 and an attribute annotation label 302 of the second training image. The attribute label is used to indicate the target attribute word of the target included in the second training image. Further, a local image 303 can be cropped from the second training image based on the attribute annotation label. Then, the pre-trained attribute recognition model is called to perform visual encoding on the local image to obtain the visual encoding feature v of the local image. In addition, text encoding can be performed on the target attribute word to obtain the text encoding feature of the target attribute word. Then, based on the visual encoding feature of the local image and the text encoding feature of the target attribute word, the matching probability between the local image and the target attribute word is obtained. Then, with the matching probability being higher than the first probability threshold as the optimization target, the pre-trained attribute recognition model is optimized to obtain the attribute recognition model.
[0119] Optionally, the attribute recognition model may include a Visual Encoder. The method of performing visual encoding on the local image may be: extracting features of the local image through the Visual Encoder to obtain the visual encoding feature of the local image.
[0120] Optionally, the attribute recognition model may include a Text Encoder, and the method for text-encoding the target attribute word may be: extracting the feature of the target attribute word through the Text Encoder to obtain the text-encoding feature of the target attribute word.
[0121] Optionally, after obtaining the matching probability between the local image and the target attribute word, a loss value L may be obtained based on the matching probability and the first probability threshold. cls . For example, the ratio of the first probability threshold to the matching probability may be used as the loss value. If the matching probability is higher than the first probability threshold, the ratio of the first probability threshold to the matching probability is less than 1. That is to say, the smaller the loss value, the greater the matching probability between the local image obtained by the attribute recognition model and the target attribute word, and the more accurate the attribute recognition model.
[0122] Optionally, in the case where the attribute recognition model includes a Text Encoder and a Visual Encoder, the specific process of training the attribute recognition model may be: adjusting the parameters of the Text Encoder.
[0123] Based on the above Figure 3 description, please refer to Figure 4 , Figure 4 which is a schematic flowchart of another method for training an attribute recognition model provided by an embodiment of the present application. The method for training the attribute recognition model may be executed by a computer device; as Figure 4 shown, the training solution of the attribute recognition model includes but is not limited to steps S401 to S406, where:
[0124] S401: Obtain a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate the target attribute word of the target included in the second training image.
[0125] In one implementation, the existing second training image and the attribute annotation label of the second training image may be obtained from a database, and the obtained second training image and the attribute annotation label of the second training image are used as the first training sample. For example, the second training image and the attribute annotation label of the second training image in the database may be sourced from VAW, where VAW includes 621 attribute annotations and can be directly used for training the attribute recognition model.
[0126] In another implementation, the second training sample may be obtained by manual annotation. For example, at least one second training image is first obtained from a memory or at least one second training image is downloaded from the Internet, and then each second training image is manually annotated to obtain the attribute annotation label of each second training image.
[0127] S402. Crop a local image from the second training image based on the attribute annotation label. The local image contains the object described by the target attribute word indicated by the attribute label.
[0128] In one implementation, the second training image can be sliced into candidate images. The candidate image contains at least the object described by the attribute word indicated by the attribute label, and the size of the candidate image is larger than the size of the detection box of the object in the second training image. Then, the candidate image is randomly cropped to obtain a local image. The size of the local image is smaller than the size of the candidate image, and the local image contains part or all of the content of the object described by the target attribute word.
[0129] For example, the local image can be cropped from the second training image through the following formula (3).
[0130] Crop(I,b i )=RandomCrop(RandomScaleCrop(I,b i ,α)β) Formula (3)
[0131] Among them, formula (3) can represent the process of cropping the object b i from the image I. First, randomly scale up the detection box of b i by 0 to α times and crop it from the image I, and then randomly crop a sub-slice from β to 1 times of the obtained image slice for training. Exemplarily, α can be taken as 0.2 and β can be taken as 0.8.
[0132] In the embodiments of the present application, since the detection box of the object is the smallest rectangle of the object, after randomly cropping after scaling up the detection box of the object, the cropped local image may contain part of the object. Because there is an object occlusion situation in the existing object recognition scenario, this cropping method has good robustness to occluded objects and can significantly improve the diversity of objects.
[0133] S403. Call the pre-trained attribute recognition model to perform visual encoding on the local image to obtain the visual encoding feature of the local image.
[0134] For example, the pre-trained attribute recognition model can be CLIP. After obtaining the local image, the local image can be scaled to 224×244 size and then sent into the visual encoder of CLIP to obtain the visual encoding feature of the local image, that is, the visual encoding feature of the object.
[0135] S404. Perform text encoding on the target attribute word to obtain the text encoding feature of the target attribute word.
[0136] For example, in order to obtain the text encoding of an attribute word for attribute classification, the target attribute word can be filled into a pre-designed template to adapt to the input of the text encoder. For example, 16 learnable prompt vectors can be designed, and then the attribute word is inserted in the middle of these 16 prompt vectors, and then this composite sentence is fed into the text encoder to extract the text encoding features of the target attribute word.
[0137] S405. Based on the visual encoding features of the local image and the text encoding features of the target attribute word, obtain the matching probability between the local image and the target attribute word.
[0138] In one implementation, the similarity between the visual encoding features of the local image and the text encoding features of the target attribute word can be obtained. Based on this similarity, the matching probability between the local image and the target attribute word is obtained. The matching probability between the local image and the target attribute word is positively correlated with this similarity. That is, the higher the similarity between the visual encoding features of the local image and the text encoding features of the target attribute word, the greater the matching probability between the local image and the target attribute word.
[0139] Exemplarily, the similarity between the visual encoding features of the local image and the text encoding features of the target attribute word can be calculated by the above formula (1).
[0140] S406. Using the matching probability higher than the first probability threshold as the optimization goal, optimize the pre-trained attribute recognition model to obtain the attribute recognition model.
[0141] Optionally, after obtaining the matching probability between the local image and the target attribute word, a loss value L can be obtained based on this matching probability and the first probability threshold. cls . For example, the ratio of the first probability threshold to this matching probability can be used as the loss value. If the matching probability is higher than the first probability threshold, the ratio of the first probability threshold to this matching probability is less than 1. That is to say, the smaller the loss value, the greater the matching probability between the local image and the target attribute word obtained by the attribute recognition model, and the more accurate the attribute recognition model.
[0142] In one implementation, the pre-trained attribute recognition model can include a visual encoder and a text encoder. The visual encoder is used to perform visual encoding on the local image, and the text encoder is used to perform text encoding on the target attribute word. Based on this, the text encoder can be optimized with the matching probability higher than the first probability threshold as the optimization goal to obtain the attribute recognition model, and the attribute recognition model includes a visual encoder and an optimized text encoder.
[0143] Exemplarily, the pre-trained attribute recognition model can be CLIP. In the embodiments of the present application, by loading the pre-trained CLIP, the visual encoding features of the local image and the text encoding features of the target attribute word are extracted using the method provided in the embodiments of the present application. Based on the visual encoding features and the text encoding features, CLIP is fine-tuned. When fine-tuning, in order to obtain stronger open vocabulary capabilities, only the text encoder part can be fine-tuned, and the parameters of the visual encoder are frozen. Through this fine-tuning, a CLIP model more suitable for open-set attribute recognition can be obtained. Because in this training process, the attributes of each target example are supervised, so we use a multi-label classification loss function (binary cross-entropy loss, BCE) for training.
[0144] In the embodiments of the present application, a second training sample is obtained. The second training sample includes a second training image and an attribute annotation label of the second training image. The attribute label is used to indicate the target attribute word of the target included in the second training image. Based on the attribute annotation label, a local image is cropped from the second training image. The local image includes the target described by the target attribute word indicated by the attribute label. The pre-trained attribute recognition model is called to perform visual encoding on the local image to obtain the visual encoding features of the local image, and text encoding is performed on the target attribute word to obtain the text encoding features of the target attribute word. Based on the visual encoding features of the local image and the text encoding features of the target attribute word, the matching probability between the local image and the target attribute word is obtained. Taking the matching probability being higher than the first probability threshold as the optimization target, the pre-trained attribute recognition model is optimized, which can ensure obtaining an attribute recognition model more suitable for open-set attribute recognition.
[0145] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of another method for training an attribute recognition model provided by the embodiments of the present application. The method for training the attribute recognition model can be executed by a computer device; as Figure 5 shown, the training scheme of the attribute recognition model includes but is not limited to steps S501 to S511, where:
[0146] S501, obtain a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate the target attribute word of the target included in the second training image.
[0147] S502, based on the attribute annotation label, crop a local image from the second training image, where the local image includes the target described by the target attribute word indicated by the attribute label.
[0148] S503, call the pre-trained attribute recognition model to perform visual encoding on the local image to obtain the visual encoding features of the local image.
[0149] S504. Perform text encoding on the target attribute word to obtain the text encoding feature of the target attribute word.
[0150] S505. Based on the visual encoding feature of the local image and the text encoding feature of the target attribute word, obtain the matching probability between the local image and the target attribute word.
[0151] S506. Using the matching probability higher than the first probability threshold as the optimization objective, optimize the pre-trained attribute recognition model to obtain the attribute recognition model.
[0152] For the steps S501 to S506 in the embodiments of the present application, reference may be made to the relevant descriptions of the steps S401 to S406 in the above embodiments, and the embodiments of the present application will not be elaborated herein.
[0153] S507. Obtain the first training sample, where the first training sample includes text-image pair data, and the text-image pair data includes the first training image and the text description information of the first training image.
[0154] S508. Invoke the attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain the visual encoding features of each candidate target.
[0155] S509. Perform text encoding on each attribute word included in the text description information to obtain the text encoding features of each attribute word.
[0156] S510. Based on the visual encoding features of each candidate target and the text encoding features of each attribute word, determine the matching results between each candidate target and each attribute word, and the matching results are used to indicate whether there is a matching relationship between each candidate target and each attribute word.
[0157] S511. Train the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation, to obtain the trained attribute recognition model.
[0158] For the steps S507 to S511 in the embodiments of the present application, reference may be made to the relevant descriptions of the steps S201 to S205 in the above embodiments, and the embodiments of the present application will not be elaborated herein.
[0159] The above steps S501 to S506 can only have good recognition ability for the attributes that appear in the dataset, and the recognition ability for other unknown attributes completely comes from the pre-trained attribute recognition model, with limited performance. Therefore, through steps S507 to S511, the attribute recognition model is further fine-tuned using the finer-grained target-attribute pairs extracted from the text-image pair data.
[0160] In the embodiments of the present application, in the first stage, training is performed on a limited attribute recognition data set, that is, the pre-trained attribute recognition model is fine-tuned using the attribute recognition annotation data, so as to better align the visual representation of the target with the attribute words and obtain a higher open-set attribute classification performance. In the second stage, training is performed using the text-image pair data that is extremely easy to obtain on the Internet, that is, weak supervision training is performed using the text-image pair data. First, a finer-grained target-attribute matching relationship is extracted from the coarser-grained matching relationship of the text-image pair, and then noise learning is introduced to obtain a good representation of the attribute, so as to ensure that the trained attribute recognition model can implement the open-set oriented attribute recognition task.
[0161] The embodiments of the present application also provide a computer storage medium, in which program instructions are stored, and when the program instructions are executed, they are used to implement the corresponding methods described in the above embodiments.
[0162] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a training device for an attribute recognition model provided by the embodiments of the present application.
[0163] In an implementation manner of the training device for the attribute recognition model of the embodiments of the present application, the training device for the attribute recognition model includes the following structure.
[0164] An acquisition unit 601, configured to acquire a first training sample; wherein, the first training sample includes text-image pair data, and the text-image pair data includes a first training image and text description information of the first training image;
[0165] A visual encoding unit 602, configured to call an attribute recognition model to perform visual encoding on each candidate target included in the first training image, so as to obtain visual encoding features of the respective candidate targets;
[0166] A text encoding unit 603, configured to perform text encoding on each attribute word included in the text description information, so as to obtain text encoding features of the respective attribute words;
[0167] A determination unit 604, configured to determine a matching result between each candidate target and each attribute word based on the visual encoding features of the respective candidate targets and the text encoding features of the respective attribute words; wherein, the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word;
[0168] A model training unit 605 is configured to train the attribute recognition model in the direction of reducing the difference between the matching results of each determined candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, so as to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on the image to be recognized.
[0169] In one embodiment, the obtaining unit 601 is further configured to obtain a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate the target attribute word of the target included in the second training image.
[0170] The above-mentioned training device for the attribute recognition model may further include:
[0171] A cropping unit 606 is configured to crop a local image from the second training image based on the attribute annotation label; wherein, the local image includes the target described by the target attribute word indicated by the attribute label.
[0172] A visual encoding unit 602 is further configured to call a pre-trained attribute recognition model to perform visual encoding on the local image to obtain the visual encoding feature of the local image.
[0173] A text encoding unit 603 is further configured to perform text encoding on the target attribute word to obtain the text encoding feature of the target attribute word.
[0174] A determining unit 604 is further configured to obtain the matching probability between the local image and the target attribute word based on the visual encoding feature of the local image and the text encoding feature of the target attribute word.
[0175] The model training unit 605 is further configured to optimize the pre-trained attribute recognition model with the matching probability being higher than a first probability threshold as the optimization target to obtain the attribute recognition model.
[0176] In one embodiment, the cropping unit 606 cropping a local image from the second training image based on the attribute label includes:
[0177] Performing image slicing on the second training image to obtain candidate images; wherein, the candidate images at least include the target described by the attribute word indicated by the attribute label, and the size of the candidate images is larger than the size of the detection frame of the target in the second training image.
[0178] Randomly crop the candidate image to obtain the local image; wherein, the size of the local image is smaller than the size of the candidate image, and the local image contains part or all of the content of the object described by the target attribute word.
[0179] In one embodiment, the pre-trained attribute recognition model includes a visual encoder and a text encoder. The visual encoder is used to perform visual encoding on the local image, and the text encoder is used to perform text encoding on the target attribute word.
[0180] The model training unit 605 optimizes the pre-trained attribute recognition model with the matching probability higher than the first probability threshold as the optimization target to obtain the attribute recognition model, including:
[0181] Optimize the text encoder with the matching probability higher than the first probability threshold as the optimization target to obtain the attribute recognition model, where the attribute recognition model includes the visual encoder and the optimized text encoder.
[0182] In one embodiment, the determination unit 604 obtains the matching probability between the local image and the attribute word indicated by the attribute annotation label based on the visual encoding feature of the local image and the text encoding feature of the target attribute word, including:
[0183] Obtain the similarity between the visual encoding feature of the local image and the text encoding feature of the target attribute word.
[0184] Based on the similarity, obtain the matching probability between the local image and the target attribute word; wherein, the matching probability between the local image and the target attribute word is positively correlated with the similarity.
[0185] In one embodiment, the visual encoding unit 602 is further configured to call the attribute recognition model to perform visual encoding on the first training image to obtain the visual encoding feature of the first training image.
[0186] The text encoding unit 603 is further configured to perform text encoding on the text description information to obtain the text encoding feature of the text description information.
[0187] The determination unit 604 is further configured to determine the matching result between the first training image and the text description information based on the visual encoding feature of the first training image and the text encoding feature of the text description information.
[0188] The model training unit 605 trains the attribute recognition model in the direction of reducing the difference between the matching results of each determined candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation, and obtains the trained attribute recognition model, including:
[0189] Train the attribute recognition model in the direction of reducing the difference between the matching result of the determined first training image and the text description information and the preset matching label of the first training image and the text description information, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the preset matching label of the first training image and the text description information is used to indicate that the first training image and the text description information have a matching relationship.
[0190] In one embodiment, the visual encoding unit 602 is further configured to call the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image;
[0191] The determination unit 604 is further configured to determine the matching results between the target unit image and each attribute word based on the visual encoding features of the target unit image and the text encoding features of each attribute word.
[0192] The model training unit 605 trains the attribute recognition model in the direction of reducing the difference between the matching results of each determined candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation, and obtains the trained attribute recognition model, including:
[0193] Train the attribute recognition model in the direction of reducing the difference between the matching results of the determined target unit image and each attribute word and the preset matching labels of the target unit image and each attribute word, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and the corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the preset matching labels of the target unit image and each attribute word are used to indicate that the target unit image and each attribute word have a matching relationship.
[0194] In one embodiment, the visual encoding unit 602 is further configured to call the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image;
[0195] The text encoding unit 603 is further configured to perform text encoding on each noun phrase included in the text description information to obtain the text encoding features of each noun phrase;
[0196] The determination unit 604 is further configured to determine the matching result between the target unit image and each noun phrase based on the visual encoding features of the target unit image and the text encoding features of each noun phrase;
[0197] The model training unit 605 trains the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word marked, to obtain the trained attribute recognition model, including:
[0198] Train the attribute recognition model in the direction of reducing the difference between the determined matching results of the target unit image and each noun phrase and the preset matching labels of the target unit image and each noun phrase, as well as the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word marked, to obtain the trained attribute recognition model; wherein, the preset matching labels of the target unit image and each noun phrase are used to indicate that the target unit image and each noun phrase have a matching relationship.
[0199] In one embodiment, the determination unit 604 is further configured to call the attribute recognition model, and based on the visual encoding features of each candidate target and the text encoding features of each attribute word, obtain the matching probabilities of each candidate target and each attribute word; compare the matching probabilities of each candidate target and each attribute word with a second probability threshold, and mark the matching labels of each candidate target and each attribute word according to the comparison result.
[0200] In the embodiment of the present application, the acquisition unit 601 acquires a first training sample, where the first training sample includes graphic-text pair data. The graphic-text pair data includes a first training image and text description information of the first training image. The visual encoding unit 602 invokes an attribute recognition model to perform visual encoding on each candidate target included in the first training image, and obtains visual encoding features of each candidate target. The text encoding unit 603 performs text encoding on each attribute word included in the text description information, and obtains text encoding features of each attribute word. The determination unit 604 determines a matching result between each candidate target and each attribute word based on the visual encoding features of each candidate target and the text encoding features of each attribute word. The matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word. The model training unit 605 trains the attribute recognition model in the direction of reducing the difference between the determined matching result of each candidate target and each attribute word and the matching label of the corresponding candidate target and the corresponding attribute word in the annotation, and obtains a trained attribute recognition model, which can implement an open-set oriented attribute recognition task.
[0201] Please refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device in the embodiment of the present application includes structures such as a power supply module, and includes a processor 701, a storage device 702, and a communication interface 703. Data can be exchanged between the processor 701, the storage device 702, and the communication interface 703, and the processor 701 implements the training method of the corresponding attribute recognition model.
[0202] The storage device 702 may include a volatile memory, such as a random-access memory (RAM); the storage device 702 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the storage device 702 may further include a combination of the above types of memories.
[0203] The processor 701 may be a central processing unit (CPU). The processor 701 may also be a combination of a CPU and a GPU. Multiple CPUs and GPUs may be included as needed to perform the training of the corresponding attribute recognition model. In one embodiment, the storage device 702 is used to store program instructions. The processor 701 may invoke the program instructions to implement various methods involved in the above embodiments of the present application.
[0204] In a first possible implementation, the processor 701 of the computer device calls the program instructions stored in the storage device 702 to obtain a first training sample; wherein, the first training sample includes image-text pair data, and the image-text pair data includes a first training image and text description information of the first training image; calls an attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain visual encoding features of each candidate target; performs text encoding on each attribute word included in the text description information to obtain text encoding features of each attribute word; determines a matching result between each candidate target and each attribute word based on the visual encoding features of each candidate target and the text encoding features of each attribute word; wherein, the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word; trains the attribute recognition model in a direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate targets and corresponding attribute words that are labeled, to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on an image to be recognized.
[0205] In one embodiment, the processor 701 is further configured to perform the following operations:
[0206] Obtain a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate a target attribute word of a target included in the second training image;
[0207] Crop a local image from the second training image based on the attribute annotation label; wherein, the local image includes the target described by the target attribute word indicated by the attribute label;
[0208] Call a pre-trained attribute recognition model to perform visual encoding on the local image to obtain visual encoding features of the local image;
[0209] Perform text encoding on the target attribute word to obtain text encoding features of the target attribute word;
[0210] Obtain a matching probability between the local image and the target attribute word based on the visual encoding features of the local image and the text encoding features of the target attribute word;
[0211] Optimize the pre-trained attribute recognition model with the matching probability being higher than a first probability threshold as an optimization target to obtain the attribute recognition model.
[0212] In one embodiment, when the processor 701 crops a local image from the second training image based on the attribute tag, the following operations may be performed:
[0213] Perform image slicing on the second training image to obtain candidate images; wherein, the candidate images at least contain the targets described by the attribute words indicated by the attribute tag, and the size of the candidate images is larger than the size of the detection box of the targets in the second training image;
[0214] Randomly crop the candidate images to obtain the local images; wherein, the size of the local images is smaller than the size of the candidate images, and the local images contain part or all of the content of the targets described by the target attribute words.
[0215] In one embodiment, the pre-trained attribute recognition model includes a visual encoder and a text encoder. The visual encoder is used to perform visual encoding on the local images, and the text encoder is used to perform text encoding on the target attribute words;
[0216] When the processor 701 optimizes the pre-trained attribute recognition model with the matching probability higher than the first probability threshold as the optimization goal to obtain the attribute recognition model, the following operations may be performed:
[0217] Optimize the text encoder with the matching probability higher than the first probability threshold as the optimization goal to obtain the attribute recognition model, and the attribute recognition model includes the visual encoder and the optimized text encoder.
[0218] In one embodiment, when the processor 701 obtains the matching probability between the local image and the attribute word indicated by the attribute annotation label based on the visual encoding feature of the local image and the text encoding feature of the target attribute word, the following operations may be performed:
[0219] Obtain the similarity between the visual encoding feature of the local image and the text encoding feature of the target attribute word;
[0220] Based on the similarity, obtain the matching probability between the local image and the target attribute word; wherein, the matching probability between the local image and the target attribute word is positively correlated with the similarity.
[0221] In one embodiment, the processor 701 is further configured to perform the following operations:
[0222] Call the attribute recognition model to perform visual encoding on the first training image to obtain the visual encoding feature of the first training image;
[0223] Perform text encoding on the text description information to obtain the text encoding features of the text description information;
[0224] Based on the visual encoding features of the first training image and the text encoding features of the text description information, determine the matching result between the first training image and the text description information;
[0225] When the processor 701 trains the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation to obtain the trained attribute recognition model, the following operations can be performed:
[0226] Train the attribute recognition model in the direction of reducing the difference between the determined matching result of the first training image and the text description information and the preset matching label of the first training image and the text description information, as well as the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the preset matching label of the first training image and the text description information is used to indicate that the first training image and the text description information have a matching relationship.
[0227] In one embodiment, the processor 701 is further configured to perform the following operations:
[0228] Call the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection frame in the first training image, and the target detection frame refers to: the detection frame with the largest size among the detection frames of at least one candidate target included in the first training image;
[0229] Based on the visual encoding features of the target unit image and the text encoding features of each attribute word, determine the matching results between the target unit image and each attribute word;
[0230] When the processor 701 trains the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation to obtain the trained attribute recognition model, the following operations can be performed:
[0231] Train the attribute recognition model in the direction of reducing the difference between the determined matching result of the target unit image and each attribute word and the preset matching label of the target unit image and each attribute word, and the difference between the matching result of each candidate target and each attribute word and the matching label of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the preset matching label of the target unit image and each attribute word is used to indicate that the target unit image and each attribute word have a matching relationship.
[0232] In one embodiment, the processor 701 is further configured to perform the following operations:
[0233] Call the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding feature of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image;
[0234] Perform text encoding on each noun phrase included in the text description information to obtain the text encoding feature of each noun phrase;
[0235] Based on the visual encoding feature of the target unit image and the text encoding feature of each noun phrase, determine the matching result of the target unit image and each noun phrase;
[0236] When the processor 701 trains the attribute recognition model in the direction of reducing the difference between the determined matching result of each candidate target and each attribute word and the matching label of the corresponding candidate target and corresponding attribute word in the annotation to obtain the trained attribute recognition model, it may perform the following operations:
[0237] Train the attribute recognition model in the direction of reducing the difference between the determined matching result of the target unit image and each noun phrase and the preset matching label of the target unit image and each noun phrase, and the difference between the matching result of each candidate target and each attribute word and the matching label of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the preset matching label of the target unit image and each noun phrase is used to indicate that the target unit image and each noun phrase have a matching relationship.
[0238] In one embodiment, the processor 701 is further configured to perform the following operations:
[0239] Invoke the attribute recognition model, and based on the visual coding features of the respective candidate targets and the text coding features of the respective attribute words, obtain the matching probabilities between the respective candidate targets and the respective attribute words;
[0240] Compare the matching probabilities between the respective candidate targets and the respective attribute words with a second probability threshold, and label the matching tags between the respective candidate targets and the respective attribute words according to the comparison results.
[0241] In an embodiment of the present application, the processor 701 obtains a first training sample, where the first training sample includes graphic-text pair data, and the graphic-text pair data includes a first training image and text description information of the first training image. Invoke the attribute recognition model to perform visual coding on each candidate target included in the first training image to obtain the visual coding features of each candidate target, perform text coding on each attribute word included in the text description information to obtain the text coding features of each attribute word, and based on the visual coding features of each candidate target and the text coding features of each attribute word, determine the matching results between each candidate target and each attribute word. The matching results are used to indicate whether there is a matching relationship between each candidate target and each attribute word. Train the attribute recognition model in the direction of reducing the difference between the determined matching results between each candidate target and each attribute word and the matching tags of the corresponding candidate target and the corresponding attribute word, and obtain the trained attribute recognition model, which can implement the open-set oriented attribute recognition task.
[0242] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. The computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, applications required for at least one function, etc.; the storage data area can store data created according to the use of the blockchain node, etc.
[0243] The above-disclosed are only some embodiments of the present application. Of course, the scope of rights of the present application cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present invention.
Claims
1. A training method for an attribute recognition model, characterized in that, Including: Obtain a first training sample; wherein, the first training sample includes text-image pair data, and the text-image pair data includes a first training image and text description information of the first training image; Call an attribute recognition model to perform visual encoding on each candidate target included in the first training image to obtain visual encoding features of each candidate target; Perform text encoding on each attribute word included in the text description information to obtain text encoding features of each attribute word; Based on the visual encoding features of each candidate target and the text encoding features of each attribute word, determine matching results of each candidate target and each attribute word; wherein, the matching results are used to indicate whether each candidate target and each attribute word have a matching relationship; Train the attribute recognition model in a direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word that are labeled, to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on an image to be recognized.
2. The method according to claim 1, wherein The method further includes: Obtain a second training sample, the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate a target attribute word of a target included in the second training image; Crop a local image from the second training image based on the attribute annotation label; wherein, the local image includes the target described by the target attribute word indicated by the attribute label; Call a pre-trained attribute recognition model to perform visual encoding on the local image to obtain visual encoding features of the local image; Perform text encoding on the target attribute word to obtain text encoding features of the target attribute word; Based on the visual encoding features of the local image and the text encoding features of the target attribute word, obtain a matching probability between the local image and the target attribute word; Optimize the pre-trained attribute recognition model with the matching probability being higher than a first probability threshold as an optimization target to obtain the attribute recognition model.
3. The method according to claim 2, wherein The cropping the local image from the second training image based on the attribute label includes: Perform image slicing on the second training image to obtain candidate images; wherein, each candidate image includes at least the target described by the attribute word indicated by the attribute label, and the size of the candidate image is larger than the size of the detection frame of the target in the second training image; Randomly crop the candidate image to obtain the local image; wherein, the size of the local image is smaller than the size of the candidate image, and the local image includes part or all of the content of the target described by the target attribute word.
4. The method according to claim 2, characterized in that, The pre-trained attribute recognition model includes a visual encoder and a text encoder, the visual encoder is used to perform visual encoding on the local image, and the text encoder is used to perform text encoding on the target attribute word; Optimizing the pre-trained attribute recognition model with the optimization objective that the matching probability is higher than the first probability threshold to obtain the attribute recognition model, including: Optimizing the text encoder with the optimization objective that the matching probability is higher than the first probability threshold to obtain the attribute recognition model, where the attribute recognition model includes the visual encoder and the optimized text encoder.
5. The method according to claim 2, wherein Based on the visual encoding features of the local image and the text encoding features of the target attribute word, obtaining the matching probability between the local image and the attribute word indicated by the attribute annotation label, including: Obtaining the similarity between the visual encoding features of the local image and the text encoding features of the target attribute word; Based on the similarity, obtaining the matching probability between the local image and the target attribute word; wherein, the matching probability between the local image and the target attribute word is positively correlated with the similarity.
6. The method according to any one of claims 1-5, characterized in that The method further includes: Invoking the attribute recognition model to perform visual encoding on the first training image to obtain the visual encoding features of the first training image; Performing text encoding on the text description information to obtain the text encoding features of the text description information; Based on the visual encoding features of the first training image and the text encoding features of the text description information, determining the matching result between the first training image and the text description information; Training the attribute recognition model in the direction of reducing the difference between the determined matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation to obtain the trained attribute recognition model, including: Training the attribute recognition model in the direction of reducing the difference between the determined matching result of the first training image and the text description information and the preset matching label of the first training image and the text description information, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation to obtain the trained attribute recognition model; wherein, the preset matching label of the first training image and the text description information is used to indicate that the first training image and the text description information have a matching relationship.
7. The method according to any one of claims 1-5, characterized in that, The method further includes: Invoking the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image; Based on the visual encoding features of the target unit image and the text encoding features of each attribute word, determining the matching results between the target unit image and each attribute word; Training the attribute recognition model in the direction of reducing the difference between the matching results of each determined candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain a trained attribute recognition model, includes: Training the attribute recognition model in the direction of reducing the difference between the matching results of the determined target unit image and each attribute word and the matching labels of the preset target unit image and each attribute word, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the matching labels of the preset target unit image and each attribute word are used to indicate that the target unit image and each attribute word have a matching relationship.
8. The method according to any one of claims 1 to 5, characterized in that The method further includes: Invoking the attribute recognition model to perform visual encoding on the target unit image to obtain the visual encoding features of the target unit image; wherein, the target unit image refers to the unit image corresponding to the target detection box in the first training image, and the target detection box refers to: the detection box with the largest size among the detection boxes of at least one candidate target included in the first training image; Performing text encoding on each noun phrase included in the text description information to obtain the text encoding features of each noun phrase; Determining the matching results of the target unit image and each noun phrase based on the visual encoding features of the target unit image and the text encoding features of each noun phrase; Training the attribute recognition model in the direction of reducing the difference between the matching results of each determined candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain a trained attribute recognition model, includes: Training the attribute recognition model in the direction of reducing the difference between the matching results of the determined target unit image and each noun phrase and the matching labels of the preset target unit image and each noun phrase, and the difference between the matching results of each candidate target and each attribute word and the matching labels of the corresponding candidate target and corresponding attribute word in the annotation, to obtain the trained attribute recognition model; wherein, the matching labels of the preset target unit image and each noun phrase are used to indicate that the target unit image and each noun phrase have a matching relationship.
9. The method according to any one of claims 1-5, characterized in that, The method further includes: Invoking the attribute recognition model to obtain the matching probabilities of each candidate target and each attribute word based on the visual encoding features of each candidate target and the text encoding features of each attribute word; Comparing the matching probabilities of each candidate target and each attribute word with a second probability threshold, and annotating the matching labels of each candidate target and each attribute word according to the comparison result.
10. A training method for an attribute recognition model, characterized in that, Includes: Obtain a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate a target attribute word of a target included in the second training image; Based on the attribute annotation label, crop a local image from the second training image; wherein, the local image includes the target described by the target attribute word indicated by the attribute label; Call a pre-trained attribute recognition model to perform visual encoding on the local image to obtain a visual encoding feature of the local image; Perform text encoding on the target attribute word to obtain a text encoding feature of the target attribute word; Based on the visual encoding feature of the local image and the text encoding feature of the target attribute word, obtain a matching probability between the local image and the target attribute word; Taking the matching probability being higher than a first probability threshold as an optimization objective, optimize the pre-trained attribute recognition model to obtain an attribute recognition model; wherein, the attribute recognition model is used to perform attribute recognition on an image to be recognized.
11. The method according to claim 10, wherein The cropping the local image from the second training image based on the attribute label includes: Perform image slicing on the second training image to obtain a candidate image; wherein, the candidate image at least includes the target described by the attribute word indicated by the attribute label, and the size of the candidate image is larger than the size of the detection frame of the target in the second training image; Randomly crop the candidate image to obtain the local image; wherein, the size of the local image is smaller than the size of the candidate image, and the local image includes part or all of the content of the target described by the target attribute word.
12. The method according to claim 10, wherein The pre-trained attribute recognition model includes a visual encoder and a text encoder, the visual encoder is used to perform visual encoding on the local image, and the text encoder is used to perform text encoding on the target attribute word; The optimizing the pre-trained attribute recognition model to obtain the attribute recognition model with the matching probability being higher than a first probability threshold includes: Taking the matching probability being higher than a first probability threshold as an optimization objective, optimize the text encoder to obtain the attribute recognition model, and the attribute recognition model includes the visual encoder and the optimized text encoder.
13. The method according to claim 10, wherein The obtaining the matching probability between the local image and the attribute word indicated by the attribute annotation label based on the visual encoding feature of the local image and the text encoding feature of the target attribute word includes: Obtain a similarity between the visual encoding feature of the local image and the text encoding feature of the target attribute word; Based on the similarity, obtain the matching probability between the local image and the target attribute word; wherein, the matching probability between the local image and the target attribute word is positively correlated with the similarity.
14. A training device for an attribute recognition model, characterized in that The device includes: An obtaining unit, configured to obtain a first training sample; wherein, the first training sample includes graphic-text pair data, and the graphic-text pair data includes a first training image and text description information of the first training image; A visual encoding unit, configured to call an attribute recognition model to perform visual encoding on each candidate target included in the first training image, so as to obtain visual encoding features of each candidate target; A text encoding unit, configured to perform text encoding on each attribute word included in the text description information, so as to obtain text encoding features of each attribute word; A determination unit, configured to determine a matching result between each candidate target and each attribute word based on the visual encoding features of each candidate target and the text encoding features of each attribute word; wherein, the matching result is used to indicate whether there is a matching relationship between each candidate target and each attribute word; A model training unit, configured to train the attribute recognition model in a direction of reducing the difference between the determined matching result between each candidate target and each attribute word and the matching label of the corresponding candidate target and the corresponding attribute word in the annotation, so as to obtain a trained attribute recognition model; wherein, the trained attribute recognition model is used to perform attribute recognition on an image to be recognized.
15. A training device for an attribute recognition model, characterized in that, The apparatus includes: An acquisition unit, configured to acquire a second training sample, where the second training sample includes a second training image and an attribute annotation label of the second training image, and the attribute label is used to indicate a target attribute word of a target included in the second training image; A cropping unit, configured to crop a local image from the second training image based on the attribute annotation label; wherein, the local image includes the target described by the target attribute word indicated by the attribute label; A visual encoding unit, configured to call a pre-trained attribute recognition model to perform visual encoding on the local image, so as to obtain visual encoding features of the local image; A text encoding unit, configured to perform text encoding on the target attribute word, so as to obtain text encoding features of the target attribute word; A determination unit, configured to obtain a matching probability between the local image and the target attribute word based on the visual encoding features of the local image and the text encoding features of the target attribute word; A model training unit, configured to optimize the pre-trained attribute recognition model with the matching probability being higher than a first probability threshold as an optimization target, so as to obtain an attribute recognition model; wherein, the attribute recognition model is used to perform attribute recognition on an image to be recognized.
16. A computer device, characterized in that, The computer device includes a processor, a storage device, and a communication interface, and the processor, the storage device, and the communication interface are connected to each other, wherein: The storage device is configured to store a computer program, and the computer program includes program instructions; The processor is configured to call the program instructions to execute the training method of the attribute recognition model according to any one of claims 1 to 13.
17. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a processor, they are used to execute the training method of the attribute recognition model according to any one of claims 1 to 13.
18. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is suitable for being loaded and executed by a processor to execute the training method of the attribute recognition model according to any one of claims 1 to 13.