Cross-domain Person Search Method Based on Text Description
By constructing a cross-domain character search network model and using gradient reverse layer to adjust feature distribution, the problem of lack of labeled data in character search based on text description is solved, and cross-domain search capabilities and target detection effects are improved.
Patent Information
- Application Number
- CN202211426869.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-11-15
AI Technical Summary
The character search task based on text description lacks large-scale annotation data, resulting in poor retrieval results and lacks methods for cross-data set effectiveness.
Using a cross-domain character search method based on text description, by constructing a cross-domain character search network model, two gradient inverse layers are used to make the text feature extractor and image feature extractor tend to be similar between the source domain and the target domain, thereby improving cross-domain search capabilities.
In the absence of labeled data, the target detection capability of the cross-domain character search network model is improved, the demand for labeled data is reduced, the dependence on human resources is reduced, and the cross-domain character search capability is realized.
Smart Images

Figure CN115712751B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer recognition, and more specifically, to a cross-domain person search method based on text description. Background Art
[0002] Thanks to the rapid development of deep learning, great progress has been made in the field of computer vision. However, the training of deep learning models mostly relies on a large amount of labeled data to achieve ideal training effects, which undoubtedly restricts the wide application of deep learning. To solve this problem, domain adaptation attempts to transfer the model from the domain rich in labeled data (source domain) to the domain lacking labeled data (target domain), thus solving this problem. The person search task based on text description aims to retrieve the target person that matches the text description from the image library through descriptive text. Solving such a fine-grained cross-modal retrieval task is extremely challenging in itself, and this task is further hindered due to the lack of large-scale datasets. Therefore, domain adaptation is considered to solve this problem.
[0003] Person search based on text description is a challenging task because it requires accurately capturing fine-grained features for recognition, such as the clothes a person wears and the accessories they wear.
[0004] Divided according to the number of models, the current work is mainly divided into two methods: (1) Single-stream model: The image and text are simultaneously input into a single model, and then the matching degree between the two is calculated and compared, and the return result is selected according to the level of the matching degree. (2) Two-stream model: The image and text are respectively input into two models to obtain the semantic information of the image and the semantic information of the text, and then the matching score is calculated to obtain the return result.
[0005] If divided according to the matching area, the current work can be roughly divided into two categories: (1) Global matching method: The global features of the image and text are obtained through a convolutional neural network and then matched. In some works, the image and text are embedded into a shared image-text space. (2) Local matching method: Focus on the local features of the image and text to capture subtle but discriminative semantic information.
[0006] There are mainly three different methods for domain adaptation: (1) Sample adaptation: The basic idea of this method is to resample the source domain samples, so that the feature distribution of the resampled source domain samples tends to be the same as that of the target domain samples, and then re-learn the classifier on the resampled sample set. This method is applicable to the case where the distribution difference between the source domain and the target domain is small. (2) Feature adaptation: Its basic idea is to project the sample features in the source domain and the target domain into a common feature space and learn the common feature representation. In the common feature space, the distributions of the source domain and the target domain should be as similar as possible. This method is applicable to the case where there are certain differences between the source domain and the target domain. (3) Model adaptation: Its basic idea is to directly perform adaptation at the model level. There are two ideas for the method of model adaptation. One is to directly build a model, but add a constraint of "close inter-domain distance" to the model; the other is to use an iterative method to gradually classify the samples in the target domain, add the samples with high confidence to the training set, and update the model. This method is applicable to the case where the difference between the source domain and the target domain is relatively large.
[0007] Due to the huge annotation cost, person search based on text description lacks a large number of annotated samples for training, and domain adaptation exactly fits the problems it is facing. So far, no relevant work has been found on applying domain adaptation to the problem of person search based on text description.
[0008] Therefore, the main problems in the existing technology are as follows: Person search based on text description is a fine-grained cross-modal retrieval task, and the annotation work of its dataset is complex and cumbersome. Therefore, there is a lack of large-scale datasets for relevant model training, resulting in very poor retrieval effects based on text description. In the case of lack of data, there is currently no method that can make the retrieval model have cross-dataset effectiveness. Summary of the Invention
[0009] In order to solve the above problems of the deficiencies and defects in the prior art, the present invention provides a cross-domain person search method based on text description, which can have the ability of cross-domain person search in the case of lack of data.
[0010] To achieve the above object of the present invention, the following technical solutions are adopted:
[0011] A cross-domain person search method based on text description, the method includes the following steps:
[0012] Construct a cross-domain person search network model based on text description, the cross-domain person search network model includes an image feature extractor for extracting image features, a text feature extractor for extracting text features, a first gradient reversal layer for gradient descent, a second gradient reversal layer for gradient descent, an image domain classifier, and a text domain classifier;
[0013] The described image feature extractor inputs the extracted image features into the image domain classifier through a first gradient reversal layer;
[0014] The image domain classifier processes the image features to obtain an image domain label, and calculates an image domain classification loss based on the image domain label;
[0015] The text feature extractor inputs the extracted text features into the text domain classifier through a second gradient reversal layer;
[0016] The text domain classifier processes the text features to obtain a text domain label, and calculates a text domain classification loss based on the text domain label;
[0017] Use the trained cross-domain person search network model to perform person search based on text description on the target domain.
[0018] Preferably, the method for specifically training the cross-domain person search network model is as follows:
[0019] Input source domain samples and target domain samples into the cross-domain person search network model for training at the same time;
[0020] For a source domain sample containing a picture and its description text segment, after the source domain sample is input into the cross-domain person search network model, a first image feature and a first text feature are obtained, and a contrast loss between the first image feature and the first text feature is calculated, so that the feature contrast loss between the matching image feature and the text segment is less than a first threshold;
[0021] At the same time, the first image feature and the first text feature are respectively input into the image domain classifier and the text domain classifier through the gradient reversal layer to obtain a first image domain label and a first text domain label, and a first image domain classification loss and a first text domain classification loss are respectively calculated, and then gradient descent is performed to update the network parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier, so as to maximize the first image domain classification loss and the first text domain classification loss of the source domain sample;
[0022] For a target domain sample, after inputting into the cross-domain person search network model, a second image feature and a second text feature are obtained. The second image feature and the second text feature are directly input into the image domain classifier and the text domain classifier through the gradient reversal layer to obtain a second image domain label and a second text domain label, and a second image domain classification loss and a second text domain classification loss are respectively calculated, and then gradient descent is performed to update the network parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier, so as to maximize the second sample image domain classification loss and the second text domain classification loss of the target domain;
[0023] By simultaneously inputting source domain samples and target domain samples to train the cross-domain person search network model, the image feature distributions and text feature distributions of the source domain samples and target domain samples tend to be similar, that is, the distance between the source domain and the target domain is less than the second threshold.
[0024] Further, the calculation formula of the contrastive loss is:
[0025]
[0026]
[0027]
[0028] Among them, is the normalized image feature; is the normalized text feature; n is the batch size, and T is the temperature coefficient; the obtained L c is the contrastive loss; li is a matrix used to represent the matching degree between each image feature and each text feature in a batch; lw is a matrix used to represent the matching degree between each text feature and each image feature; li ii and lw ii represent the elements on the diagonal of the matrix, because the elements on the diagonal are the feature matching degrees of the corresponding pictures and texts.
[0029] Further, the empirical estimation of H-divergence is used to represent the distance between the source domain and the target domain:
[0030]
[0031] Among them, I[a] is the indicator function, which outputs 1 if the content in the brackets is true, otherwise it returns 0; η is a binary classification function that outputs the domain classification result after inputting the features; n represents the number of source domain samples, n′ represents the number of target domain samples, and N is the sum of the number of source domain samples and the number of target domain samples.
[0032] Furthermore, if the image domain classification label and text domain classification label of the source domain are both 0, the calculation methods of the image domain classification loss and text domain classification loss of the source domain samples are:
[0033]
[0034] Among them, c is the domain classification result output by the classifier.
[0035] Furthermore, due to the effect of the gradient reversal layer, the gradients backpropagated to the image classifier and text classifier are both:
[0036]
[0037] The gradients backpropagated to the image feature extractor and the text feature extractor are respectively:
[0038]
[0039] Among them, θ c is the coefficient of the classifier, θ f is the coefficient of the feature extractor, λ is the regularization parameter used to prevent overfitting, and L d represents the domain classification loss.
[0040] Preferably, the image feature extractor adopts Visual Transformer, loads the pre-trained model vit_base_patch16_384 into the image feature extractor, and the image features output by the image feature extractor are tensors of size batch_size*768.
[0041] Preferably, the text feature extractor adopts Bert, loads the pre-trained model bert-base-uncased into the text feature extractor; the text features output by the text feature extractor are tensors of size batch_size*768.
[0042] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the cross-domain person search method based on text description as described are implemented.
[0043] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the cross-domain person search method based on text description as described are implemented.
[0044] The beneficial effects of the present invention are as follows:
[0045] A cross-domain person search method based on text description provided by the present invention, through two gradient reverse layers, makes the feature distributions extracted by the text feature extractor and the image feature extractor from source domain samples and target domain samples tend to be similar, thereby improving the object detection ability of the cross-domain person search network model in the unlabeled domain, reducing the demand of the cross-domain person search network model for labeled data, reducing the dependence on human resources, and having the cross-domain person search ability in the case of lack of labeled data. Description of the Drawings
[0046] Figure 1 is a flowchart of the steps of the cross-domain person search method based on text description of the present invention.
[0047] Figure 2It is the principle block diagram of the cross - domain person search network model based on text description of the present invention. Detailed implementation manners
[0048] The present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0049] Embodiment 1
[0050] As Figure 1 shown, a cross - domain person search method based on text description, the method includes the following steps:
[0051] Construct a cross - domain person search network model, the cross - domain person search network model includes an image feature extractor for extracting image features, a text feature extractor for extracting text features, a first gradient reversal layer for gradient descent, a second gradient reversal layer for gradient descent, an image domain classifier, and a text domain classifier;
[0052] The image feature extractor inputs the extracted image features into the image domain classifier through the first gradient reversal layer;
[0053] The image domain classifier processes the image features to obtain an image domain label, and calculates an image domain classification loss according to the image domain label;
[0054] The text feature extractor inputs the extracted text features into the text domain classifier through the second gradient reversal layer;
[0055] The text domain classifier processes the text features to obtain a text domain label, and calculates a text domain classification loss according to the text domain label;
[0056] Use the trained cross - domain person search network model to perform text - description - based person search on the target domain.
[0057] In this embodiment, using the trained cross - domain person search network model to perform text - description - based person search on the target domain is as follows: The input text enters the text feature extractor to obtain the features of the text, and then all images are input into the image feature extractor to obtain image features. The text feature vector and the image feature matrix are multiplied to obtain the matching degree of each image feature with respect to the text feature. The image corresponding to the image feature with the highest matching degree is the search result, thereby realizing finding the image that best matches the text.
[0058] In a specific embodiment, the method for specifically training the cross - domain person search network model is as follows:
[0059] Input source domain samples and target domain samples into the cross - domain person search network model for training at the same time;
[0060] For a source domain sample containing an image and its descriptive text segment, after the source domain sample is input into the cross-domain person search network model, a first image feature and a first text feature are obtained, and the contrast loss between the first image feature and the first text feature is calculated, such that the contrast loss between the matching image feature and the text segment feature is less than a first threshold;
[0061] At the same time, the first image feature and the first text feature are respectively input into an image domain classifier and a text domain classifier through a gradient reversal layer to obtain a first image domain label and a first text domain label, and the first image domain classification loss and the first text domain classification loss are respectively calculated, and then gradient descent is performed to update the network parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier, so as to maximize the first image domain classification loss and the first text domain classification loss of the source domain sample;
[0062] For a target domain sample, after being input into the cross-domain person search network model, a second image feature and a second text feature are obtained. The second image feature and the second text feature are directly input into an image domain classifier and a text domain classifier through a gradient reversal layer to obtain a second image domain label and a second text domain label, and the second image domain classification loss and the second text domain classification loss are respectively calculated, and then gradient descent is performed to update the network parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier, so as to maximize the second sample image domain classification loss and the second text domain classification loss of the target domain.
[0063] By simultaneously inputting source domain samples and target domain samples to train the cross-domain person search network model, the image feature distributions and text feature distributions of the source domain samples and the target domain samples are made to tend to be similar, that is, the distance between the source domain and the target domain is made less than a second threshold.
[0064] In this embodiment, the source domain has labeled training samples, for example, a pair of text and image for a certain person; the target domain samples are unlabeled training samples. This embodiment hopes that the trained cross-domain person search network model can also achieve good results on the target domain; the contrast loss is to calculate the features between text and image and is only used for source domain samples. The text domain classifier is a domain classifier for text and is used to judge whether the input text belongs to the source domain or the target domain during training; the image domain classifier is a domain classifier for images and is used to judge whether the input image belongs to the source domain or the target domain during training.
[0065] In this embodiment, the purpose of calculating the contrast loss for source domain samples is to make the text corresponding to the input image match, so that the image feature extractor and the text feature extractor extract matching image features and text features with corresponding distributions.
[0066] In this embodiment, after obtaining the first image feature and the first text feature, the contrast loss between the two is calculated and then gradient descent is performed, so that the features of the matching image and text segment tend to be similar, thereby achieving the purpose of training a cross-domain person search based on text description using the source domain.
[0067] In a specific embodiment, the calculation formula of the contrast loss is as follows:
[0068]
[0069]
[0070]
[0071] where is the normalized image feature; is the normalized text feature; n is the batch size, and T is the temperature coefficient; the obtained L c is the contrast loss; li is a matrix used to represent the matching degree between each image feature and each text feature in a batch; lw is a matrix used to represent the matching degree between each text feature and each image feature; li ii and lw ii represent the elements on the diagonal of the matrix, because the elements on the diagonal are the feature matching degrees of the corresponding pictures and texts.
[0072] In a specific embodiment, the image feature and the text feature respectively pass through the gradient reversal layer and enter the image domain classifier and the text domain classifier to obtain the image domain label and the text domain label. After calculating their respective losses, gradient descent is performed, thereby narrowing the distance between the source domain and the target domain. The purpose of this embodiment is to minimize the distance between the source domain and the target domain. This embodiment uses empirical estimation of the H-divergence to represent the distance between the source domain and the target domain:
[0073]
[0074] where I[a] is the indicator function, which outputs 1 if the content in the parentheses is true, otherwise returns 0; η is a binary classification function that outputs the domain classification result after inputting the feature; n represents the number of source domain samples, n′ represents the number of target domain samples, and N is the sum of the number of source domain samples and the number of target domain samples.
[0075] In a specific embodiment, if the image domain classification label and the text domain classification label of the source domain are both 0, the calculation methods of the image domain classification loss and the text domain classification loss of the source domain samples are as follows:
[0076]
[0077] Among them, c is the domain classification result output by the classifier.
[0078] In a specific embodiment, due to the role of the gradient reversal layer, the gradients backpropagated to the image classifier and the text classifier are both:
[0079]
[0080] The gradients backpropagated to the image feature extractor and the text feature extractor are respectively:
[0081]
[0082] Among them, θ c is the coefficient of the classifier, θ f is the coefficient of the feature extractor, λ is the regularization parameter used to prevent overfitting, and L d represents the domain classification loss.
[0083] In this embodiment, the parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier are updated in reverse by gradient descent through the gradient reversal layer, so that the domain classification error of the cross-domain person search network model will continue to increase.
[0084] In a specific embodiment, the image feature extractor uses Visual Transformer, and the image feature extractor is loaded into the pre-trained model vit_base_patch16_384. The image features output by the image feature extractor are tensors of size batch_size*768.
[0085] In a specific embodiment, the text feature extractor uses Bert, and the text feature extractor is loaded into the pre-trained model bert-base-uncased; the text features output by the text feature extractor are tensors of size batch_size*768.
[0086] Generally, the pre-trained model has been effectively trained on a specific dataset and has a good effect on feature extraction. Loading the pre-trained model can reduce the training time cost and have a better training effect.
[0087] Embodiment 2
[0088] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the cross-domain person search method based on text description as described in Embodiment 1.
[0089] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0090] Embodiment 2
[0091] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the cross-domain person search method based on text description as described in Embodiment 1 are implemented.
[0092] That is, those skilled in the art can understand that all or part of the steps of implementing the method in the above embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which are various media that can store program codes.
[0093] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A cross - domain person search method based on text description, characterized in that: The method described above includes the following steps: Construct a cross-domain person search network model based on text descriptions. The cross-domain person search network model includes an image feature extractor for extracting image features, a text feature extractor for extracting text features, a first gradient reversal layer for gradient descent, a second gradient reversal layer for gradient descent, an image domain classifier, and a text domain classifier; The image feature extractor inputs the extracted image features into the image domain classifier through the first gradient reversal layer; The image domain classifier processes the image features to obtain an image domain label and calculates an image domain classification loss based on the image domain label; The text feature extractor inputs the extracted text features into the text domain classifier through the second gradient reversal layer; The text domain classifier processes the text features to obtain a text domain label and calculates a text domain classification loss based on the text domain label; Use the trained cross-domain person search network model to perform a person search based on text descriptions in the target domain; The specific method for training the cross-domain person search network model is as follows: Input source domain samples and target domain samples into the cross-domain person search network model for training at the same time; For a source domain sample containing a picture and its description text segment, after the source domain sample is input into the cross-domain person search network model, obtain a first image feature and a first text feature, and calculate a contrast loss between the first image feature and the first text feature, so that the contrast loss between the matching image feature and the text segment feature is less than a first threshold; At the same time, input the first image feature and the first text feature into the image domain classifier and the text domain classifier respectively through the gradient reversal layer to obtain a first image domain label and a first text domain label, calculate a first image domain classification loss and a first text domain classification loss respectively, and then perform gradient descent to update the network parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier, so as to maximize the first image domain classification loss and the first text domain classification loss of the source domain samples; For target domain samples, after inputting them into the cross-domain person search network model, obtain a second image feature and a second text feature, directly input the second image feature and the second text feature into the image domain classifier and the text domain classifier respectively through the gradient reversal layer to obtain a second image domain label and a second text domain label, calculate a second image domain classification loss and a second text domain classification loss respectively, and then perform gradient descent to update the network parameters of the image feature extractor, the text feature extractor, the image domain classifier, and the text domain classifier, so as to maximize the second sample image domain classification loss and the second text domain classification loss of the target domain; Train the cross-domain person search network model by simultaneously inputting source domain samples and target domain samples, so that the image feature distributions and text feature distributions of the source domain samples and the target domain samples tend to be similar, that is, the distance between the source domain and the target domain is less than a second threshold.
2. The cross - domain person search method based on text description according to claim 1, characterized in that: The calculation formula for the contrast loss is as follows: Among them, is the normalized image feature; is the normalized text feature; n is the batch size, and T is the temperature coefficient; the obtained is the contrastive loss; is a matrix used to represent the matching degree between each image feature and each text feature in a batch; is a matrix used to represent the matching degree between each text feature and each image feature; and represent the elements on the diagonal of the matrix, because the elements on the diagonal are the feature matching degrees of the corresponding pictures and texts.
3. The cross - domain person search method based on text description according to claim 1, characterized in that: Use empirical estimation of H-divergence to represent the distance between the source domain and the target domain: Among them, is an indicator function that outputs 1 if the content in the parentheses is true, otherwise returns 0; is a binary classification function that outputs the domain classification result after inputting features; represents the number of source domain samples, represents the number of target domain samples, N is the sum of the number of source domain samples and the number of target domain samples.
4. The cross - domain person search method based on text description according to claim 2, characterized in that: If the classification labels of the image domain and the text domain in the source domain are both 0, the calculation methods of the classification losses of the image domain and the text domain of the source domain samples are as follows: Among them, c is the domain classification result output by the classifier.
5. The cross - domain person search method based on text description according to claim 4, characterized in that: Due to the effect of the gradient reversal layer, the gradients backpropagated to the image classifier and the text classifier are both: The gradients backpropagated to the image feature extractor and the text feature extractor are respectively: Among them, is the coefficient of the classifier, is the coefficient of the feature extractor, is the regularization parameter for preventing overfitting, represents the domain classification loss.
6. The cross - domain person search method based on text description according to claim 1, characterized in that: The image feature extractor adopts Visual Transformer, and the pre-trained model vit_base_patch16_384 is loaded into the image feature extractor. The image features output by the image feature extractor are tensors of size batch_size*768.
7. The cross-domain person search method based on text description according to claim 1, characterized in that: The text feature extractor adopts Bert, and the pre-trained model bert-base-uncased is loaded into the text feature extractor; the text features output by the text feature extractor are tensors of size batch_size*768.
8. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the cross-domain person search method based on text description according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by the processor, the steps of the cross-domain person search method based on text description according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Information retrieval method and system based on voice recognition and storage medium
CN114282090A
Generating active content for assistant system
CN114930363A