A text description-driven pedestrian search method
By using a dual-stream model of symmetric Transformer in pedestrian search, the image and text features are extracted using the Transformer encoder, combined with cross-modal comparison learning and approximate data generation methods, the problems of slow computing speed and difficult training in the existing technology are solved, and efficient and accurate pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202211138480.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The existing deep learning-based pedestrian search method is slow to process multiple query texts, and the heterogeneity of convolutional neural networks and recurrent neural networks makes training difficult, computational complexity and difficult to deploy.
The dual-stream model of symmetric Transformer is adopted, and the image global features and text global features are extracted through the visual Transformer encoder and the text Transformer encoder, and pedestrian search is realized through cosine similarity calculation, combining cross-modal contrast learning and approximate data generation methods for training, and using the sharpness-aware optimizer SAM for optimization.
It improves the computing efficiency and accuracy of pedestrian search, simplifies the model architecture, reduces the computational complexity, and makes the model easier to deploy and implement.
Smart Images

Figure CN115455226B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer recognition, and more specifically, to a pedestrian search method driven by text description. Background Art
[0002] Text-based pedestrian search refers to a method that takes a text description of a pedestrian as the input of an algorithm, and the algorithm is required to find the specified pedestrian from pictures / videos. Compared with the image-based pedestrian re-identification method (which requires a picture to be given in advance as the input), this method is more flexible and has broad application prospects in video surveillance fields such as searching for suspects or missing children.
[0003] This problem is a fine-grained recognition problem that requires accurate identification of external features such as the clothing and carried items of people in the image. Therefore, existing methods all solve this problem based on deep learning, and there are roughly two categories of methods to solve this problem: (1) Single-stream model: The image and text are simultaneously input into a single model to calculate the matching score between the two. After comparing the text with the images one by one, the k images with the highest similarity to the text are calculated as the return results; (2) Two-stream model: Two independent models are used, one for calculating the semantic information of the image and one for calculating the semantic information of the text, and then the scores of the semantic information of the two modalities are calculated to obtain the k images most similar to the text as the return results. Most methods use convolutional neural networks to extract image semantic information and recurrent neural networks to extract text semantic information.
[0004] In addition, in recent years, in order to improve the recognition accuracy of the algorithm, many methods use image segmentation as the local features of the image, extract noun phrases in the sentence as the local features of the text by using an algorithm, and then design a multi-branch network to match the two local features. There are also methods that introduce other algorithms such as human pose estimation and human instance segmentation as external information to help the neural network model for training.
[0005] Disadvantages of the prior art:
[0006] The disadvantage of the single-stream model is that if there is a large amount of input text (that is, when multiple people need to be queried), the query speed is very slow. For example, if there are N query texts and M images, then NxM calculations are required at this time, while the two-stream model requires N + M calculations, and it has an obvious advantage when N is relatively large.
[0007] Convolutional neural networks are suitable for extracting global features of images, and recurrent neural networks are suitable for extracting text sequence features. However, the calculation methods of these two models are not the same. Therefore, in order to train these two models, additional network structures are often required to implement, and the effect is usually not good.
[0008] The introduction of local features first requires the design of a very complex model structure, which introduces more complex calculations, making the model not easy to deploy and implement. And introducing external information such as human pose information requires another model to predict human information, which also introduces an additional model and is not conducive to the deployment of the model.
[0009] Although the current network models already have a certain degree of accuracy, their performance is still not good and there is still a large room for improvement. Summary of the Invention
[0010] In order to solve the problems of the deficiencies and defects of the above prior art, the present invention provides a text description-driven pedestrian search method, which solves the problem of difficult training caused by the heterogeneity of convolutional neural networks and recurrent neural networks in the prior art by constructing a two-stream model, effectively improves the performance of the two-stream model, and realizes more accurate pedestrian recognition.
[0011] In order to achieve the above object of the present invention, the following technical solutions are adopted:
[0012] A text description-driven pedestrian search method, the method comprising the following steps:
[0013] Construct a two-stream model based on symmetric Transformer, the two-stream model comprising a visual Transformer encoder and a text Transformer encoder;
[0014] For an image, first evenly divide the image, perform a linear mapping on each image patch and add the position encoding as the first encoding vector input to the visual Transformer encoder; the first encoding vector is N1 D-dimensional vectors; wherein, N1 represents the number of image patches;
[0015] For text, first perform vector quantization encoding on each word, then perform a linear mapping of the word and add the position encoding as the second encoding vector input to the text Transformer encoder; the second encoding vector is N2 D-dimensional vectors, wherein, N2 represents the number of words;
[0016] Perform cosine similarity calculation on the image global feature output by the visual Transformer encoder and the text global feature output by the text Transformer encoder one by one, and find the image with the highest cosine similarity to the corresponding text, so as to realize text-based pedestrian search.
[0017] Preferably, in order to extract global features of different features, while the first encoding vector is input to the visual Transformer encoder, an additional vector is added as an output of a global feature for subsequent cosine similarity judgment;
[0018] Similarly, while the second encoded vector is input into the text Transformer encoder, an additional vector is also added as a global feature output for subsequent cosine similarity judgment.
[0019] Preferably, the calculation formula of the cosine similarity is expressed as:
[0020]
[0021] where f T represents the text global feature, f I represents the image global feature, represents the L2 norm.
[0022] Preferably, before using the two-stream model for recognition, the two-stream model is trained using a cross-modal contrastive learning method.
[0023] Furthermore, the cross-modal contrastive learning method is specifically as follows:
[0024] When training the two-stream model, in one batch, N image-text pairs are given. The text global feature f T of each pair is used as a query, and the image global feature f I is used as a keyword. Then, during training, it is expected that the similarity between the text global feature and the image global feature that match each other is high, while the image global feature that is not similar to f T is used as a negative sample, thereby constructing the first loss function for text searching for images;
[0025] Similarly, using the image feature as a query and the text global feature as a keyword, the two-stream model is trained in turn with the image to search for text, thereby constructing the second loss function for image searching for text;
[0026] The sum of the first loss function and the second loss function is used to construct the final loss function of the two-stream model.
[0027] Still further, the expression of the first loss function is as follows:
[0028]
[0029] The expression of the second loss function is as follows:
[0030]
[0031] The expression of the final loss function is as follows:
[0032]
[0033] where, represents the j-th text global feature, represents the j-th image global feature, represents the temperature coefficient.
[0034] Furthermore, in order to prevent overfitting of the two-stream model, an approximate data generation method is adopted to construct a multi-modal mixed sample for training;
[0035] One or more of the methods including random consecutive multi-word deletion, random word deletion, and random synonym replacement are used to generate approximate texts;
[0036] Two different pictures are interpolated to obtain a new picture, and two different sentences are concatenated to obtain a new sentence. The new picture and the new sentence are used as a multi-modal mixed sample as a training sample to be input into the two-stream model for training.
[0037] Furthermore, the sharpness-aware optimizer SAM is used to replace the traditional Adam optimizer for training the two-stream model. The sharpness-aware optimizer SAM adopts a second-order gradient optimization method.
[0038] Furthermore, the second-order gradient optimization method is specifically as follows:
[0039] Based on the python and PyTorch deep learning frameworks, the sharpness-aware optimizer SAM is adopted, with an initial learning rate of 0.001. A slow start strategy is adopted in the first 10 rounds, and then the learning rate decays to 1 / 10 of the original every 20 rounds; training continues until the two-stream model converges.
[0040] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the pedestrian search method driven by the text description as described are implemented.
[0041] The beneficial effects of the present invention are as follows:
[0042] The present invention proposes a two-stream model of symmetric Transformer. This method uses 2 Transformer encoders to extract image global features and text global features respectively, and its computational complexity is low. At the same time, for both text input and image input, this method uses Transformer to process, solving the problem of difficult training caused by heterogeneity in the existing technology using convolutional neural networks and recurrent neural networks. In addition, without introducing additional local information and external information, this method achieves the best performance compared with the existing methods, which is very friendly for the implementation and deployment of the algorithm.
[0043] Compared with the prior art, the dual-stream model of the present invention has a simple architecture, is easy to implement and deploy. Moreover, the computational complexity of the present invention is relatively low and will not become difficult to calculate as the image gallery increases. Description of the Drawings
[0044] Figure 1 It is a schematic diagram of the text description-driven pedestrian search method of the present invention.
[0045] Figure 2 It is a schematic diagram of approximate text generation of the present invention.
[0046] Figure 3 It is a schematic diagram of multi-modal hybrid sample generation of the present invention. Detailed Embodiments
[0047] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0048] Embodiment 1
[0049] As Figure 1 shown, a text description-driven pedestrian search method includes the following steps:
[0050] Construct a dual-stream model based on symmetric Transformer, where the dual-stream model includes a visual Transformer encoder and a text Transformer encoder;
[0051] For an image, first evenly divide the image, linearly map each image patch and add the position encoding to obtain the first encoding vector input to the visual Transformer encoder; the first encoding vector is N1 D-dimensional vectors; where N1 represents the number of image patches;
[0052] For text, first vectorize and encode each word, then perform linear mapping of the word and add the position encoding to obtain the second encoding vector input to the text Transformer encoder; the second encoding vector is N2 D-dimensional vectors, where N2 represents the number of words;
[0053] Calculate the cosine similarity one by one between the global image feature output by the visual Transformer encoder and the global text feature output by the text Transformer encoder, and find the image with the highest cosine similarity to the corresponding text, so as to achieve text-based pedestrian search.
[0054] This embodiment proposes a two-stream model of a symmetric Transformer. This method uses two Transformer encoders to extract the global features of images and texts respectively. It is a two-stream model framework with low computational complexity. At the same time, this method uses Transformer to process both text input and image input, solving the problem of difficult training caused by the heterogeneity of convolutional neural networks and recurrent neural networks.
[0055] In this embodiment, the structures of the text Transformer encoder and the visual Transformer encoder both have 12 Transformer layers.
[0056] The text Transformer encoder and the visual Transformer encoder are both composed of 12 Transformer structures. Each Transformer structure includes a multi-head attention mechanism (MSA) and a multi-layer linear perceptron (MLP), which can be expressed by mathematical formulas as:
[0057]
[0058] where z l represents the encoded vector of the l th Transformer layer, and layer normalization needs to be performed on each layer. z0 represents the input encoded vector, which is the initial value. represents the encoded vector output by the l th Transformer layer.
[0059] In a specific embodiment, in order to extract the global features of different features, while the first encoded vector is input into the visual Transformer encoder, an additional vector is added as a global feature output for subsequent cosine similarity judgment.
[0060] Similarly, while the second encoded vector is input into the text Transformer encoder, an additional vector is also added as a global feature output for subsequent cosine similarity judgment.
[0061] In this embodiment, the visual Transformer encoder and the text Transformer encoder both finally output the global features of a D-dimensional vector. Only by performing L2 normalization on the text global features and the image global features and then taking the dot product can the cosine similarity be obtained. The calculation formula of the cosine similarity is expressed as:
[0062]
[0063] where, f T represents the text global feature, fI Represents the global image feature, Represents the L2 norm.
[0064] Therefore, under the described two-stream model, it is only necessary to calculate the cosine similarity one by one between the global image feature output by the vision Transformer encoder and the global text feature output by the text Transformer encoder, and find the image with the highest similarity to the corresponding text, then the text-based pedestrian search can be realized.
[0065] Embodiment 2
[0066] Based on the text description-driven pedestrian search method described in Embodiment 1, on this basis, before using the described two-stream model for recognition, in order to improve the recognition accuracy, the two-stream model is trained as follows:
[0067] Construct the final loss function for training the two-stream model, and use the cross-modal contrast learning method to train the two-stream model;
[0068] Adopt the approximate data generation method to generate training samples;
[0069] Use the Sharpness-Aware Minimizer (SAM) optimizer to replace traditional optimizers such as Adam for training.
[0070] In a specific embodiment, the cross-modal contrast learning method is used to train the two-stream model. The cross-modal contrast learning method is specifically as follows:
[0071] When training the two-stream model, in one batch, N image-text pairs are given. The global text feature f of each pair T is used as the query, and the global image feature f I is used as the keyword. Then, during training, it is expected that the similarity between the global text feature and the global image feature that match each other is high, while the global image feature that is not similar to f T is used as a negative sample, thereby constructing the first loss function for text searching for images; Describes this process. When the similarity of the positive sample pairs is high and the similarity of the negative sample pairs is low, this value is relatively small, and vice versa, the loss function is larger.
[0072] Similarly, using the image feature as the query and the global text feature as the keyword, use the image to search for text to train the two-stream model in turn, thereby constructing the second loss function for image searching for text.
[0073] Construct the final loss function of the two-stream model by summing the first loss function and the second loss function.
[0074] In this embodiment, the expression of the first loss function is as follows:
[0075]
[0076] The expression of the second loss function is as follows:
[0077]
[0078] The expression of the final loss function is as follows:
[0079]
[0080] Among them, represents the j-th text global feature, represents the j-th image global feature, represents the temperature coefficient.
[0081] In addition, it is worth mentioning that in order to enable the two-stream model to automatically discover hard samples (some pictures with very high similarity between positive samples) during training, setting a smaller temperature coefficient τ is beneficial for the two-stream model to pay more attention to hard samples, so that the two-stream model can be better trained to achieve better performance.
[0082] In a specific embodiment, in order to prevent the two-stream model from overfitting and thus improve the performance of the model, an approximate data generation method is used to construct multi-modal mixed samples for training;
[0083] Such as Figure 2 shown, one or more of the methods including random consecutive multi-word deletion, random word deletion, and random synonym replacement are used to generate approximate texts; the generated approximate texts and a corresponding picture are input into the two-stream model for training.
[0084] The semantics of the approximate texts generated in this embodiment are generally unchanged, which can play the role of expanding the data set, thereby improving the performance of the two-stream model.
[0085] Two different pictures are interpolated to obtain a new picture, two different sentences are concatenated to obtain a new sentence, and the new picture and the new sentence are used as multi-modal mixed samples as training samples and input into the two-stream model for training.
[0086] In this embodiment, the generation of multi-modal mixed samples is referred to as shown in Figure 3 shown: where X1 and X2 represent two different pictures, and S1 and S2 represent two different sentences. After interpolation and concatenation respectively, the new picture and the new sentence are used as multi-modal mixed samples as training samples and input into the two-stream model, allowing the two-stream model to learn the mixed samples, which can improve the performance of the two-stream model.
[0087] In a specific embodiment, the Sharpness-Aware Minimizer (SAM) is used to replace the traditional Adam optimizer for training the two-stream model. The Sharpness-Aware Minimizer (SAM) adopts a second-order gradient optimization method, which can effectively alleviate the overfitting situation of the two-stream model described in this embodiment under the condition of insufficient training data during the training of the two-stream model, so that the two-stream model can achieve better performance.
[0088] In a specific embodiment, the second-order gradient optimization method is as follows:
[0089] Based on the Python and PyTorch deep learning frameworks, the Sharpness-Aware Minimizer (SAM) is adopted with an initial learning rate of 0.001. A slow start strategy is adopted in the first 10 epochs, and then the learning rate decays to 1 / 10 of the original every 20 epochs; training continues until the two-stream model converges. In this embodiment, after about 70 data epochs, the two-stream model converges to the optimal effect. Optimal means that during the training process, the validation set is verified in each epoch, and the model with the best performance on the validation set is taken as the final result. Convergence means that the loss no longer decreases significantly and tends to be stable.
[0090] In addition, this embodiment fully considers the fine-grained feature of pedestrians in pedestrian search and proposes a cross-modal contrast learning method. Using this contrast learning to train the two-stream model makes the two-stream model perform very well.
[0091] In addition, this embodiment fully considers the overfitting phenomenon caused by the small amount of training data. Therefore, an approximate data generation and second-order gradient optimization method are proposed to prevent the two-stream model from overfitting and further improve the performance of the two-stream model we trained.
[0092] Based on the method described in this embodiment, through experimental verification, the method described in this embodiment achieves the optimal detection performance on three publicly available pedestrian search datasets: CUHK-PEDES, ICFG-PEDES, and RSTPReid.
[0093] Embodiment 3
[0094] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the text description-driven pedestrian search method are as follows:
[0095] Construct a two-stream model based on the symmetric Transformer. The two-stream model includes a visual Transformer encoder and a text Transformer encoder;
[0096] For an image, first evenly divide the image, and after linearly mapping each image block and adding the position encoding, use it as the first encoding vector input to the vision Transformer encoder; the first encoding vector is N1 D-dimensional vectors; where N1 represents the number of image blocks;
[0097] For text, first perform vector quantization encoding on each word, and then use it as the second encoding vector input to the text Transformer encoder after linear mapping of the word and adding the position encoding; the second encoding vector is N2 D-dimensional vectors, where N2 represents the number of words;
[0098] Calculate the cosine similarity one by one between the global image features output by the vision Transformer encoder and the global text features output by the text Transformer encoder, and find the image with the highest cosine similarity to the corresponding text, then the pedestrian search based on text can be realized.
[0099] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0100] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A text description-driven pedestrian search method, characterized in that: The method described above includes the following steps: Construct a two-stream model based on a symmetric Transformer. The two-stream model includes a vision Transformer encoder and a text Transformer encoder; For an image, first evenly divide the image. After linearly mapping each image patch and adding the position encoding, it is used as the first encoding vector input to the vision Transformer encoder. The first encoding vector is N1 D-dimensional vectors, where N1 represents the number of image patches; For text, first vectorize and encode each word, and then after linear mapping of the word and adding the position encoding, it is used as the second encoding vector input to the text Transformer encoder. The second encoding vector is N2 D-dimensional vectors, where N2 represents the number of words; Calculate the cosine similarity one by one between the image global feature output by the vision Transformer encoder and the text global feature output by the text Transformer encoder, and find the image with the highest cosine similarity to the corresponding text, so as to achieve pedestrian search based on text; Before using the two-stream model for recognition, use a cross-modal contrast learning method to train the two-stream model; The cross-modal contrast learning method is specifically as follows: When training the two-stream model, in a batch, given N image-text pairs, the text global feature f of each pair T is used as the query, and the image global feature f I is used as the keyword. Then, during training, it is expected that the similarity between the text global feature and the image global feature that are supposed to match each other is high, while the image global features that are not similar to f T are used as negative samples, thereby constructing the first loss function for text to search for images; Similarly, use the image feature as the query and the text global feature as the keyword, and use the image to search for text to train the two-stream model in turn, thereby constructing the second loss function for image searching text; Construct the final loss function of the two-stream model by adding the first loss function and the second loss function; 2. The text description-driven pedestrian search method according to claim 1, characterized in that: In order to extract the global features of different features, when the first encoding vector is input to the vision Transformer encoder, an additional vector is added as a global feature output for subsequent cosine similarity judgment; Similarly, when the second encoding vector is input to the text Transformer encoder, an additional vector is also added as a global feature output for subsequent cosine similarity judgment; 3. The text description-driven pedestrian search method according to claim 1, characterized in that: The calculation formula of the cosine similarity is expressed as: Among them, f T represents the global text feature, and f I represents the global image feature, which represents the L2 norm.
4. The text description-driven pedestrian search method according to claim 1, characterized in that: The expression of the first loss function is as follows: The expression of the second loss function is as follows: The expression of the final loss function is as follows: Among them, represents the j-th text global feature, represents the j-th image global feature, represents the temperature coefficient.
5. The text description-driven pedestrian search method according to claim 1, characterized in that: In order to prevent the two-stream model from overfitting, an approximate data generation method is used to construct a multi-modal mixed sample for training; Use one or more of the methods including random consecutive multi-word deletion, random word deletion, and random synonym replacement to generate approximate text; input the generated approximate text and a corresponding image into the two-stream model for training; Interpolate two different images to obtain a new image, splice two different sentences to obtain a new sentence, and use the new image and the new sentence as a multi-modal mixed sample as a training sample to input into the two-stream model for training; 6. The text description-driven pedestrian search method according to claim 1, characterized in that: Use the Sharpness-Aware Minimizer (SAM) to replace the traditional Adam optimizer to train the two-stream model. The Sharpness-Aware Minimizer (SAM) uses a second-order gradient optimization method; 7. The text description-driven pedestrian search method according to claim 6, characterized in that: The second-order gradient optimization method is specifically as follows: Based on the Python and PyTorch deep learning frameworks, the Sharpness-Aware Minimizer (SAM) optimizer is adopted, with an initial learning rate of 0.
001. A slow start strategy is used in the first 10 epochs, and then the learning rate decays to 1 / 10 of the original every 20 epochs; training continues until the two-stream model converges.
8. A computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that,When the processor executes the computer program, it implements the steps of the text description-driven pedestrian search method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Near infrared-visible light cross-modal double-current pedestrian re-identification method and system
CN114220124A
Image description generation method and device based on position guidance Transform and computer equipment
CN114782698A