Pedestrian re-identification method and system based on feature fusion and splicing
By adopting feature fusion splicing method in pedestrian recognition, combining deep residual network and multi-head attention mechanism network, the feature noise problem is solved and the accuracy and performance of the model is improved.
Patent Information
- Application Number
- CN202411946181.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The existing unsupervised learning methods have feature-containing noisy information in pedestrian re-identification, which affects the performance of the model.
Using a method based on feature fusion splicing, the deep residual network and multi-head attention mechanism network are constructed, global and local features are extracted, and feature index lists are merged through feature fusion algorithm, local feature coefficients are calculated using interleaving ratios, and weighted splicing is performed, and finally input into the full connection layer for model training.
It improves the accuracy of pedestrian re-identification, improves the performance of the model, and reduces the impact of feature noise.
Smart Images

Figure CN119942637A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a pedestrian re-identification method and system based on feature fusion and splicing. Background Art
[0002] With the rapid development of artificial intelligence, managing and processing massive amounts of video and image data has become a major challenge. Due to the limitations of manual processing of large-scale data, computer-based person re-identification technology is usually used for recognition.
[0003] The pedestrian re-identification system mainly uses a convolutional neural network to search the database for matching person images given a person's head portrait. Traditional supervised learning methods require large-scale manual annotation, which is not in line with reality. To solve this problem, unsupervised learning methods can perform pedestrian re-identification tasks without manually annotating images. In most existing unsupervised training methods, convolutional neural networks or convolutional networks with attention mechanisms are usually used for feature extraction. Although the performance of the model has been improved to a certain extent, the extracted features contain noise information that will affect the final model performance. Summary of the invention
[0004] In view of this, the present invention provides a pedestrian re-identification method and system based on feature fusion and splicing, so as to improve the accuracy of re-identification and enhance the performance of the model.
[0005] In a first aspect, the present invention provides a method for pedestrian re-identification based on feature fusion and splicing, the method comprising: Step 1: Use the collected data set to obtain pedestrian images, and preprocess the pedestrian images to obtain a pedestrian re-identification data set; Step 2: Construct two network models. The first network model is a deep residual network, and the second network model adds a multi-head attention mechanism based on the first network model. Step 3: Pass the person re-identification dataset through the first network model and the second network model respectively, and obtain the corresponding global features and local features respectively. According to the similarity sorting of feature vectors, four corresponding feature index lists of the first k nearest neighbors are obtained respectively; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs respectively to obtain the merged global feature index list and a list of local feature indices ; Step 4: Use the global feature index list and a list of local feature indices The intersection and union ratio of ; Use the feature coefficient to weight the local features obtained by the second network model to obtain the weighted local features , and then concatenate the weighted local features in the channel dimension to obtain the weighted global features ; Step 5: The weighted global features Input the fully connected layer to get the category probability , then use the cross entropy loss function to train the model, and use the trained model for pedestrian re-identification.
[0006] Optionally, step 1 includes: Pedestrian images include different backgrounds and lighting. First, the pedestrian images are converted into the required size, and then the pedestrian images are enhanced by horizontal flipping, padding, and random cropping operations to obtain a pedestrian re-identification dataset.
[0007] Optionally, the step 2 includes: The first network model is a deep residual network, including multiple residual blocks, each residual block contains multiple residual units, each residual unit includes two 3×3 convolutional layers and activation function ReLU; its main structure is input layer, initial convolutional layer, maximum pooling layer, first residual block, second residual block, third residual block, fourth residual block, global average pooling layer, fully connected layer, output layer; The second network model adds a multi-head attention mechanism after each residual block based on the first network model. Specifically: , where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively; given h heads, each head has a corresponding query Q, key K, value V weight matrix; for the input feature The corresponding query Q, key K, value V weight matrix is obtained through convolution operation; the self-attention of each head is calculated, and the specific formula is: ,in , , are the query Q, key K, and value V weight matrices corresponding to the i-th attention head, is the dimension of the key, which is used to scale the dot product to prevent gradient explosion; then the self-attention of all heads are spliced together and merged into the final weight matrix through a linear transformation, and the weight matrix is then combined with the input feature Multiply to get the final weighted feature .
[0008] Optionally, the step 3 includes: The pedestrian re-identification dataset passes through the first network model to obtain global features and local features ; Based on the similarity sorting between feature vectors, the feature index lists of the first k nearest neighbors are obtained, which are represented by , The pedestrian re-identification dataset is passed through the second network model to obtain global features and local features , Similarly, the feature index lists of the first k nearest neighbors are expressed as , ; Through the feature fusion algorithm, for two global feature index lists and , two local feature index lists and , perform feature fusion in pairs; the feature fusion algorithm is to first take out the first k indexes from the index list corresponding to each feature, assign values from the first to the last index, assign the first index k and then decrease 1 in sequence; secondly, merge the two lists according to the index assignment and sort them from large to small; the merging method is that if there are the same indexes, add the assignments and keep one index; if the assignments of different indexes are the same, the different indexes are sorted in the order of the indexes; finally, take out the first k indexes of the merged list to form the final index list; the merged global feature index list is obtained through the feature fusion algorithm and a list of local feature indices .
[0009] Optionally, step 4 includes: By calculating the global feature index list and a list of local feature indices The intersection and union ratio of the local features is obtained ; The intersection-combination ratio formula is: ; in and Represents the global feature index list and a list of local feature indices ; Then the local features obtained by the second network model are weighted by the feature coefficients , get the weighted local features , the specific formula is: , i=1, 2, 3; then concatenate the weighted local features in the channel dimension , and obtain the weighted global features .
[0010] Optionally, step 5 includes: The weighted global features Then add a fully connected layer to get the original score , the specific formula is: , where U is the weight matrix of the fully connected layer and b is the bias term; Then, the original score The activation function Softmax is used to perform activation operation and obtain the category probability ; For the category probability, the cross entropy loss function is used to optimize the model. The specific formula is: ,in is the true label of the sample, and K is the total number of categories.
[0011] In a second aspect, the present invention provides a pedestrian re-identification system based on feature fusion and splicing, the system comprising: A data acquisition and processing module is used to obtain pedestrian images using the collected data set, and pre-process the pedestrian images to obtain a pedestrian re-identification data set; The network construction module is used to construct two network models. The first network model is a deep residual network, and the second network model adds a multi-head attention mechanism on the basis of the first network model. The feature fusion module is used to pass the pedestrian re-identification data set through the first network model and the second network model respectively, and obtain the corresponding global features and local features respectively, and obtain four corresponding feature index lists of the first k nearest neighbors according to the similarity sorting of feature vectors; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs to obtain the merged global feature index list and a list of local feature indices ; Feature splicing module for utilizing global feature index lists and a list of local feature indices The intersection and union ratio of ; Use the feature coefficient to weight the local features obtained by the second network model to obtain the weighted local features , and then concatenate the weighted local features in the channel dimension to obtain the weighted global features ; Data training module, used to transform weighted global features Input the fully connected layer to get the category probability , then use the cross entropy loss function to train the model, and use the trained model for pedestrian re-identification.
[0012] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the pedestrian re-identification method based on feature fusion and splicing in the first aspect or any possible implementation of the first aspect.
[0013] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute the pedestrian re-identification method based on feature fusion and splicing in the first aspect or any possible implementation of the first aspect.
[0014] In the technical solution provided by the present invention, the method includes using the collected data set to obtain pedestrian images, and preprocessing the pedestrian images to obtain a pedestrian re-identification data set; constructing two network models, the first network model is a deep residual network, and the second network model adds a multi-head attention mechanism on the basis of the first network model; the pedestrian re-identification data set is passed through the first network model and the second network model respectively, and the corresponding global features and local features are obtained respectively, and four corresponding feature index lists of the top k nearest neighbors are obtained respectively according to the similarity sorting of feature vectors; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs respectively to obtain a merged global feature index list and a local feature index list; using the full The intersection and union ratio of the local feature index list and the local feature index list is calculated to obtain the local feature coefficient; the local features obtained by the second network model are weighted by the feature coefficient to obtain the weighted local features, and then the weighted local features are spliced in the channel dimension to obtain the weighted global features; the weighted global features are input into the fully connected layer to obtain the category probability, and then the model is trained using the cross entropy loss function, and pedestrian re-identification is performed using the trained model. This method introduces a feature fusion splicing algorithm, selects appropriate global feature index lists and local feature index lists, and uses the coefficients obtained by calculating the intersection and union ratio to weight the local features to obtain the weighted global features. Finally, the global features are used for model training, which improves the accuracy of re-identification and improves the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 A flowchart of a pedestrian re-identification method provided by an embodiment of the present invention; Figure 2 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.
[0020] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0021] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0022] Figure 1 A flowchart of a pedestrian re-identification method provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes: Step 1: Use the collected data set to obtain pedestrian images, and preprocess the pedestrian images to obtain a pedestrian re-identification data set.
[0023] In the embodiment of the present invention, step 1 includes: Pedestrian images include different backgrounds and lighting. First, the pedestrian images are converted into the required size, and then the pedestrian images are enhanced by horizontal flipping, padding, and random cropping operations to obtain a pedestrian re-identification dataset.
[0024] Step 2: Construct two network models. The first network model is a deep residual network, and the second network model adds a multi-head attention mechanism based on the first network model.
[0025] In the embodiment of the present invention, step 2 includes: The first network model is a deep residual network, including multiple residual blocks, each residual block contains multiple residual units, each residual unit includes two 3×3 convolutional layers and activation function ReLU; its main structure is input layer, initial convolutional layer, maximum pooling layer, first residual block, second residual block, third residual block, fourth residual block, global average pooling layer, fully connected layer, output layer; The second network model adds a multi-head attention mechanism after each residual block based on the first network model. Specifically: , where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively; given h heads, each head has a corresponding query Q, key K, value V weight matrix; for the input feature The corresponding query Q, key K, value V weight matrix is obtained through convolution operation; the self-attention of each head is calculated, and the specific formula is: ,in , , are the query Q, key K, and value V weight matrices corresponding to the i-th attention head, is the dimension of the key, which is used to scale the dot product to prevent gradient explosion; then the self-attention of all heads are spliced together and merged into the final weight matrix through a linear transformation, and the weight matrix is then combined with the input feature Multiply to get the final weighted feature .
[0026] Step 3: Pass the person re-identification dataset through the first network model and the second network model respectively, and obtain the corresponding global features and local features respectively. According to the similarity sorting of feature vectors, four corresponding feature index lists of the first k nearest neighbors are obtained respectively; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs respectively to obtain the merged global feature index list and a list of local feature indices .
[0027] In the embodiment of the present invention, step 3 includes: The pedestrian re-identification dataset passes through the first network model to obtain global features and local features ; Based on the similarity sorting between feature vectors, the feature index lists of the first k nearest neighbors are obtained, which are represented by , The pedestrian re-identification dataset is passed through the second network model to obtain global features and local features , Similarly, the feature index lists of the first k nearest neighbors are expressed as , ; Through the feature fusion algorithm, for two global feature index lists and , two local feature index lists and , perform feature fusion in pairs; the feature fusion algorithm is to first take out the first k indexes from the index list corresponding to each feature, assign values from the first to the last index, assign the first index k and then decrease 1 in sequence; secondly, merge the two lists according to the index assignment and sort them from large to small; the merging method is that if there are the same indexes, add the assignments and keep one index; if the assignments of different indexes are the same, the different indexes are sorted in the order of the indexes; finally, take out the first k indexes of the merged list to form the final index list; the merged global feature index list is obtained through the feature fusion algorithm and a list of local feature indices .
[0028] Step 4: Use the global feature index list and a list of local feature indices The intersection and union ratio of ; Use the feature coefficient to weight the local features obtained by the second network model to obtain the weighted local features , and then concatenate the weighted local features in the channel dimension to obtain the weighted global features .
[0029] In the embodiment of the present invention, step 4 includes: By calculating the global feature index list and a list of local feature indices The intersection and union ratio of the local features is obtained ; The intersection-combination ratio formula is: ; in and Represents the global feature index list and a list of local feature indices ; Then the local features obtained by the second network model are weighted by the feature coefficients , get the weighted local features , the specific formula is: , i=1, 2, 3; then concatenate the weighted local features in the channel dimension , and obtain the weighted global features .
[0030] Step 5: The weighted global features Input the fully connected layer to get the category probability , then use the cross entropy loss function to train the model, and use the trained model for pedestrian re-identification.
[0031] In the embodiment of the present invention, step 5 includes: The weighted global features Then add a fully connected layer to get the original score , the specific formula is: , where U is the weight matrix of the fully connected layer and b is the bias term; Then, the original score The activation function Softmax is used to perform activation operation and obtain the category probability ; For the category probability, the cross entropy loss function is used to optimize the model. The specific formula is: ,in is the true label of the sample, and K is the total number of categories.
[0032] The present invention provides a pedestrian re-identification system based on feature fusion and splicing, the system comprising: Data acquisition processing module 1, network construction module 2, feature fusion module 3, feature splicing module 4 and data training module 5; the data acquisition processing module 1 is connected to the network construction module 2, the network construction module 2 is connected to the feature fusion module 3, the feature fusion module 3 is connected to the feature splicing module 4, and the feature splicing module 4 is connected to the data training module 5.
[0033] The data acquisition processing module 1 is used to obtain pedestrian images using the collected data set, and pre-process the pedestrian images to obtain a pedestrian re-identification data set; the network construction module 2 is used to construct two network models, the first network model is a deep residual network, and the second network model adds a multi-head attention mechanism on the basis of the first network model; the feature fusion module 3 is used to pass the pedestrian re-identification data set through the first network model and the second network model respectively, and obtain the corresponding global features and local features respectively, and obtain four corresponding feature index lists of the top k nearest neighbors according to the similarity sorting of the feature vectors; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs to obtain a merged global feature index list and a list of local feature indices ; Feature splicing module 4 is used to use the global feature index list and a list of local feature indices The intersection and union ratio of ; Use the feature coefficient to weight the local features obtained by the second network model to obtain the weighted local features , and then concatenate the weighted local features in the channel dimension to obtain the weighted global features ; Data training module 5 is used to convert the weighted global features Input the fully connected layer to get the category probability , then use the cross entropy loss function to train the model, and use the trained model for pedestrian re-identification.
[0034] The cross entropy loss function is used to measure the difference between the predicted probability distribution and the true label, continuously improving the model performance.
[0035] Traditional pedestrian re-identification systems usually cannot distinguish the quality of features obtained by adding attention and features without adding attention, or only use features obtained by adding attention for training, ignoring the different importance of different areas of the image. The present invention introduces a feature fusion splicing algorithm, selects appropriate global feature index lists and local feature index lists, and uses the coefficients obtained by calculating their intersection and union ratio to weight local features to obtain weighted global features. Finally, the global features are used for model training to continuously improve accuracy.
[0036] In the technical solution provided by the present invention, the method includes using the collected data set to obtain pedestrian images, and preprocessing the pedestrian images to obtain a pedestrian re-identification data set; constructing two network models, the first network model is a deep residual network, and the second network model adds a multi-head attention mechanism on the basis of the first network model; the pedestrian re-identification data set is passed through the first network model and the second network model respectively, and the corresponding global features and local features are obtained respectively, and four corresponding feature index lists of the top k nearest neighbors are obtained respectively according to the similarity sorting of feature vectors; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs respectively to obtain a merged global feature index list and a local feature index list; using the full The intersection and union ratio of the local feature index list and the local feature index list is calculated to obtain the local feature coefficient; the local features obtained by the second network model are weighted by the feature coefficient to obtain the weighted local features, and then the weighted local features are spliced in the channel dimension to obtain the weighted global features; the weighted global features are input into the fully connected layer to obtain the category probability, and then the model is trained using the cross entropy loss function, and pedestrian re-identification is performed using the trained model. This method introduces a feature fusion splicing algorithm, selects appropriate global feature index lists and local feature index lists, and uses the coefficients obtained by calculating the intersection and union ratio to weight the local features to obtain the weighted global features. Finally, the global features are used for model training, which improves the accuracy of re-identification and improves the performance of the model.
[0037] Each step of the embodiment of the present invention may be performed by an electronic device, which includes but is not limited to a mobile phone, a tablet computer, a portable PC, a desktop computer, etc.
[0038] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the electronic device where the computer-readable storage medium is located is controlled to execute the above-mentioned embodiment of the pedestrian re-identification method based on feature fusion and splicing.
[0039] Figure 2 A schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 2 As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, the pedestrian re-identification method based on feature fusion and splicing in the embodiment is implemented. To avoid repetition, they are not described one by one here.
[0040] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art will appreciate that Figure 2It is only an example of the electronic device 21 and does not constitute a limitation of the electronic device 21. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0041] The processor 211 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0042] The memory 212 may be an internal storage unit of the electronic device 21, such as a hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card (FlashCard), etc. equipped on the electronic device 21. Further, the memory 212 may also include both an internal storage unit of the electronic device 21 and an external storage device. The memory 212 is used to store computer programs and other programs and data required by network devices. The memory 212 may also be used to temporarily store data that has been output or is to be output.
[0043] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A pedestrian re-identification method based on feature fusion and splicing, characterized in that: The method comprises: Step 1: Use the collected data set to obtain pedestrian images, and preprocess the pedestrian images to obtain a pedestrian re-identification data set; Step 2: Construct two network models. The first network model is a deep residual network, and the second network model adds a multi-head attention mechanism based on the first network model. Step 3: Pass the person re-identification dataset through the first network model and the second network model respectively, and obtain the corresponding global features and local features respectively. According to the similarity sorting of the feature vectors, four corresponding feature index lists of the top k nearest neighbors are obtained respectively; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs respectively to obtain the merged global feature index list and a list of local feature indices ; Step 4: Use the global feature index list and a list of local feature indices The intersection and union ratio of ; Use the feature coefficient to weight the local features obtained by the second network model to obtain the weighted local features , and then concatenate the weighted local features in the channel dimension to obtain the weighted global features ; Step 5: The weighted global features Input the fully connected layer to get the category probability , then use the cross entropy loss function to train the model, and use the trained model to perform pedestrian re-identification.
2. The method according to claim 1, characterized in that The step 1 comprises: Pedestrian images include different backgrounds and lighting. First, the pedestrian images are converted into the required size, and then the pedestrian images are enhanced by horizontal flipping, padding, and random cropping operations to obtain a pedestrian re-identification dataset.
3. The method according to claim 1, characterized in that The step 2 comprises: The first network model is a deep residual network, including multiple residual blocks, each residual block contains multiple residual units, each residual unit includes two 3×3 convolutional layers and activation function ReLU; its main structure is input layer, initial convolutional layer, maximum pooling layer, first residual block, second residual block, third residual block, fourth residual block, global average pooling layer, fully connected layer, output layer; The second network model adds a multi-head attention mechanism after each residual block based on the first network model. Specifically: , where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively; given h heads, each head has a corresponding query Q, key K, value V weight matrix; for the input feature The corresponding query Q, key K, value V weight matrix is obtained through convolution operation; the self-attention of each head is calculated, and the specific formula is: ,in , , are the query Q, key K, and value V weight matrices corresponding to the i-th attention head, is the dimension of the key, which is used to scale the dot product to prevent gradient explosion; then the self-attention of all heads are spliced together and merged into the final weight matrix through a linear transformation, and the weight matrix is then combined with the input feature Multiply to get the final weighted feature .
4. The method according to claim 1, characterized in that: The step 3 comprises: The pedestrian re-identification dataset passes through the first network model to obtain global features and local features ; Based on the similarity sorting between feature vectors, the feature index lists of the first k nearest neighbors are obtained, which are represented by , The pedestrian re-identification dataset is passed through the second network model to obtain global features and local features , Similarly, the feature index lists of the first k nearest neighbors are expressed as , ; Through the feature fusion algorithm, for two global feature index lists and , two local feature index lists and , perform feature fusion in pairs; the feature fusion algorithm is to first take out the first k indexes from the index list corresponding to each feature, assign values from the first to the last index, assign the first index k and then decrease 1 in sequence; secondly, merge the two lists according to the index assignment and sort them from large to small; the merging method is that if there are the same indexes, add the assignments and keep one index; if the assignments of different indexes are the same, the different indexes are sorted in the order of the indexes; finally, take out the first k indexes of the merged list to form the final index list; the merged global feature index list is obtained through the feature fusion algorithm and a list of local feature indices .
5. The method according to claim 1, characterized in that The step 4 comprises: By calculating the global feature index list and a list of local feature indices The intersection and union ratio of the local features is obtained ; The intersection-combination ratio formula is: ; in and Represents the global feature index list and a list of local feature indices ; Then the local features obtained by the second network model are weighted by the feature coefficients , get the weighted local features , the specific formula is: , i=1, 2, 3; then concatenate the weighted local features in the channel dimension , and obtain the weighted global features .
6. The method according to claim 1, characterized in that The step 5 comprises: The weighted global features Then add a fully connected layer to get the original score , the specific formula is: , where U is the weight matrix of the fully connected layer and b is the bias term; Then, the original score The activation function Softmax is used to perform activation operation and obtain the category probability ; For the category probability, the cross entropy loss function is used to optimize the model. The specific formula is: ,in is the true label of the sample, and K is the total number of categories.
7. A pedestrian re-identification system based on feature fusion and splicing, characterized in that: The system comprises: A data acquisition and processing module is used to obtain pedestrian images using the collected data set, and pre-process the pedestrian images to obtain a pedestrian re-identification data set; The network construction module is used to construct two network models. The first network model is a deep residual network, and the second network model adds a multi-head attention mechanism on the basis of the first network model. The feature fusion module is used to pass the pedestrian re-identification data set through the first network model and the second network model respectively, and obtain the corresponding global features and local features respectively, and obtain four corresponding feature index lists of the first k nearest neighbors according to the similarity sorting of feature vectors; through the feature fusion algorithm, the two global feature index lists and the two local feature index lists are fused in pairs to obtain the merged global feature index list and a list of local feature indices ; Feature splicing module for utilizing global feature index lists and a list of local feature indices The intersection and union ratio of ; Use the feature coefficient to weight the local features obtained by the second network model to obtain the weighted local features , and then concatenate the weighted local features in the channel dimension to obtain the weighted global features ; Data training module, used to transform weighted global features Input the fully connected layer to get the category probability , then use the cross entropy loss function to train the model, and use the trained model to perform pedestrian re-identification.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the pedestrian re-identification method based on feature fusion and splicing according to any one of claims 1 to 6.
9. An electronic device, characterized in that: include: one or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to perform the pedestrian re-identification method based on feature fusion and splicing as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Foreground guiding and texture focusing pedestrian re-identification model establishing method and application thereof
CN112163498A
Video pedestrian re-identification method based on prior knowledge
CN115050050A
Global feature and stepped local feature fused pedestrian re-identification method and device
CN115171165A
Pedestrian re-identification method and system based on second-order attention mechanism, and related equipment
CN117437657A
Object image re-identification method based on multi-feature information capture and correlation analysis
WO2023273290A1