Target re-identification method and device based on foreground segmentation
Through a foreground segmentation-based method, dual-stream network and feature fusion technology are used to solve the problem of background interference in traditional re-identification technology, and achieve higher target re-identification accuracy.
Patent Information
- Application Number
- CN202510710370.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional re-identification technology is difficult to accurately identify and match targets in complex backgrounds and dynamic environments, especially affected by background interference and occlusion by other vehicles/peeders in dynamic environments.
The foreground segmentation method is adopted to extract the features of the image to be identified and the foreground image respectively through a dual-stream network, build feature channel weights for feature fusion, and use global and local feature extraction methods to calculate feature similarity in combination with the trained prediction model.
Effectively eliminate background interference, improve the accuracy of target re-identification, and can more accurately match the target to be identified and the target in the image to be compared.
Smart Images

Figure CN120580435A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and in particular to a method and device for target re-identification based on foreground segmentation. Background Art
[0002] With the widespread adoption of smart campuses and intelligent access control systems, the management and control of visitors, such as vehicles and pedestrians, has become increasingly important. Traditional recognition technologies rely primarily on image recognition from a single camera, making them incapable of tracking and identifying targets across multiple devices. To address this issue, re-identification technology has emerged.
[0003] Re-ID technology is a computer vision-based target tracking and recognition technology. Its primary task is to accurately identify and match the same target at different times and locations. This technology has important applications in smart campuses and intelligent access control systems, enabling precise control of visitors and free personnel, preventing unauthorized access to sensitive areas, and ensuring internal campus security.
[0004] However, re-identification technology faces many technical challenges. Complex backgrounds and other vehicles / pedestrians in dynamic environments may block the target, thereby affecting the target detection and feature extraction process. Summary of the Invention
[0005] To address the above issues, this application provides a method for target re-identification based on foreground segmentation, which includes the following:
[0006] In a first aspect, the present application provides a method for target re-identification based on foreground segmentation, the method comprising:
[0007] Acquire an image to be identified, and perform foreground segmentation processing on the image to be identified to obtain a foreground image;
[0008] Using a dual-stream network to extract features from the image to be identified and the foreground image, respectively, to obtain a feature map of the image to be identified and a feature map of the foreground image;
[0009] Constructing feature channel weights based on the feature map of the foreground image, and performing feature fusion on the feature channel weights and the feature map of the image to be identified to obtain a fused feature map;
[0010] Performing feature extraction on the fused feature map using a global feature extraction method and a local feature extraction method respectively to obtain global features and local features;
[0011] splicing the global features and the local features together as the target to be identified in the image to be identified;
[0012] The trained prediction model is used to calculate the similarity of features between the target to be identified and all images to be compared.
[0013] Optionally, the processing of the image to be identified to obtain a foreground image includes:
[0014] Processing the image to be identified using foreground segmentation technology to obtain a binary mask of the target object;
[0015] The binary mask is fused with the image to be identified to obtain the foreground image.
[0016] Optionally, constructing feature channel weights based on the feature map of the foreground image includes:
[0017] Performing a global average pooling operation on the feature map of the foreground image;
[0018] The compressed feature vector is input into the 1x1 convolution layer to reduce the number of feature channels;
[0019] The ReLu activation function is used to enhance the nonlinear characteristics of the channel, and then the fully connected layer is input to restore the number of feature channels;
[0020] Use the sigmoid activation function to excite the channel and obtain the enhanced feature channel weights.
[0021] Optionally, when using the trained prediction model to calculate the similarity of features between the target to be identified and all images to be compared, the method includes:
[0022] Calculate the Euclidean distance between the target to be identified and each image to be compared;
[0023] Sort all the images to be compared according to the calculated Euclidean distance to obtain a similarity sorted list;
[0024] The final search results are returned in a list sorted by similarity, with images with high similarity at the front.
[0025] Optionally, the model training process includes:
[0026] Perform foreground segmentation on the input image to obtain the target area image, and extract the original image features and foreground area features through a two-stream network respectively;
[0027] Generating channel weights based on the foreground area features, and performing weighted fusion with the original image features to obtain fused features;
[0028] The fusion features are extracted through global features and local features to extract global feature vectors and local feature vectors;
[0029] The cross entropy loss function is used to optimize the global features and local feature vectors, and the model parameters are adjusted through iterative training until convergence to complete the model training.
[0030] In a second aspect, the present application provides a device for object re-identification based on foreground segmentation, the device comprising:
[0031] a foreground image extraction unit, configured to obtain an image to be identified and perform foreground segmentation processing on the image to be identified to obtain a foreground image;
[0032] a first feature extraction unit, configured to perform feature extraction on the image to be identified and the foreground image respectively using a dual-stream network to obtain a feature map of the image to be identified and a feature map of the foreground image;
[0033] a fusion unit, configured to construct a feature channel weight based on the feature map of the foreground image, and perform feature fusion on the feature channel weight and the feature map of the image to be identified to obtain a fused feature map;
[0034] A second feature extraction unit is used to extract features from the fused feature map using a global feature extraction method and a local feature extraction method to obtain global features and local features;
[0035] A feature splicing unit, configured to splice the global features and the local features as an object to be identified in the image to be identified;
[0036] The comparison unit is used to calculate the similarity of features between the target to be identified and all images to be compared using the trained prediction model.
[0037] Optionally, the foreground image extraction unit is specifically configured to:
[0038] Processing the image to be identified using foreground segmentation technology to obtain a binary mask of the target object;
[0039] The binary mask is fused with the image to be identified to obtain the foreground image.
[0040] Optionally, the fusion unit constructing feature channel weights based on the feature map of the foreground image includes:
[0041] Performing a global average pooling operation on the feature map of the foreground image;
[0042] The compressed feature vector is input into the 1x1 convolution layer to reduce the number of feature channels;
[0043] The ReLu activation function is used to enhance the nonlinear characteristics of the channel, and then the fully connected layer is input to restore the number of feature channels;
[0044] Use the sigmoid activation function to excite the channel and obtain the enhanced feature channel weights.
[0045] Optionally, when the comparison unit uses the trained prediction model to calculate the similarity of features between the target to be identified and all images to be compared, the method includes:
[0046] Calculate the Euclidean distance between the target to be identified and each image to be compared;
[0047] Sort all the images to be compared according to the calculated Euclidean distance to obtain a similarity sorted list;
[0048] The final search results are returned in a list sorted by similarity, with images with high similarity at the front.
[0049] Optionally, the device further includes a model training unit for model training, and the model training process includes:
[0050] Perform foreground segmentation on the input image to obtain the target area image, and extract the original image features and foreground area features through a two-stream network respectively;
[0051] Generating channel weights based on the foreground area features, and performing weighted fusion with the original image features to obtain fused features;
[0052] The fusion features are extracted through global features and local features to extract global feature vectors and local feature vectors;
[0053] The cross entropy loss function is used to optimize the global features and local feature vectors, and the model parameters are adjusted through iterative training until convergence to complete the model training.
[0054] In a third aspect, the present application provides a device comprising a memory and a processor, wherein the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device performs the foreground segmentation-based target re-identification method introduced in any implementation of the first aspect.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a code. When the code is executed, the device executing the code implements the target re-identification method based on foreground segmentation introduced in any implementation method of the first aspect.
[0056] The present application provides a target re-identification method based on foreground segmentation. When executing the method, first obtain the image to be identified, perform foreground segmentation processing on the image to be identified to obtain a foreground image, then use a two-stream network to extract features from the image to be identified and the foreground image respectively, obtain a feature map of the image to be identified and a feature map of the foreground image, construct feature channel weights based on the feature map of the foreground image, perform feature fusion on the feature channel weights and the feature map of the image to be identified, obtain a fused feature map, respectively use a global feature extraction method and a local feature extraction method to extract features from the fused feature map, obtain global features and local features, splice the global features and the local features together as the target to be identified in the image to be identified, and use the trained prediction model to calculate the similarity of features between the target to be identified and all images to be compared.
[0057] This method uses foreground segmentation to precisely focus on the target area in the image to be identified, effectively eliminating background interference. A dual-stream network is used to extract features from both the original image and the foreground image. Feature fusion is then performed using feature channel weights constructed based on the foreground image feature map. This fused feature map contains both global information and highlights key local features of the target. Global and local feature extraction methods are then used to extract and combine global and local features, providing a comprehensive and detailed description of the target to be identified. This significantly improves the accuracy of target re-identification and enables more precise matching of the target to be identified with the target in the image to be compared. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the following briefly introduces the drawings required for use in the embodiment or the prior art description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0059] Figure 1 A flowchart of a method for target re-identification based on foreground segmentation provided in an embodiment of the present application;
[0060] Figure 2 A schematic diagram of obtaining a foreground image provided in an embodiment of the present application;
[0061] Figure 3 A schematic diagram of the structure of a dual-flow network provided in an embodiment of the present application;
[0062] Figure 4 A schematic diagram of the structure of a feature fusion module provided in an embodiment of the present application;
[0063] Figure 5 A schematic diagram of generating a feature vector provided in an embodiment of the present application;
[0064] Figure 6 A schematic diagram of comparing the similarity of image features using a trained prediction model provided in an embodiment of the present application;
[0065] Figure 7 A schematic diagram of a model training process provided in an embodiment of the present application;
[0066] Figure 8 A schematic diagram of the structure of an object re-identification device based on foreground segmentation provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solutions and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0068] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0069] Figure 1 This is a flow chart of a target re-identification method based on foreground segmentation provided in an embodiment of the present application. Figure 1 As shown, the target re-identification method based on foreground segmentation provided in the embodiment of the present application may include:
[0070] S101 , obtaining an image to be identified, and performing foreground segmentation processing on the image to be identified to obtain a foreground image.
[0071] First, an original image to be recognized is obtained, wherein the original image contains a background and a foreground, i.e., a target object, such as a pedestrian or a vehicle.
[0072] The original image is processed using image processing techniques such as YOLACT and Mask R-CNN for foreground segmentation algorithms to separate the foreground target object from the background. The result of foreground segmentation is a binary mask image in which the foreground region (i.e., the target object) is marked as white or a specific color, and the background region is marked as black or another color. Based on the foreground segmentation results, the portion of the original image corresponding to the foreground region is extracted to generate a foreground image containing only the target object. This step helps reduce background interference, allowing subsequent feature extraction and recognition to focus more on the target object itself.
[0073] In one implementation of the embodiment of the present application, the processing of the image to be identified to obtain a foreground image includes: processing the image to be identified using a foreground segmentation technique to obtain a binary mask of the target object; and fusing the binary mask with the image to be identified to obtain the foreground image.
[0074] The specific implementation process is as follows Figure 2 As shown, Figure 2 A schematic diagram of obtaining a foreground image provided in an embodiment of the present application, first, the acquired image to be identified is resized, all images are scaled to a uniform size such as 256×128 pixels, and the size of the input image is standardized to ensure consistency and efficiency of subsequent processing. The resized image is processed using a foreground segmentation network, and the image is processed by a YOLACT network in an embodiment of the present application. The YOLACT network can efficiently identify and separate target objects such as vehicles, pedestrians, etc. from the background in the image. After processing, a binary mask image with the same size as the input image is generated. In the binary mask, the area where the target object is located is marked as white, and the background area is marked as black. The generated binary mask is then projected pixel by pixel onto the original resized image. The final foreground area image is as shown below. Figure 2 As shown in the figure, background interference is removed and only the visual information of the target object is retained, providing clear and accurate input data for subsequent feature extraction and image re-identification tasks.
[0075] S102 , using a dual-stream network to perform feature extraction on the image to be identified and the foreground image respectively, to obtain a feature map of the image to be identified and a feature map of the foreground image.
[0076] The two-stream network is a deep learning network architecture. Its core idea is to process input data in parallel through two independent branches (i.e., "streams") to capture different types of information or features, and finally fuse the results of the two branches to improve the performance of the model. Figure 3This is a schematic diagram of the structure of a dual-stream network provided in an embodiment of the present application. In this embodiment of the present application, the network that uses the image to be identified as input is called the original branch, and the network that uses the foreground image as input is called the foreground branch. For the original branch, the SENet-50 network that provides channel attention is used as the backbone network to extract the feature map F1 of the image to be identified. The information between the channels of this network is correlated with each other, and then clues that are helpful for identification are discovered from the target object area. For the foreground branch, ResNet-50 is used as the backbone network of this branch to extract the discriminative foreground area feature map F2.
[0077] S103 , constructing feature channel weights based on the feature map of the foreground image, and performing feature fusion on the feature channel weights and the feature map of the image to be identified to obtain a fused feature map.
[0078] In the embodiments of this application, feature channel weights are constructed using the SE (Squeeze-and-Excitation) Block. The SE Block is a module used to enhance network representation capabilities. It adjusts the channel weights of the feature map by explicitly modeling the relationship between channels. After obtaining the feature channel weights w, the calculated feature channel weights are applied to the feature map F1 of the image to be recognized to obtain a new feature map F. This process is typically achieved by multiplying the weights by the corresponding channels in the feature map, that is, using the calculation F = F1 × w.
[0079] This weighted fusion combines information from the foreground image to optimize the feature representation of the image to be identified, allowing the fused feature map to focus more on features that are helpful for the recognition task. After feature fusion, a new feature map is obtained. This step allows for more effective utilization of foreground image information to enhance the feature representation of the image to be identified, thereby improving the accuracy of subsequent recognition tasks.
[0080] In one implementation of an embodiment of the present application, constructing feature channel weights based on the feature map of the foreground image includes: performing a global average pooling operation on the feature map of the foreground image; inputting the compressed feature vector into a 1x1 convolutional layer to reduce the number of feature channels; using a ReLu activation function to enhance the nonlinear characteristics of the channel, and then inputting the fully connected layer to restore the number of feature channels; using a sigmoid activation function to excite the channel to obtain the enhanced feature channel weights.
[0081] Figure 4A structural diagram of a feature fusion module provided in an embodiment of the present application, first performs a global average pooling (GAP) operation on the feature map F2 of the foreground image. This step compresses the spatial information of each channel into a scalar value to generate a compressed feature vector. Global average pooling helps to reduce computational complexity and retain the global information of the channel. The compressed feature vector is then input into a 1x1 convolutional layer, and the number of feature channels is transformed from c to c / r, where c is 2048 and r is 4. The purpose of this step is to reduce the number of feature channels so that subsequent processing is more efficient. The 1x1 convolutional layer achieves a reduction in the number of channels by adjusting the number of convolution kernels.
[0082] The ReLU (Rectified Linear Unit) activation function is then used to perform a nonlinear transformation on the features processed by the 1x1 convolutional layer. The ReLU activation function can enhance the nonlinear expression ability of the network, enabling the network to learn more complex feature relationships. The features processed by the ReLU activation function are input into a fully connected layer to restore the number of feature channels. The purpose of this step is to adjust the number of feature channels back to the number of channels of the original feature map for subsequent channel weight application. Finally, the Sigmoid activation function is used to excite the features output by the fully connected layer to obtain a set of enhanced feature channel weights w. The Sigmoid function limits the weight of each channel to between 0 and 1, indicating the relative importance of each channel.
[0083] S104 , performing feature extraction on the fused feature map using a global feature extraction method and a local feature extraction method respectively, to obtain global features and local features.
[0084] In order to capture information at different levels of the image more comprehensively and meticulously, this application uses global feature extraction methods and local feature extraction methods to extract features from the fused feature map, where global features provide overall semantic and structural understanding, and local features focus on details and key areas. The combination of the two makes the model's understanding of the image more in-depth and accurate.
[0085] Figure 5 A schematic diagram of generating a feature vector provided in the embodiment of the present application, combined with Figure 5 The process of obtaining global features and local features in the embodiments of the present application is described in detail.
[0086] During the feature segmentation stage, this solution divides the fused features into blocks equally in the horizontal direction in the local branch. The more blocks there are, the finer the feature granularity is, and it can focus on local but detailed information of the image; the fewer blocks there are, the coarser the feature granularity is, and it can provide global but rough information of the image. Since when the number of feature blocks is too small, the granularity of the local features after segmentation is coarse, and its representation ability is close to that of the global features; when there are too many segmented parts, the granularity of the local features after segmentation is fine, which is not enough to better describe the target features. Therefore, in this application, the number of blocks in the local branch is selected as 6 to achieve a better balance between local feature granularity and feature representation ability. It can ensure that the local features extracted by each block are detailed enough to capture the subtle features and changes of the target object in the local area of the image, so as to better describe the local structure, texture and other detailed information of the target and avoid the loss of local information due to too few blocks; it can also prevent too many blocks from making the information contained in each block too fragmented, making it difficult to form an effective description of the overall characteristics of the target, ensuring that the local features can still maintain a certain degree of relevance and integrity, thereby improving the accuracy and effectiveness of feature segmentation and providing more discriminative and expressive feature information for subsequent image processing tasks. In order to conveniently represent the features after segmentation, the features after the segmentation stage are defined as F i,j , where i represents global / local features. When i = g, it represents global features, i.e., features obtained after global branch processing. These features cover the entire image. When i = l, it represents local features, i.e., features of each block obtained after local branch processing. These features focus on a local area of the image. j is used as the index of the block feature. In the local branch, j ranges from 1 to 6, corresponding to six different block features.
[0087] For the global branch, the fused feature F is not segmented to maintain the integrity of the feature map, thereby extracting the global features of the image, which can reflect the overall semantic and structural information of the image.
[0088] S105: Concatenate the global features and the local features as the target to be identified in the image to be identified.
[0089] In tasks like image recognition, global and local features each possess unique information value. Combining these two allows us to integrate feature information at different levels and scales to more comprehensively and accurately describe the target in the image, providing a richer and more effective feature representation for subsequent recognition, classification, and detection tasks.
[0090] Concatenating the global and local features refers to combining the global and local feature vectors in the feature space according to certain rules. Since both global and local features are represented as vectors, they are simply concatenated dimensionally to form a new, higher-dimensional feature vector. Assuming the global feature vector has dimension m and the local feature vector has dimension n, the concatenated feature vector has dimension m+n.
[0091] By combining global and local features, we can leverage the strengths of both to improve the descriptive capabilities of the target being identified. In subsequent recognition tasks, these fused features can provide richer information, helping the model more accurately identify the target's category, location, posture, and other attributes, thereby improving the accuracy of recognition tasks.
[0092] S106: Calculate the similarity of features between the target to be identified and all images to be compared using the trained prediction model.
[0093] In step S105 above, the global and local features of the target in the image to be identified are combined to obtain a feature representation of the target. Similarly, feature representations must be extracted for all images to be compared. These feature representations form the basis for subsequent similarity calculations.
[0094] The features of the target to be identified and the features of all the images to be compared are sequentially input into the trained prediction model. The model will process these features based on its internal structure and parameter settings.
[0095] The prediction model contains specific similarity calculation modules, such as cosine similarity calculation, Euclidean distance calculation, etc. The model will calculate the similarity value between the features of the target to be identified and the features of each image to be compared according to the input features and the preset similarity calculation method. In the embodiment of the present application, the similarity of the features is calculated by Euclidean distance, and then the similarity value of the features between the target to be identified and each image to be compared is obtained. These values can be sorted in order from high to low, and according to the specific task requirements, several images to be compared with the highest similarity are selected as the results most relevant to the target to be identified. In the image retrieval task in the embodiment of the present application, the images with the highest similarity ranking are returned. For example Figure 6 As shown, Figure 6 A schematic diagram of an embodiment of the present application provides a method for comparing the similarity of image features using a trained prediction model. The image to be queried and the set of images to be compared are input into a pre-trained prediction model. After the prediction model performs a similarity comparison, the image to be searched is obtained. The result includes multiple images for the user to select.
[0096] The following is an introduction to the training process of the pre-trained prediction model in this application. Figure 7 A flow chart of a model training process provided in an embodiment of the present application, according to Figure 7 , the model training process includes:
[0097] S701. Perform foreground segmentation on the input image to obtain a target area image, and extract original image features and foreground area features through a dual-stream network.
[0098] The goal of foreground segmentation is to separate the target region of interest (i.e., the foreground) from the background region in the input image. For example, in a photo containing a person, the foreground is the person, and the background is everything else. Using foreground segmentation techniques, we can obtain an image containing only the target region, known as the target region image.
[0099] A two-stream network is a common deep learning architecture that typically consists of two parallel subnetworks. In this step, one subnetwork processes the original input image and extracts its features, while the other subnetwork processes the foreground image and extracts its features. These two subnetworks can share some parameters or structures, or they can be designed independently to obtain image feature information from different perspectives.
[0100] S702: Generate channel weights based on the foreground area features, and perform weighted fusion with the original image features to obtain fused features.
[0101] Channel weights are parameters used to adjust the importance of different channels in a feature map. Generating channel weights based on foreground features means assigning a weight to each channel of the original image features based on the characteristic information of the foreground image. These weights reflect the importance of different channels in describing the target features.
[0102] The generated channel weights are weighted and fused with the original image features. This involves multiplying the weights by the corresponding feature channels and then integrating the features of all channels. This aims to highlight the parts of the original image features that are relevant to the foreground area and suppress the irrelevant parts, thereby obtaining fused features that are more focused on the target area.
[0103] S703 , extracting a global feature vector and a local feature vector from the fused features through global feature extraction and local feature extraction respectively.
[0104] Global feature extraction aims to extract feature vectors that can reflect the overall information of the image from the fused features, such as the overall color distribution, shape outline, etc. Local feature extraction focuses on local details of the image, such as texture, edges, etc., and extracts local feature vectors.
[0105] The specific implementation process of steps S701-S703 is similar to the implementation process in the above embodiment and will not be repeated here.
[0106] S704: Use the cross entropy loss function to optimize the global features and local feature vectors, adjust the model parameters through iterative training until convergence, and complete the model training.
[0107] The cross-entropy loss function is a commonly used loss function that measures the difference between a model's predictions and the true labels. In image recognition tasks, the true labels are often the image's category information. Model performance can be evaluated by calculating the cross-entropy loss between the predicted global and local feature vectors and the true labels.
[0108] In the embodiment of the present application, the loss function is used to treat each car or person's identity as a category, and the network parameters are optimized to classify all car images or person images. In the embodiment of the present application, the car recognition is used as an example to introduce each car's identity as a category. Assuming that there are a total of P different car IDs, all the global and local feature vectors H after dimensionality reduction are processed. i,j Input into a separate classifier to predict the label score for that feature i is used to distinguish global or local features, i=g represents global features, i=l represents local features, and j is used to index the slice feature in the case of local features. The calculation formula for the label score is as follows:
[0109]
[0110] in represents the predicted label score; W i,j is the vector H i,j Corresponding to the learning weights of the fully connected layer, it is used to transfer features to H i,j Mapped to different category spaces; P represents the total number of car IDs, Denotes the eigenvector H i,j The set of weights corresponding to category c.
[0111] Then, the ID loss is calculated based on these label scores, and the classification accuracy of the model for car images is improved by optimizing the network parameters. The corresponding sum of cross entropy losses:
[0112]
[0113] Where N represents the total number of images in the current batch, and y represents the true label of the input image. Based on the calculated ID loss, the model parameters are adjusted using the backpropagation algorithm to gradually reduce the loss. This process is repeated until the model performance reaches a satisfactory level, which is when the loss function converges. When the model converges, it means that it has learned effective feature representation and classification capabilities and can be used for practical image recognition tasks. Assume that the current batch has N = 2 images. For the first image, it has a global feature vector and two local feature vectors (i is global and local, j is 1 and 2). Assume that the calculated cross-entropy losses are L1,1 = 0.2 and L1,2 = 0.3, and the cross-entropy loss for the global feature vector is Lg,1 = 0.1. For the second image, it also has a global feature vector and two local feature vectors. Assume that the cross-entropy losses are L2,1 = 0.4 and L2,2 = 0.5, and the cross-entropy loss for the global feature vector is Lg,2 = 0.2. Then the ID loss LID = 21 × ((0.1 + 0.2 + 0.3) + (0.2 + 0.4 + 0.5)) = 21 × (0.6 + 1.1) = 0.85. The model will gradually reduce this loss value by adjusting the parameters.
[0114] The above are some specific implementations of a method for object re-identification based on foreground segmentation provided in the embodiment of the present application. Based on this, the present application also provides a corresponding device. The device provided in the embodiment of the present application will be introduced from the perspective of functional modularization.
[0115] Figure 8 This is a schematic diagram of the structure of a target re-identification device based on foreground segmentation provided in an embodiment of the present application. Figure 8 As shown, the object re-identification device 800 based on foreground segmentation provided in an embodiment of the present application includes:
[0116] The foreground image extraction unit 810 is used to obtain an image to be identified and perform foreground segmentation processing on the image to be identified to obtain a foreground image;
[0117] A first feature extraction unit 820 is configured to perform feature extraction on the image to be identified and the foreground image using a dual-stream network to obtain a feature map of the image to be identified and a feature map of the foreground image;
[0118] A fusion unit 830 is configured to construct feature channel weights based on the feature map of the foreground image, and perform feature fusion on the feature channel weights and the feature map of the image to be identified to obtain a fused feature map;
[0119] A second feature extraction unit 840 is configured to extract features from the fused feature map using a global feature extraction method and a local feature extraction method to obtain global features and local features;
[0120] A feature splicing unit 850 is configured to splice the global features and the local features as an object to be identified in the image to be identified;
[0121] The comparison unit 860 is configured to use the trained prediction model to calculate the similarity of features between the target to be identified and all images to be compared.
[0122] In one implementation of the embodiment of the present application, the foreground image extraction unit is specifically configured to:
[0123] Processing the image to be identified using foreground segmentation technology to obtain a binary mask of the target object;
[0124] The binary mask is fused with the image to be identified to obtain the foreground image.
[0125] In one implementation of the embodiment of the present application, the fusion unit constructing feature channel weights based on the feature map of the foreground image includes:
[0126] Performing a global average pooling operation on the feature map of the foreground image;
[0127] The compressed feature vector is input into the 1x1 convolution layer to reduce the number of feature channels;
[0128] The ReLu activation function is used to enhance the nonlinear characteristics of the channel, and then the fully connected layer is input to restore the number of feature channels;
[0129] Use the sigmoid activation function to excite the channel and obtain the enhanced feature channel weights.
[0130] In one implementation of the embodiment of the present application, when the comparison unit uses the trained prediction model to calculate the similarity of features between the target to be identified and all images to be compared, the method includes:
[0131] Calculate the Euclidean distance between the target to be identified and each image to be compared;
[0132] Sort all the images to be compared according to the calculated Euclidean distance to obtain a similarity sorted list;
[0133] The final search results are returned in a list sorted by similarity, with images with high similarity at the front.
[0134] In one implementation of the embodiment of the present application, the apparatus further includes a model training unit for model training, and the model training process includes:
[0135] Perform foreground segmentation on the input image to obtain the target area image, and extract the original image features and foreground area features through a two-stream network respectively;
[0136] Generating channel weights based on the foreground area features, and performing weighted fusion with the original image features to obtain fused features;
[0137] The fusion features are extracted through global features and local features to extract global feature vectors and local feature vectors;
[0138] The cross entropy loss function is used to optimize the global features and local feature vectors, and the model parameters are adjusted through iterative training until convergence to complete the model training.
[0139] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.
[0140] The device includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the method described in any embodiment of the present application.
[0141] The computer storage medium stores code, and when the code is executed, the device executing the code implements the method described in any embodiment of the present application.
[0142] Through the description of the above implementation methods, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment methods can be implemented by means of software plus a general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment of the present application or certain parts of the embodiments.
[0143] It is understandable that in the specific implementation of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved, when the above embodiments of this application are applied to specific products or technologies, need to obtain user permission or consent, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0144] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0145] It should also be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and apparatus embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0146] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A target re-identification method based on foreground segmentation, characterized in that: The method comprises: Acquire an image to be identified, and perform foreground segmentation processing on the image to be identified to obtain a foreground image; Using a dual-stream network to extract features from the image to be identified and the foreground image, respectively, to obtain a feature map of the image to be identified and a feature map of the foreground image; Constructing feature channel weights based on the feature map of the foreground image, and performing feature fusion on the feature channel weights and the feature map of the image to be identified to obtain a fused feature map; Performing feature extraction on the fused feature map using a global feature extraction method and a local feature extraction method respectively to obtain global features and local features; splicing the global features and the local features together as the target to be identified in the image to be identified; The trained prediction model is used to calculate the similarity of features between the target to be identified and all images to be compared.
2. The method according to claim 1, characterized in that The foreground image obtained by processing the image to be identified includes: Processing the image to be identified using foreground segmentation technology to obtain a binary mask of the target object; The binary mask is fused with the image to be identified to obtain the foreground image.
3. The method according to claim 1, characterized in that The constructing feature channel weights based on the feature map of the foreground image includes: Performing a global average pooling operation on the feature map of the foreground image; The compressed feature vector is input into the 1x1 convolution layer to reduce the number of feature channels; The ReLu activation function is used to enhance the nonlinear characteristics of the channel, and then the fully connected layer is input to restore the number of feature channels; Use the sigmoid activation function to excite the channel and obtain the enhanced feature channel weights.
4. The method according to claim 1, wherein When using the trained prediction model to calculate the similarity of features between the target to be identified and all images to be compared, it includes: Calculate the Euclidean distance between the target to be identified and each image to be compared; Sort all the images to be compared according to the calculated Euclidean distance to obtain a similarity sorted list; The final search results are returned in a list sorted by similarity, with images with high similarity at the front.
5. The method according to claim 1, wherein The model training process includes: Perform foreground segmentation on the input image to obtain the target area image, and extract the original image features and foreground area features through a two-stream network respectively; Generating channel weights based on the foreground area features, and performing weighted fusion with the original image features to obtain fused features; The fusion features are extracted through global features and local features to extract global feature vectors and local feature vectors; The cross entropy loss function is used to optimize the global features and local feature vectors, and the model parameters are adjusted through iterative training until convergence to complete the model training.
6. A target re-identification device based on foreground segmentation, characterized in that: The device comprises: a foreground image extraction unit, configured to obtain an image to be identified and perform foreground segmentation processing on the image to be identified to obtain a foreground image; a first feature extraction unit, configured to perform feature extraction on the image to be identified and the foreground image respectively using a dual-stream network to obtain a feature map of the image to be identified and a feature map of the foreground image; a fusion unit, configured to construct a feature channel weight based on the feature map of the foreground image, and perform feature fusion on the feature channel weight and the feature map of the image to be identified to obtain a fused feature map; A second feature extraction unit is used to extract features from the fused feature map using a global feature extraction method and a local feature extraction method to obtain global features and local features; A feature splicing unit, configured to splice the global features and the local features as an object to be identified in the image to be identified; The comparison unit is used to calculate the similarity of features between the target to be identified and all images to be compared using the trained prediction model.
7. The device according to claim 6, characterized in that The foreground image extraction unit is specifically used for: Processing the image to be identified using foreground segmentation technology to obtain a binary mask of the target object; The binary mask is fused with the image to be identified to obtain the foreground image.
8. The device according to claim 7, characterized in that The fusion unit constructs feature channel weights based on the feature map of the foreground image, including: Performing a global average pooling operation on the feature map of the foreground image; The compressed feature vector is input into the 1x1 convolution layer to reduce the number of feature channels; The ReLu activation function is used to enhance the nonlinear characteristics of the channel, and then the fully connected layer is input to restore the number of feature channels; Use the sigmoid activation function to excite the channel and obtain the enhanced feature channel weights.
9. A computing device, characterized in that The computing device includes: a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1 to 5 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Ultrasonic image measurement key detection method and device and electronic equipment
CN121527058A