Pedestrian re-identification method and device

By acquiring and fusing the enhanced features of visible light and infrared images, the problem of low recognition accuracy in pedestrians being blocked or complex scenes is solved, and a higher pedestrian re-identification accuracy is achieved.

CN120388329APending Publication Date: 2025-07-29DAWNING INT INFORMATION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510438127.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing pedestrian re-identification method based on visible light-infrared images is not ideal for pedestrians being blocked or complex scenes, resulting in a decrease in recognition accuracy.

Method used

By acquiring the enhancement features of the first image and the enhancement features of the second image, we obtain the fusion features, make full use of the complementary information of the two images, suppress noise interference, and enhance the response of the key information of the visible area.

Benefits of technology

The accuracy of pedestrian re-identification is improved, and effective information in the image can be extracted and utilized more effectively to generate high-quality discriminant features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388329A_ABST
    Figure CN120388329A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and discloses a pedestrian re-recognition method and device. The method comprises the steps that a first image and a second image are acquired, pedestrians in the first image are not shielded, the second image is obtained by processing the first image, and the pedestrians in the second image are partially shielded; obtaining an enhanced feature of the first image according to the image feature of the first image and the global attention feature of the first image; and according to the image feature of the second image and the global attention feature of the second image, obtaining an enhanced feature of the second image, obtaining a fusion feature at least according to the enhanced feature of the first image and the enhanced feature of the second image; and determining an image matched with the fusion feature in the image library as a pedestrian re-identification result. Thus, through mutual learning of the enhanced features of the first image and the enhanced features of the second image, effective fusion features are obtained, and then the accuracy of pedestrian re-identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a pedestrian re-identification method and device. Background Art

[0002] Pedestrian re-identification aims to match an image in an image library that has the same identity as the pedestrian in the image. Currently, it is mainly divided into pedestrian re-identification under a single data source and pedestrian re-identification under multi-source data. Among them, the effect of pedestrian re-identification under a single data source is limited by the discriminable information of the single data source. Pedestrian re-identification under multi-source data usually performs pedestrian re-identification based on visible light-infrared data. Visible light images have rich texture and color features, and infrared images can capture effective information under low-light conditions. The two are complementary. Therefore, pedestrian re-identification based on visible light-infrared images widely exists in the field of police investigation. The police can obtain the identity information of a pedestrian from a gallery (visible light image / infrared image) by matching the detected image (infrared image / visible light image) of the pedestrian.

[0003] However, in the existing research on pedestrian re-identification based on visible light-infrared images, there are often problems such as pedestrians being deliberately blocked or blocked by a complex scene environment, resulting in unsatisfactory recognition effects. For example, when investigating a case, criminals usually commit crimes at night and cover their bodies with obstacles, resulting in the camera capturing an infrared blocked image of the criminal; or, in a crowded public place with a chaotic background, pedestrians are blocked by obstacles such as other people, vehicles, and chairs in the scene. When a pedestrian in an image is blocked by an obstacle, the information of the pedestrian in the blocked image is lost, and this information loss is irreversible, significantly reducing the effective information in the finally extracted image features, thereby reducing the recognition accuracy. Summary of the Invention

[0004] Embodiments of the present invention provide a pedestrian re-identification method and device. By mutually learning the enhanced features of a first image and the enhanced features of a second image, effective fusion features are obtained, thereby improving the accuracy of pedestrian re-identification.

[0005] In a first aspect, an embodiment of the present invention provides a person re-identification method, which includes: obtaining a first image and a second image, where the pedestrian in the first image is not occluded, and the second image is obtained by processing the first image, and the pedestrian in the second image is partially occluded; obtaining an enhanced feature of the first image according to the image feature and the global attention feature of the first image; and obtaining an enhanced feature of the second image according to the image feature and the global attention feature of the second image; obtaining a fusion feature at least according to the enhanced feature of the first image and the enhanced feature of the second image; and determining an image in the image library that matches the fusion feature as the result of the person re-identification.

[0006] In the above manner, by obtaining a fusion feature according to the enhanced feature of the first image and the enhanced feature of the second image, it is possible to fully extract the effective information of the first image and the second image, facilitate suppressing noise interference, enhance the response of key information in the visible region, generate high-quality discriminative features, and make full use of the complementary information of the first image and the second image, thereby improving the accuracy of person re-identification.

[0007] In an optional implementation manner, the obtaining a fusion feature at least according to the enhanced feature of the first image and the enhanced feature of the second image includes: obtaining a fusion feature according to the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image.

[0008] In the above manner, by obtaining a fusion feature according to the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image, it is possible to combine the effective information of the first image and the second image and extract discriminative features.

[0009] In an optional implementation manner, the obtaining a fusion feature according to the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image includes: multiplying the enhanced feature of the second image and the global attention feature of the second image, and adding the result of the multiplication to the enhanced feature of the first image to obtain the fusion feature.

[0010] By using the above method, multiplying the enhanced feature of the second image and the global attention feature of the second image can effectively suppress noise interference and extract the features of the unoccluded part of the pedestrian in the second image, and then adding it to the enhanced feature of the first image can fully extract the effective information in the first image and the second image.

[0011] In an alternative embodiment, the method further includes: determining a loss function according to the cross-entropy loss of the fused features, the triplet loss of the fused features, and the cross-entropy loss of the result obtained by multiplying the fused features and the global attention features of the first image, where the loss function is used to optimize the sample distribution in the feature space so that the positive sample features are aggregated and the negative sample features are separated.

[0012] In an alternative embodiment, the loss function satisfies the following formula: L = ω1·L cls + ω2·L M-cls + ω3·L tri , where L cls is the cross-entropy loss of the fused features, L tri is the triplet loss of the fused features, L M-cls is the cross-entropy loss of the result obtained by multiplying the fused features and the global attention features of the first image, and ω1, ω2, and ω3 are the weights of L cls , L M-cls and L tri respectively.

[0013] In an alternative embodiment, the first image belongs to a first type, the second image belongs to a second type, and the image library includes a plurality of images belonging to the first type and a plurality of images belonging to the second type; wherein, the first type is a visible light image and the second type is an infrared image.

[0014] In a second aspect, an embodiment of the present invention provides a person re-identification device, which includes: an acquisition module for acquiring a first image and a second image, where the pedestrian in the first image is not blocked, and the second image is processed from the first image, and the pedestrian in the second image is partially blocked; a processing module for obtaining enhanced features of the first image according to the image features of the first image and the global attention features of the first image; and obtaining enhanced features of the second image according to the image features of the second image and the global attention features of the second image; and obtaining fused features at least according to the enhanced features of the first image and the enhanced features of the second image; a matching module for determining, as the result of the person re-identification, the image in the image library that matches the fused features.

[0015] In an alternative embodiment, the processing module is further configured to obtain fused features according to the enhanced features of the first image, the enhanced features of the second image, and the global attention features of the second image.

[0016] In an alternative embodiment, the processing module is specifically configured to perform a dot product on the enhanced features of the second image and the global attention features of the second image, and add the result of the dot product to the enhanced features of the first image to obtain the fusion features.

[0017] In an alternative embodiment, the processing module is further configured to determine a loss function according to the cross-entropy loss of the fusion features, the triplet loss of the fusion features, and the cross-entropy loss of the result of performing a dot product on the fusion features and the global attention features of the first image. The loss function is used to optimize the sample distribution in the feature space so that the positive sample features are aggregated and the negative sample features are separated.

[0018] In an alternative embodiment, the loss function satisfies the following formula: L = ω1·L cls + ω2·L M-cls + ω3·L tri ; where L cls is the cross-entropy loss of the fusion features, L tri is the triplet loss of the fusion features, L M-cls is the cross-entropy loss of the result of performing a dot product on the fusion features and the global attention features of the first image, and ω1, ω2, and ω3 are the weights of L cls , L M-cls and L tri respectively.

[0019] In an alternative embodiment, the first image belongs to a first type, the second image belongs to a second type, and the image library includes a plurality of images belonging to the first type and a plurality of images belonging to the second type; where the first type is a visible light image and the second type is an infrared image.

[0020] In a third aspect, an embodiment of the present invention provides a pedestrian re-identification device, including: a memory for storing a computer program; a processor for, when executing the computer program stored on the memory, performing the method described in the first aspect above according to the obtained program.

[0021] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored. When a computer reads and executes the computer program, the method described in the first aspect above is executed.

[0022] In a fifth aspect, an embodiment of the present invention provides a computer program product. When a computer reads and executes the computer program product, the method described in the first aspect above is executed. Description of the Drawings

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 It is a system architecture diagram provided by an embodiment of the present application;

[0025] Figure 2 It is the overall architecture diagram of the person re-identification method provided by an embodiment of the present application;

[0026] Figure 3 It is the flowchart corresponding to the person re-identification method provided by an embodiment of the present application;

[0027] Figure 4 It is the flowchart of the first feature enhancement module provided by an embodiment of the present application;

[0028] Figure 5 It is the flowchart of the second feature enhancement module provided by an embodiment of the present application;

[0029] Figure 6 It is the structural schematic diagram of the modality information fusion module provided by an embodiment of the present application;

[0030] Figure 7 It is the structural schematic diagram of the modality information fusion module provided by an embodiment of the present application;

[0031] Figure 8 It is the structural schematic diagram of the modality information fusion module provided by an embodiment of the present application;

[0032] Figure 9 It is the structural schematic diagram of a person re-identification device provided by an embodiment of the present application;

[0033] Figure 10 It is the structural schematic diagram of a person re-identification device provided by an embodiment of the present application. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. In the embodiments of the present invention, "a plurality of" means two or more. Terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be construed as indicating or implying relative importance, nor can they be construed as indicating or implying an order.

[0035] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices. The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to that element.

[0036] In the research of person re-identification, both the global context information and the local fine-grained information of images are crucial. Existing person re-identification methods can be divided into two categories. The first category of methods is person re-identification based on visible light images, and the second category of methods is person re-identification based on visible light-infrared images.

[0037] Person re-identification methods based on visible light images usually adopt the following two implementation methods, including Implementation Method 1 and Implementation Method 2. Implementation Method 1 ignores the occluded body regions and extracts fine-grained features in the unoccluded body regions. This implementation method usually uses pre-trained auxiliary models (such as pose estimation and human parsing) because they can effectively prevent the occluded parts from being misidentified as part of the body. However, this implementation method using auxiliary models has a disadvantage that they are easily affected by the domain gap. Especially in the case of severe occlusion, incorrect key points or segmentation results are likely to mask the effective fine-grained features. Implementation Method 2 is to extract complete body features by obtaining information for predicting the occluded body regions. Usually, the node relationship of the visible region or the information of the nearest neighbor image is used to recover the information of the occluded region. However, although the entire image is retained in this implementation method, the features recovered from the occluded regions are often incomplete, lacking consistency and reliability.

[0038] Visible light image-based pedestrian re-identification utilizes the complementarity between visible light and infrared images, which can improve the performance of pedestrian re-identification. However, existing visible light-infrared occluded pedestrian re-identification faces some problems, resulting in inaccurate recognition results. For example, since obstacles will occlude some parts of the pedestrian, the effective information in the final features is significantly reduced. Another example is that obstacles will introduce blurred noise, which will have a negative impact during the feature extraction process and lead to semantic misalignment during the matching process.

[0039] Based on this, the embodiments of this application provide a pedestrian re-identification method. By obtaining the fusion features according to the enhanced features of the first image and the enhanced features of the second image, and determining the image in the image library that matches the fusion features as the result of pedestrian re-identification, the complementary information of the first image and the second image can be fully utilized to improve the accuracy of pedestrian re-identification.

[0040] Figure 1 This is a system architecture diagram provided by the embodiments of this application, including a terminal device 101 and a server 102. Among them, the terminal device 101 is used to obtain the first image and the second image, and the server 102 is used to process the first image and the second image and train the pedestrian re-identification model. The terminal device 101 pre-installs a business application, where the business application is a client application, a web version application, a mini-program application, etc. The terminal device 101 can be a laptop, a desktop computer, etc., but is not limited thereto. The server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal device 101 and the server 102 can be connected by wire or wirelessly.

[0041] Figure 2 This is the overall architecture diagram of the pedestrian re-identification method provided by the embodiments of this application. As Figure 2As shown in the figure, the overall architecture of the method includes a preprocessing module, a first branch, a second branch, and a modality information fusion module. Among them, the first branch includes: a first feature extraction module and a first feature enhancement module; the second branch includes a second feature extraction module and a second feature enhancement module. The first feature extraction module in the first branch extracts features from the first image to obtain the image features of the first image and the global attention features of the first image; the first feature enhancement module in the first branch obtains the enhanced features of the first image according to the image features of the first image and the global attention features of the first image. The preprocessing module processes the first image to obtain a second image, and the second feature extraction module in the second branch extracts features from the second image to obtain the image features of the second image and the global attention features of the second image; the second feature enhancement module in the second branch obtains the enhanced features of the second image according to the image features of the second image and the global attention features of the second image. The modality information fusion module obtains fusion features at least according to the enhanced features of the first image and the enhanced features of the second image, matches the fusion features obtained by the modality information fusion module with the images in the image library, and determines the image after matching as the pedestrian re-identification result.

[0042] The method provided by the embodiments of the present application will be described in detail below in conjunction with specific embodiments.

[0043] Figure 3 It is a flowchart corresponding to the pedestrian re-identification method provided by the embodiments of the present application. As Figure 3 shown, the method includes the following steps:

[0044] Step 301, obtain a first image and a second image.

[0045] Among them, the first image can be a visible light image including a pedestrian, and the pedestrian in the first image is not occluded. The second image can be an infrared image obtained by processing the first image, and the pedestrian in the second image is partially occluded. Among them, processing the first image to obtain the second image can be performed by the above-mentioned Figure 2 preprocessing module. The implementation of the preprocessing module is as follows:

[0046] S1: Extract occlusion blocks from the public dataset. Among them, the occlusion blocks are occluders in the public dataset or other pedestrians other than non-target pedestrians, and the public dataset can be Occluded-SYSU-MM01 and Occluded-RegDB.

[0047] S2: Paste the occlusion block in the specified area of the first image according to a preset rule to generate an occluded image of the first image. For example, if the occlusion block extracted from the public dataset is a vehicle or a billboard, etc., paste the occlusion in the lower area of the first image to generate an occluded image of the first image; if the occlusion block extracted from the public dataset is an umbrella, paste the umbrella in the upper area of the first image to generate an occluded image of the first image; if the occlusion block extracted from the public dataset is other pedestrians other than the target pedestrian, paste the other pedestrians other than the non-target pedestrian in the left area or the right area of the first image to generate an occluded image of the first image.

[0048] S3: Perform color conversion and expansion operations on the occluded image of the first image to obtain a second image. Among them, perform color conversion on the occluded image of the first image so that the converted occluded image gradually approaches the pedestrian image captured by an infrared camera under a night scene. In order to increase the complexity and diversity of the generated second image, an expansion operation can be performed on the image after color conversion to obtain the second image. There are various expansion operations, such as random cropping and random erasing.

[0049] It can be understood that, in order to ensure the consistency of information between the first image and the second image and promote information interaction between the two branches, the expansion operation in the embodiments of this application may not include flipping the image. In this way, the obtained second image can help the pedestrian re-identification model of this application better learn and align the features of the first image and the second image.

[0050] Step 302: Obtain the enhanced feature of the first image according to the image feature of the first image and the global attention feature of the first image; and obtain the enhanced feature of the second image according to the image feature of the second image and the global attention feature of the second image.

[0051] Exemplarily, the first feature extraction module in the first branch extracts features from the first image to obtain the image feature of the first image and the global attention feature of the first image; the first feature enhancement module in the first branch obtains the enhanced feature of the first image according to the image feature of the first image and the global attention feature of the first image. And, the second feature extraction module in the second branch extracts features from the second image to obtain the image feature of the second image and the global attention feature of the second image; the second feature enhancement module in the second branch obtains the enhanced feature of the second image according to the image feature of the second image and the global attention feature of the second image.

[0052] The first branch and the second branch are introduced below respectively.

[0053] (1) First branch

[0054] Extract features from the first image through the first feature extraction module in the first branch to obtain the image features of the first image and the global attention features of the first image. The specific implementation is as follows:

[0055] A1: Divide the first image into multiple image patches of the same size with overlap according to a preset formula. The preset formula is as follows:

[0056]

[0057] where P is the side length of the image patch of the first image, S is the step size of the sliding window, H is the height of the first image, and W is the width of the first image.

[0058] For example, assume the first image is 128×128, the step size S is 8, and the side length of the image patch of the first image is 16. At this time, D calculated based on formula (1) is 255, that is, the first image is divided into 255 16×16 image patches.

[0059] A2: Use linear projection to map each of the multiple image patches of the first image of the same size into an embedding vector with a fixed length. The specific formula is as follows:

[0060] F0 = [X cls ; τ(X1); τ(X2); …; τ(X D )] + P E + δC E (2)

[0061] where X cls is an additional learnable classification marker at the front of the embedding vector, and this marker can represent the features of the entire image; P E is the position information, C E is the camera information, and δ is the weight of C E .

[0062] A3: Pass the embedding vector obtained in A2 through the ViT encoder to obtain the image features of the first image. The image features of the first image include the global features and local features of the first image. The global features are suitable for capturing the body structure and overall contour of the entire pedestrian in the first image, and the local features can reflect the pedestrian's texture, color, and other minute detail information. Strategies that rely only on global features may not be able to capture key local detail information, and these local information is crucial for refined recognition or for distinguishing different pedestrians with similar appearances in complex scenarios.

[0063] A4: Process the intermediate features obtained from some layers of the ViT encoder to obtain the first global attention feature of the image. For example, the intermediate features F2, F4, F 10 , F 12 are processed to obtain the global attention feature of the first image. Among them, the formula for the global attention feature is as follows:

[0064] Mask = σ(Avgpool(Γ(F ll ))), l = 2, 4, 10, 12 (3)

[0065]

[0066] where Mask represents the global attention feature of the first image, ψ is the reshape operation, σ is the activation function, Γ is the convolutional layer, and F l is the intermediate feature obtained from the l-th layer.

[0067] Since the global feature is easily affected by occluders and complex backgrounds and has poor adaptability to changes in the image (such as object size and pose changes). Especially when pedestrians appear in different ways in different images, the global feature representation lacking local information may lead to a decline in recognition performance. Therefore, the comprehensive mining of the global and local features of the image is crucial for constructing a highly discriminative feature representation. However, most existing Transformer-based methods often only learn global features, local features, or simply concatenate the two, but these methods usually fail to fully explore the potential connection between global and local information in pedestrian images, resulting in the generated feature representation lacking sufficient discriminability and being easily affected by occlusion noise. Therefore, in this application, the first feature enhancement module in the first branch obtains the enhanced feature of the first image according to the global feature, local feature, and global attention feature of the first image. Figure 4 This is the flowchart of the first feature enhancement module provided by the embodiment of this application, as Figure 4 shown, and the specific implementation is as follows:

[0068] B1: Expand the dimension of the global attention feature of the first image through a fully connected layer and a Sigmoid activation layer to obtain the local attention feature of the first image. For example, the dimension of the global attention feature of the first image is 1×D, the dimension after passing through the fully connected layer is 1×N, and the dimension of the local attention feature of the first image obtained after passing through the Sigmoid activation layer is 1×N×D. The dimension of the local attention feature of the first image after dimension expansion is the same as the dimension of the local feature of the first image.

[0069] B2: Multiply the global attention feature of the first image and the global feature of the first image, and use a residual module to add the feature after multiplication to the global feature of the first image to obtain the enhanced global feature of the first image.

[0070] B3: Multiply the local attention feature of the first image and the local feature of the first image, and use a residual module to add the feature after multiplication to the local feature of the first image to obtain the enhanced local feature of the first image.

[0071] B4: Concatenate the enhanced global feature of the first image and the enhanced local feature of the first image, and use a module containing T Transformer layers to obtain the enhanced feature of the first image.

[0072] (2) The second branch

[0073] Extract features from the second image through the second feature extraction module in the second branch to obtain the image feature of the second image and the global attention feature of the second image. The specific implementation is as follows:

[0074] C1: Cut the second image into multiple image patches of the same size and with overlap of the first image according to the above formula (1). For example, assume the second image is 128×128, the stride S is 8, and the side length of the image patch of the second image is 16. At this time, D calculated based on formula (1) is 255, that is, the second image is cut into 255 16×16 image patches.

[0075] C2: Use linear projection to map each of the multiple image patches of the second image of the same size into an embedding vector with a fixed length.

[0076] C3: Pass the embedding vector obtained in C2 through the ViT encoder to obtain the image feature of the second image. The image feature of the second image includes the global feature and the local feature of the second image. The global feature is suitable for capturing the body structure and overall contour of pedestrians in the second image, and the local feature can reflect the obligation texture, color, and other tiny detail information of pedestrians.

[0077] C4: Process the intermediate features obtained from some layers of the ViT encoder to obtain the global attention feature of the second image. For example, the intermediate features F2, F4, F 10 、F 12 obtained from the 2nd, 4th, 10th, and 12th layers of the ViT encoder can be processed to obtain the global attention feature of the second image.

[0078] The second feature enhancement module in the second branch obtains the enhanced feature of the second image according to the image feature of the second image and the global attention feature of the second image. Figure 5It is a flowchart of the second feature enhancement module provided by the embodiments of this application. As Figure 5 shown, the specific implementation is as follows:

[0079] D1: Dimension-expand the global attention feature of the second image through a fully-connected layer and a Sigmoid activation layer to obtain the local attention feature of the second image. For example, the dimension of the global attention feature of the second image is 1×D, the dimension after passing through the fully-connected layer is 1×N, and the dimension of the local attention feature of the second image obtained after passing through the Sigmoid activation layer is 1×N×D. The dimension of the local attention feature of the second image after dimension expansion is the same as the dimension of the local feature of the second image.

[0080] D2: Perform dot multiplication on the global attention feature of the second image and the global feature of the second image, and use a residual module to add the feature after dot multiplication to the global feature of the second image to obtain the enhanced global feature of the second image.

[0081] D3: Perform dot multiplication on the local attention feature of the second image and the local feature of the second image, and use a residual module to add the feature after dot multiplication to the local feature of the second image to obtain the enhanced local feature of the second image.

[0082] D4: Concatenate the enhanced global feature of the second image and the enhanced local feature of the second image, and use a structure containing T Transformer layers to obtain the enhanced feature of the second image.

[0083] Step 303, obtain a fusion feature based on at least the enhanced feature of the first image and the enhanced feature of the second image.

[0084] Exemplarily, a fusion feature can be obtained based on the enhanced feature of the first image and the enhanced feature of the second image; or, a fusion feature can be obtained based on the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image; or, a fusion feature can be obtained based on the enhanced feature of the first image, the enhanced feature of the second image, the global attention feature of the first image, and the global attention feature of the second image.

[0085] During the training of the person re-identification model, the model is optimized through a loss function, and the network parameters are dynamically adjusted through backpropagation to optimize the sample distribution in the feature space, so that the positive sample features are aggregated and the negative sample features are separated. The larger the value of the loss function, the greater the distance between the fusion feature and the positive sample, and the less accurate the fusion feature extracted by the model. The gradient descent algorithm can be used to perform backpropagation on the gradient of the model parameters and adjust the model parameters, thereby training the model. By minimizing the value of the loss function, the accuracy of the model can be improved. Different loss functions are set for different fusion features.

[0086] The following is an explanation for different situations respectively.

[0087] (1) Obtain a fusion feature based on the enhanced feature of the first image and the enhanced feature of the second image.

[0088] Figure 6 The following is a schematic structural diagram of the modality information fusion module provided by the embodiment of the present application. As Figure 6 shown, add the enhanced feature of the first image and the enhanced feature of the second image to obtain a fusion feature.

[0089] Regarding the above-mentioned obtaining of the fusion feature based on the enhanced feature of the first image and the enhanced feature of the second image, at this time, determine the loss function according to the cross-entropy loss of the fusion feature and the triplet loss of the fusion feature. Before calculating the cross-entropy, the fusion feature can be normalized and passed through a fully connected layer. The loss function satisfies the following formula:

[0090] L = α1·L cls +α2·L tri (4)

[0091] where, L cls is the cross-entropy loss of the fusion feature, L tri is the triplet loss of the fusion feature, and α1 and α2 are the weights of L cls and L tri respectively. α1 and α3 can be set according to actual needs, and α1 and α2 can both be set to 0.5.

[0092] The cross-entropy loss L cls of the fusion feature satisfies the following formula:

[0093]

[0094] where, P ic is the category to which the image x i belongs to c, and y(x ic ) is the label value of the image x i .

[0095] The triplet loss L tri of the fusion feature satisfies the following formula:

[0096]

[0097] where, is the anchor image, represents the positive sample image with the same identity as (that is, the pedestrians in the images are the same person), represents the one with They are negative sample images of different identities (i.e., the pedestrians in the images are different people).

[0098] (2) Obtain a fused feature based on the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image.

[0099] Figure 7 It is a schematic structural diagram of the modality information fusion module provided by an embodiment of the present application. As Figure 7 shown, multiply the enhanced feature of the second image and the global attention feature of the second image, and add the result of the multiplication to the enhanced feature of the first image to obtain a fused feature. By multiplying the enhanced feature of the second image and the global attention feature of the second image, the features of the pedestrians in the second image that are not occluded can be obtained. Adding them to the enhanced feature of the first image to obtain a fused feature can more effectively extract the feature information of the pedestrians.

[0100] Regarding obtaining the fused feature based on the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image, at this time, determine the loss function according to the cross-entropy loss of the fused feature, the triplet loss of the fused feature, and the cross-entropy loss of the result after multiplying the fused feature and the global attention feature of the first image. The fused feature can be normalized and passed through a fully connected layer before calculating the cross-entropy. The loss function satisfies the following formula:

[0101] L = ω1·L cls + ω2·L M-cls + ω3·L tri (7)

[0102] where, L cls is the cross-entropy loss of the fused feature, L tri is the triplet loss of the fused feature, L M-cls is the cross-entropy loss of the result after multiplying the fused feature and the global attention feature of the first image (i.e., obtained by substituting the result of multiplying the fused feature and the global attention feature of the first image into the formula for calculation), and ω1, ω2, and ω3 are the weights of L cls , L M-cls , and L tri , respectively. ω1, ω2, and ω3 are set to satisfy ω1 + ω2 = ω3. For example, ω1 is set to 0.5, ω2 is set to 0.5, and ω3 is set to 1.

[0103] (3) Obtain a fused feature based on the enhanced feature of the first image, the enhanced feature of the second image, the global attention feature of the first image, and the global attention feature of the second image.

[0104] Figure 8The structural schematic diagram of the modal information fusion module provided by the embodiment of the present application is as follows. Figure 8 As shown, the enhanced feature of the first image and the global attention feature of the first image are multiplied point by point, the enhanced feature of the second image and the global attention feature of the second image are multiplied point by point, and the result of the point multiplication of the first image is added to the result of the point multiplication of the second image to obtain a fusion feature.

[0105] For obtaining the fusion feature according to the enhanced feature of the first image, the enhanced feature of the second image, the global attention feature of the first image, and the global attention feature of the second image, at this time, according to the cross-entropy loss of the fusion feature and the triplet loss of the fusion feature, a loss function is determined. Before calculating the cross-entropy, the fusion feature can be normalized and passed through a fully connected layer. The loss function satisfies the following formula:

[0106] L = μ1·L cls + μ2·L tri (8)

[0107] Wherein, L cls is the cross-entropy loss of the fusion feature, L tri is the triplet loss of the fusion feature, and μ1 and μ2 are the weights of L cls and L tri respectively. α1 and α3 can be set according to actual needs, and μ1 and μ2 can both be set to 0.5.

[0108] Step 304: Determine the image in the image library that matches the fusion feature as the result of pedestrian re-identification.

[0109] Exemplarily, the obtained fusion feature can be matched with the images in the image library, and the images with a similarity greater than a preset threshold are selected as the result of pedestrian re-identification.

[0110] Optionally, the first image belongs to the first type, the second image belongs to the second type, and the image library includes multiple images belonging to the first type and multiple images belonging to the second type; the first type is a visible light image, and the second type is an infrared image. For example, the first image is a visible light image, the second image is an infrared image, and the image library includes multiple visible light images and multiple infrared images.

[0111] Using the above method, first, a second image is obtained by performing a series of processes on the first image, expanding the samples of occluded data in the training set; second, based on the image features of the first image and the global attention features of the first image, the enhanced features of the first image are obtained; and, based on the image features of the second image and the global attention features of the second image, the enhanced features of the second image are obtained; it can suppress noise interference, enhance the response of key information in the visible area, and generate high-quality discriminative features. In addition, the modality information fusion module obtains fusion features at least based on the enhanced features of the first image and the enhanced features of the second image, realizes the interaction of information between the two branches, and improves the ability to extract effective features. Compared with the methods in the prior art that separately use global or local information, or simply splice these two types of information, the method of this application helps to extract more discriminative features, thereby improving the accuracy of pedestrian re-identification.

[0112] Based on the same technical concept, an embodiment of this application also provides a pedestrian re-identification device 9000, Figure 9 which is a schematic diagram of the pedestrian re-identification device provided by an embodiment of this application, as Figure 9 described, the device includes:

[0113] An acquisition module 901, configured to acquire a first image and a second image, where the pedestrian in the first image is not occluded, and the second image is obtained by processing the first image, and the pedestrian in the second image is partially occluded;

[0114] A processing module 902, configured to obtain the enhanced features of the first image based on the image features of the first image and the global attention features of the first image; and, obtain the enhanced features of the second image based on the image features of the second image and the global attention features of the second image; obtain fusion features at least based on the enhanced features of the first image and the enhanced features of the second image;

[0115] A matching module 903, configured to determine the image in the image library that matches the fusion features as the result of the pedestrian re-identification.

[0116] In an alternative embodiment, the processing module 902 is further configured to obtain fusion features based on the enhanced features of the first image, the enhanced features of the second image, and the global attention features of the second image.

[0117] In an alternative embodiment, the processing module 902 is further configured to perform a dot product on the enhanced features of the second image and the global attention features of the second image, and add the result of the dot product to the enhanced features of the first image to obtain the fusion features.

[0118] In an alternative embodiment, the processing module is further configured to determine a loss function according to the cross-entropy loss of the fusion features, the triplet loss of the fusion features, and the cross-entropy loss of the result after multiplying the fusion features by the global attention features of the first image. The loss function is used to optimize the sample distribution in the feature space, so that the positive sample features are aggregated and the negative sample features are separated.

[0119] In an alternative embodiment, the loss function satisfies the following formula: L = ω1·L cls +·2·L M-cls + ω3·L tri , where L cls is the cross-entropy loss of the fusion features, L tri is the triplet loss of the fusion features, L M-cls is the cross-entropy loss of the result after multiplying the fusion features by the global attention features of the first image, and ω1, ω2, and ω3 are the weights of L cls , L M-cls , and L tri , respectively.

[0120] In an alternative embodiment, the first image belongs to a first type, the second image belongs to a second type, and the image library includes multiple images belonging to the first type and multiple images belonging to the second type; wherein, the first type is a visible light image and the second type is an infrared image.

[0121] Based on the same technical concept, the embodiment of the present application further provides a schematic structural diagram of a person re-identification device, as shown in Figure 10 . The device 10000 includes at least one processor 1001 and a memory 1002 connected to at least one processor 1001. In the embodiment of the present application, the specific connection medium between the processor 1001 and the memory 1002 is not limited. Figure 10 Taking the example that the processor 1001 and the memory 1002 are connected by a bus. The bus can be divided into an address bus, a data bus, a control bus, etc. In the embodiment of the present application, the memory 1002 stores instructions executable by at least one processor 1001. By executing the instructions stored in the memory 1002, the at least one processor 1001 can implement the steps of the above person re-identification method.

[0122] Among them, the processor 1001 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 1002 and calling the data stored in the memory 1002, resource settings can be performed. Optionally, the processor 1001 may include one or more processing units. The processor 1001 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1001. In some embodiments, the processor 1001 and the memory 1002 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.

[0123] The processor 1001 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0124] The memory 1002, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1002 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 1002 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1002 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0125] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer program product includes: computer program code, which when run on a computer, causes the computer to execute any of the person re-identification methods described above. Since the principle of solving problems by the above computer-readable storage medium is similar to that of the person re-identification method, the implementation of the above computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0126] Based on the same inventive concept, an embodiment of the present application further provides a computer program product. The computer program product includes: computer program code, which when run on a computer, causes the computer to execute any of the person re-identification methods described above. Since the principle of solving problems by the above computer program product is similar to that of the person re-identification method, the implementation of the above computer program product can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0128] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0129] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0131] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A pedestrian re-identification method, characterized in that, The method includes: Obtaining a first image and a second image, where the pedestrian in the first image is not occluded, and the second image is obtained by processing the first image, and the pedestrian in the second image is partially occluded; Obtaining an enhanced feature of the first image according to the image feature and the global attention feature of the first image; and obtaining an enhanced feature of the second image according to the image feature and the global attention feature of the second image; Obtaining a fusion feature at least according to the enhanced feature of the first image and the enhanced feature of the second image; Determining the image in the image library that matches the fusion feature as the result of the pedestrian re-identification.

2. The method according to claim 1, characterized in that The obtaining a fusion feature at least according to the enhanced feature of the first image and the enhanced feature of the second image includes: Obtaining a fusion feature according to the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image.

3. The method according to claim 2, characterized in that, The obtaining a fusion feature according to the enhanced feature of the first image, the enhanced feature of the second image, and the global attention feature of the second image includes: Performing a dot product on the enhanced feature of the second image and the global attention feature of the second image, and adding the result of the dot product to the enhanced feature of the first image to obtain the fusion feature.

4. The method according to claim 2 or 3, characterized in that, The method further includes: Determining a loss function according to the cross-entropy loss of the fusion feature, the triplet loss of the fusion feature, and the cross-entropy loss of the result after performing a dot product on the fusion feature and the global attention feature of the first image, where the loss function is used to optimize the sample distribution in the feature space so that the positive sample features are aggregated and the negative sample features are separated.

5. The method according to claim 4, characterized in that, The loss function satisfies the following formula: L = ω1·L cls + ω2·L M-cls + ω3·L tri ; Among them, L cls is the cross-entropy loss of the fusion feature, L tri is the triplet loss of the fusion feature, L M-cls is the cross-entropy loss of the result after multiplying the fusion feature and the global attention feature of the first image, and ω1, ω2, and ω3 are respectively L cls , L M-cls and L tri 's weights.

6. The method according to claim 1, wherein The first image belongs to a first type, the second image belongs to a second type, and the image library includes multiple images belonging to the first type and multiple images belonging to the second type; Wherein, the first type is a visible light image, and the second type is an infrared image.

7. A pedestrian re-identification device, characterized in that, The device includes: An obtaining module, configured to obtain a first image and a second image, where the pedestrian in the first image is not occluded, and the second image is obtained by processing the first image, and the pedestrian in the second image is partially occluded; A processing module, configured to obtain an enhanced feature of the first image according to the image feature and the global attention feature of the first image; and obtain an enhanced feature of the second image according to the image feature and the global attention feature of the second image; obtain a fusion feature at least according to the enhanced feature of the first image and the enhanced feature of the second image; A matching module, configured to determine the image in the image library that matches the fusion feature as the result of the pedestrian re-identification.

8. A pedestrian re-identification device, characterized in that, The device includes: A memory, configured to store program instructions; A processor, configured to call the program instructions stored in the memory and execute the steps included in the method according to any one of claims 1-6 according to the obtained program instructions.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a computer, the method according to any one of claims 1-6 is executed.

10. A computer program product, characterized in that, The computer program product includes computer program code. When the computer program code runs on a computer, any one of claims 1-6 is executed.