A cross-modal pedestrian re-identification method, device and equipment and storage medium

By constructing a heterogeneous target graph and discarding redundant features, the problem of low retrieval accuracy in cross-modal pedestrian re-identification was solved, enabling more accurate intelligent tracking and capture, and promoting the development of intelligent security.

CN114708611BActive Publication Date: 2026-04-14SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2022-03-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing cross-modal pedestrian re-identification methods suffer from insufficient retrieval accuracy because the modal differences between near-infrared and visible light images lead to deep learning networks learning features containing modality-specific, identity-irrelevant information.

Method used

By acquiring a set of pedestrian images to be identified, performing feature extraction, constructing a target heterogeneous map, determining target features, and generating target pedestrian images, redundant features such as clothing color are discarded, thus achieving cross-modal pedestrian re-identification.

Benefits of technology

It has improved the accuracy of cross-modal pedestrian re-identification, enabled accurate nighttime intelligent tracking and personnel pursuit, and promoted the construction of intelligent security systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708611B_ABST
    Figure CN114708611B_ABST
Patent Text Reader

Abstract

The application discloses a cross-modal pedestrian re-identification method, device and equipment and a storage medium. The method comprises the following steps: acquiring a to-be-identified pedestrian image set, wherein the to-be-identified pedestrian image set comprises to-be-identified pedestrian images in at least two modes; performing feature extraction on the to-be-identified pedestrian images in the to-be-identified pedestrian image set to obtain feature information corresponding to each to-be-identified pedestrian image in the to-be-identified pedestrian image set; constructing a target heterogeneous graph according to each to-be-identified pedestrian image in the to-be-identified pedestrian image set and the feature information corresponding to each to-be-identified pedestrian image in the to-be-identified pedestrian image set; determining a target feature according to the target heterogeneous graph; and generating a target pedestrian image according to the target feature. The scheme can be added to the back of any existing cross-modal pedestrian re-identification method as post-processing, improves the accuracy of the existing cross-modal pedestrian re-identification method, makes the intelligent tracking and personnel pursuit at night more accurate, and promotes the intelligent security construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pedestrian re-identification technology, and in particular to a cross-modal pedestrian re-identification method, apparatus, device, and storage medium. Background Technology

[0002] Cross-modal pedestrian re-identification plays a significant role in smart city security monitoring scenarios. The cross-modal pedestrian re-identification method enables mutual retrieval between human images captured in visible light and those captured in near-infrared imaging. The goal is to find an image matching the queried image given a visible light human image, and vice versa, to find an image matching the queried image given a near-infrared image.

[0003] Most existing cross-modal person re-identification methods use labels as a medium and employ statistical learning methods (currently represented by neural networks) to enable network models to extract effective information representing the identity from visible light and near-infrared pedestrian images. However, because near-infrared and visible light images have significant modal differences, the features learned by deep learning networks still contain modality-specific, identity-irrelevant information (such as color information in visible light images), resulting in insufficient accuracy in cross-modal retrieval. Summary of the Invention

[0004] This invention provides a cross-modal pedestrian re-identification method, apparatus, device, and storage medium to solve the problem of insufficient retrieval accuracy in cross-modal pedestrian re-identification in the prior art, achieving more accurate intelligent tracking and personnel pursuit at night, and greatly promoting the construction of intelligent security.

[0005] According to one aspect of the present invention, a cross-modal pedestrian re-identification method is provided, the method comprising:

[0006] Obtain a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images in at least two modalities;

[0007] Feature extraction is performed on the pedestrian images to be identified in the set of pedestrian images to be identified to obtain the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0008] A target heterogeneous graph is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified;

[0009] Determine target features based on the target heterogeneity graph;

[0010] Generate a target pedestrian image based on the target features.

[0011] According to another aspect of the present invention, a cross-modal pedestrian re-identification device is provided, the device comprising:

[0012] The acquisition module is used to acquire a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images to be identified in at least two modalities;

[0013] The feature extraction module is used to extract features from the pedestrian images to be identified in the set of pedestrian images to be identified, and obtain the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0014] The construction module is used to construct a target heterogeneous graph based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0015] The determination module is used to determine target features based on the target heterogeneity map;

[0016] The generation module is used to generate a target pedestrian image based on the target features.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the cross-modal pedestrian re-identification method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the cross-modal pedestrian re-identification method according to any embodiment of the present invention.

[0022] The technical solution of this invention involves acquiring a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images in at least two modalities; extracting features from the pedestrian images to be identified in the set of pedestrian images to obtain feature information corresponding to each pedestrian image in the set of pedestrian images to be identified; constructing a target heterogeneous map based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified; determining target features based on the target heterogeneous map; and generating target pedestrian images based on the target features. This invention can be added to any existing cross-modal pedestrian re-identification method as post-processing to improve the accuracy of existing cross-modal pedestrian re-identification methods, making intelligent tracking and personnel pursuit more accurate at night, and greatly promoting the construction of intelligent security.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a cross-modal pedestrian re-identification method provided in Embodiment 1 of the present invention;

[0026] Figure 2a This is a schematic diagram of a first heterogeneous graph provided according to Embodiment 1 of the present invention;

[0027] Figure 2b This is a schematic diagram of a method for noise removal from a first heterogeneous graph according to Embodiment 1 of the present invention;

[0028] Figure 3 This is a schematic diagram of a cross-modal pedestrian re-identification method according to Embodiment 1 of the present invention;

[0029] Figure 4 This is a schematic diagram of the structure of a cross-modal pedestrian re-identification device according to Embodiment 2 of the present invention;

[0030] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the cross-modal pedestrian re-identification method of this invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first," "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] Example 1

[0034] Figure 1 This is a flowchart of a cross-modal pedestrian re-identification method according to Embodiment 1 of the present invention. This embodiment is applicable to cross-modal pedestrian re-identification scenarios. The method can be executed by a cross-modal pedestrian re-identification device, which can be implemented in hardware and / or software. This cross-modal pedestrian re-identification device can be integrated into any electronic device that provides cross-modal pedestrian re-identification functionality. Figure 1 As shown, the method includes:

[0035] S101. Obtain the set of pedestrian images to be identified.

[0036] It should be noted that the set of pedestrian images to be identified refers to the collection of existing pedestrian images to be identified. For example, pedestrian images can be images of pedestrians captured by surveillance cameras.

[0037] The set of pedestrian images to be identified includes pedestrian images in at least two modalities.

[0038] It should be noted that, in this embodiment, the pedestrian images to be identified in at least two modes can be understood as pedestrian images to be identified taken under different conditions. For example, they can be visible light pedestrian images to be identified taken during the day, or near-infrared pedestrian images to be identified taken at night.

[0039] Specifically, by collecting surveillance images and other methods, at least two existing pedestrian images in different modalities are obtained to form a set of pedestrian images to be identified.

[0040] S102. Extract features from the pedestrian images to be identified in the set of pedestrian images to be identified, and obtain the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified.

[0041] It should be noted that feature extraction can be understood as extracting information from an image and determining whether each point in the image belongs to an image feature. For example, it could be extracting facial information of pedestrians in an image, or it could be extracting clothing information of pedestrians in an image.

[0042] In this embodiment, the feature information may be, for example, the facial features, hairstyle features, or clothing features of pedestrians in the image.

[0043] For example, a convolutional neural network can be used for feature extraction. Specifically, the images of pedestrians to be identified in the acquired set of images to be identified can be input into the convolutional neural network to obtain the feature information corresponding to each image of pedestrians to be identified in the set of images to be identified.

[0044] S103. Construct a target heterogeneous graph based on the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified.

[0045] It should be explained that the target heterogeneous graph can be understood as a model of the relationship between different entities constructed from the feature information of each pedestrian image in the set of pedestrian images to be identified and the corresponding pedestrian images in the set of pedestrian images to be identified. It is modeled in a graph manner to analyze the topological relationship between different entities.

[0046] Specifically, a target heterogeneous graph is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified.

[0047] Because the constructed graph network contains the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified in different modalities, this graph network is a heterogeneous graph. Heterogeneous graph methods are used to model the relationships between different types of entities. Modeling is done through a graph, where different types of entities are represented as nodes in the heterogeneous graph, and the edges connecting nodes represent the interaction relationships between different entities. Through learning, the features of nodes or edges in the graph can represent the topological relationships between entities, or information can be transferred between nodes based on the graph's topological relationships.

[0048] S104. Determine target features based on the target heterogeneity graph.

[0049] It should be noted that the target feature can be understood as the specific pedestrian features that can represent the semantics of pedestrian identity in the process of pedestrian re-identification, such as the pedestrian's facial features.

[0050] Specifically, target features are determined based on the target heterogeneous map, that is, specific pedestrian features (such as pedestrian facial features) that can represent the semantics of pedestrian identity during the pedestrian re-identification process are determined, and redundant features (such as pedestrian clothing color) in the pedestrian image to be identified are discarded.

[0051] S105. Generate a target pedestrian image based on the target features.

[0052] It should be noted that the target pedestrian image refers to the image of the target person generated during the pedestrian re-identification process based on the target features.

[0053] Specifically, a target pedestrian image is generated based on the target features, and a search and matching process is performed in the image library based on the target pedestrian image to find an image with the same identity semantics as the target pedestrian image, thereby achieving cross-modal pedestrian re-identification.

[0054] The technical solution of this invention involves acquiring a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images in at least two modalities; extracting features from the pedestrian images to be identified in the set of pedestrian images to obtain feature information corresponding to each pedestrian image in the set of pedestrian images to be identified; constructing a target heterogeneous map based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified; determining target features based on the target heterogeneous map; and generating target pedestrian images based on the target features. This invention can be added to any existing cross-modal pedestrian re-identification method as post-processing to improve the accuracy of existing cross-modal pedestrian re-identification methods, making intelligent tracking and personnel pursuit more accurate at night, and greatly promoting the construction of intelligent security.

[0055] Optionally, target features are determined based on the target heterogeneity graph, including:

[0056] Based on the target heterogeneous graph, information is aggregated from at least two modalities of the pedestrian image to be identified, resulting in at least two modal aggregated features.

[0057] In this embodiment, information aggregation refers to aggregating the pedestrian images to be identified in the target heterogeneous map according to different modalities. For example, it could be to aggregate information from all visible light modal pedestrian images to be identified, or to aggregate information from all near-infrared modal pedestrian images to be identified.

[0058] It should be noted that modal aggregation features refer to the aggregation features obtained by aggregating the pedestrian images to be identified in the target heterogeneous map according to different modalities. For example, it can be that the information of the pedestrian images to be identified in all visible light modalities is aggregated to obtain visible light modal aggregation features, and the information of the pedestrian images to be identified in all near-infrared modalities is aggregated to obtain near-infrared modal aggregation features.

[0059] Specifically, the pedestrian images to be identified in at least two modalities in the target heterogeneous image are aggregated according to the modal category to obtain at least two modal aggregation features. For example, the pedestrian images to be identified in the visible light modality are aggregated to obtain visible light modal aggregation features, and the pedestrian images to be identified in the near-infrared modality are aggregated to obtain near-infrared modal aggregation features.

[0060] The target feature is obtained by fusing at least two modal aggregated features.

[0061] It should be noted that fusion refers to fusing at least two modal aggregated features obtained after information aggregation.

[0062] Specifically, the target features are obtained by fusing the aggregated features of at least two modalities of the pedestrian image to be identified obtained by aggregating information from at least two modalities based on the target heterogeneous map.

[0063] Optionally, information aggregation is performed on the pedestrian images to be identified in at least two modalities based on the target heterogeneous map to obtain at least two modal aggregated features, including:

[0064] Modal partitioning is performed on the nodes in the target heterogeneous graph to obtain a node set corresponding to at least two modes.

[0065] In this embodiment, a node refers to each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified.

[0066] It should be explained that modal partitioning refers to dividing the nodes in the target heterogeneous graph according to different modes. For example, the nodes in the target heterogeneous graph can be divided according to visible light mode and near-infrared mode.

[0067] It should be noted that the node set refers to the set of nodes with the same mode after the nodes in the target heterogeneous graph are divided according to different modes.

[0068] Specifically, nodes in the target heterogeneous graph are modally divided according to their different modes, resulting in node sets corresponding to at least two modes. For example, nodes in the target heterogeneous graph can be divided into visible light modes and near-infrared modes, resulting in node sets corresponding to visible light modes and node sets corresponding to near-infrared modes.

[0069] Information is aggregated from the nodes in the node sets corresponding to at least two modalities to obtain aggregated features for at least two modalities.

[0070] Specifically, information aggregation is performed on the nodes in the node sets corresponding to at least two modes obtained after modal segmentation to obtain at least two modal aggregation features. For example, information aggregation can be performed on the nodes in the node set corresponding to the visible light mode to obtain visible light modal aggregation features, and information aggregation can be performed on the nodes in the node set corresponding to the near-infrared mode to obtain near-infrared modal aggregation features.

[0071] In practical operation, all nodes in the target heterogeneous graph can be taken as the central node. The nearest neighbors of the central node contain nodes of different modalities. Generally speaking, due to the differences in the original feature distribution of the pedestrian image to be identified in different modalities, most of the nearest neighbors are of the same modality as the central node. This embodiment of the invention proposes to divide the nearest neighbors according to modality and perform feature aggregation to obtain corresponding modal aggregated features. For example, the aggregation of nearest neighbors in visible light yields visible light modal aggregated features, and the aggregation of nearest neighbors in near-infrared light yields near-infrared modal aggregated features. Modal aggregated features The specific calculation formula is as follows:

[0072]

[0073] in, The calculated modal aggregation features are given by σ, which is a nonlinear activation function, and x. i N represents the nearest neighbor node. α (x k ) represents x k The set of nearest neighbors of the central node. u represents the calculated attention weights. i x represents i The corresponding original features.

[0074] Among them, the calculated attention weights The calculation formula is as follows:

[0075]

[0076] Among them, a α (x k ,x i ) represents the degree of association between different nodes, a α (x k ,x i The specific formula for calculating ) is as follows:

[0077]

[0078] Among them, A α and W α This represents the learnable parameters, which can be trained under supervised conditions using classification loss functions and triplet loss functions, thereby improving the learnable parameters A. α and W α To learn.

[0079] Optionally, at least two modality aggregation features are fused to obtain the target features, including:

[0080] At least two modal aggregated features are mapped to at least two subspaces respectively to obtain the aggregated features corresponding to each subspace.

[0081] In this embodiment, the space containing at least two modal aggregation features can be regarded as the full space, and the subspace can be understood as a part of the space with a dimension smaller than the full space.

[0082] Specifically, at least two modal aggregated features can be mapped to at least two subspaces using linear mapping, resulting in aggregated features for each subspace. In practice, the mapping process is equivalent to processing the input at least two modal aggregated features to obtain processed aggregated features for each subspace.

[0083] The target feature is obtained by linearly mapping the aggregated features corresponding to at least two subspaces.

[0084] As we know, linear mapping refers to a process in which the input is a feature vector of length K, and the output is multiple feature vectors of length M by multiplying them with multiple matrices, where K and M can be set as needed.

[0085] Specifically, at least two modal aggregated features are mapped to at least two subspaces respectively, and the aggregated features corresponding to each subspace are obtained. Then, the aggregated features corresponding to the at least two subspaces are mapped again through linear mapping to obtain Q. k K k and Z k Three characteristics, Q k and K k The attention map is obtained by performing the inner product, and then compared with Z. k The target features are obtained by multiplication, and this process can refer to the existing technology called transformer.

[0086] Optionally, a target heterogeneous map is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified, including:

[0087] A first heterogeneous graph is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified.

[0088] It should be noted that the first heterogeneous map refers to the original noisy heterogeneous map constructed based on the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified.

[0089] Specifically, the first heterogeneous graph is constructed by using the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified as nodes.

[0090] The first heterogeneous graph is subjected to noise removal to obtain the target heterogeneous graph.

[0091] Specifically, noise removal is performed on the first heterogeneous image, that is, filtering out the pedestrian images in the set of pedestrian images to be identified whose identity semantic information is different from that of the target pedestrian in the first heterogeneous image, and the feature information corresponding to the pedestrian images in the set of pedestrian images to be identified, to obtain the target heterogeneous image.

[0092] Optionally, noise removal is performed on the first heterogeneous graph to obtain the target heterogeneous graph, including:

[0093] Obtain the confidence score of each node in the first heterogeneous graph.

[0094] In this embodiment, the confidence level of each node can be assigned manually based on the actual situation and can be adjusted according to performance. The confidence level of each node can be in exponential form, such as λ, where 0 < λ < 1.

[0095] Specifically, each pedestrian image in the set of pedestrian images to be identified and the corresponding feature information of each pedestrian image in the set of pedestrian images to be identified are taken as nodes. Each node is taken as the center node, and the nodes connected to it are searched in a recursive manner. A certain confidence level is assigned to the connected nodes to obtain the confidence level of each node in the first heterogeneous graph.

[0096] For example, Figure 2a This is a schematic diagram of a first heterogeneous graph provided according to Embodiment 1 of the present invention. Figure 2a As shown, there are a total of 9 pedestrian images in this set of images to be identified. After feature extraction of the 9 pedestrian images, the feature information corresponding to the 9 pedestrian images to be identified is obtained. The 9 pedestrian images to be identified and the feature information corresponding to the 9 pedestrian images to be identified are used as nodes to construct the first heterogeneous graph. The 9 nodes are labeled as nodes 1, 2, 3, 4, 5, 6, 7, 8 and 9.

[0097] Figure 2b This is a schematic diagram of a method for noise removal from a first heterogeneous graph according to Embodiment 1 of the present invention. In actual operation, for example, node 1 can be used as the central node to search for directly connected first-order nodes (nodes 2, 3, 4, 6, 7). Nodes in this stage will be assigned a certain confidence level λ. 1 , where 0 < λ 1 <1. Next, each node in the first phase continues to search for nodes directly connected to it (for example, using node 2 as the central node, it searches for nodes 3, 8, and 9 directly connected to it; using node 3 as the central node, it searches for nodes 2, 4, and 8 directly connected to it; using node 4 as the central node, it searches for nodes 3, 5, and 6 directly connected to it; using node 6 as the central node, it searches for nodes 4 and 5 directly connected to it; using node 7 as the central node, it finds no directly connected nodes). Nodes in this phase are also assigned a certain confidence level λ. 2 This process continues for N rounds. It's worth noting that at each stage of the search, the same node can be searched multiple times. Ultimately, the total confidence of each node is the sum of all the confidence scores assigned to it. Based on an observation, if a node truly shares the same identity as the central node, it will be searched in multiple search stages during the search process.

[0098] In actual operation, when constructing a heterogeneous graph, due to the low robustness of the initial features, if there are incorrect connections between nodes, it will lead to errors in subsequent steps. In this embodiment of the invention, the nearest neighbor search method is used to automatically filter the nearest neighbor nodes to enhance robustness.

[0099] Nodes with confidence scores below the confidence threshold in the first heterogeneous graph are deleted to obtain the target heterogeneous graph.

[0100] The confidence threshold can be a pre-set value of the node's confidence level based on the actual situation, and can be adjusted based on experience. This embodiment does not limit this.

[0101] Specifically, after obtaining the confidence level of each node in the first heterogeneous graph, nodes with confidence levels lower than the confidence threshold in the first heterogeneous graph are deleted to obtain the target heterogeneous graph.

[0102] In practice, you can also select the K nodes with the highest total confidence (K can be an appropriate integer or adjusted according to the specific effect) to keep, and filter out other noisy nodes with low confidence.

[0103] Optionally, a first heterogeneous graph is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified, including:

[0104] Each pedestrian image in the set of pedestrian images to be identified and the corresponding feature information of each pedestrian image in the set of pedestrian images to be identified are used as nodes.

[0105] Specifically, after obtaining the set of pedestrian images to be identified, each pedestrian image in the set of pedestrian images to be identified and the corresponding feature information of each pedestrian image in the set of pedestrian images to be identified are used as nodes in the first heterogeneous graph.

[0106] Get the similarity between any two nodes.

[0107] It should be noted that the images of pedestrians to be identified in the acquired set of images to be identified are input into the convolutional neural network to obtain the feature information corresponding to each image of pedestrians to be identified in the set of images to be identified. Since the image data is transformed into features (represented by vectors) through the convolutional neural network, the similarity calculation between any two nodes is equivalent to calculating the cosine similarity between the two vectors.

[0108] Specifically, the similarity between any two nodes is calculated to obtain the similarity between any two nodes.

[0109] Connect two nodes whose similarity is greater than a similarity threshold using an edge.

[0110] The similarity threshold can be any value of similarity between two nodes that is preset according to the actual situation. It can be adjusted based on experience, and this embodiment does not limit it.

[0111] Specifically, calculate the similarity between any two nodes, and connect the nodes with a similarity greater than the similarity threshold by connecting them with an edge, thus constructing the first heterogeneous graph.

[0112] As an exemplary description of an embodiment of the present invention Figure 3 This is a schematic diagram of a cross-modal pedestrian re-identification method provided in Embodiment 1 of the present invention.

[0113] like Figure 3 As shown, the first step is to obtain a set of pedestrian images to be identified. Figure 3The set of pedestrian images to be identified consists of 7 images. Next, features are extracted from these 7 images to obtain feature information corresponding to each image. Then, a first heterogeneous graph is constructed based on the feature information of each image. Noise removal is performed on the first heterogeneous graph to obtain a target heterogeneous graph. Nodes in the target heterogeneous graph are modally segmented to obtain at least two sets of nodes corresponding to different modalities. Information aggregation is performed on the nodes in each of these sets to obtain at least two modal aggregation features. These features are then fused to obtain target features. Finally, a target pedestrian image is generated based on the target features, and cross-modal pedestrian re-identification is performed based on the target pedestrian image.

[0114] Table 1 compares the performance parameters of the cross-modal pedestrian re-identification method with denoising capability provided in the embodiments of the present invention with some existing pedestrian re-identification methods.

[0115] Table 1

[0116]

[0117] As shown in Table 1, the first six methods are all pedestrian re-identification methods that do not use heterogeneous graphs for post-processing enhancement. Compared with the method proposed in the embodiments of this invention, their cross-modal re-identification accuracy is lower. The seventh method is an existing general heterogeneous graph attention method, which does not optimize for the problem of modal distribution imbalance in cross-modal pedestrian re-identification tasks. Its accuracy is lower than that of the method proposed in the embodiments of this invention. Compared with the eighth method, the method proposed in the embodiments of this invention has the ability to remove noise, which can effectively avoid the adverse effects of erroneous connections in heterogeneous graphs.

[0118] Training with conventional data cannot accurately predict attention. This invention introduces a heterogeneous graph structure to model the structural relationships between different images in the dataset, expanding the scope to multiple images and generating more accurate attention. The features extracted from each image are post-processed to further enhance the features that represent identity semantics in each image and suppress redundant modal features (such as clothing color), thereby achieving more accurate cross-modal person re-identification.

[0119] Example 2

[0120] Figure 4 This is a schematic diagram of the structure of a cross-modal pedestrian re-identification device according to Embodiment 2 of the present invention. Figure 4As shown, the device includes: an acquisition module 201, a feature extraction module 202, a construction module 203, a determination module 204, and a generation module 205.

[0121] The acquisition module 201 is used to acquire a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images to be identified in at least two modalities;

[0122] Feature extraction module 202 is used to extract features from the pedestrian images to be identified in the set of pedestrian images to be identified, and obtain feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0123] Construction module 203 is used to construct a target heterogeneous map based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0124] Determining module 204 is used to determine target features based on the target heterogeneity map;

[0125] The generation module 205 is used to generate a target pedestrian image based on the target features.

[0126] Optionally, the determining module 204 includes:

[0127] An information aggregation unit is used to aggregate information from at least two modalities of the pedestrian image to be identified based on the target heterogeneous graph, and obtain at least two modal aggregation features.

[0128] A fusion unit is used to fuse the at least two modal aggregated features to obtain the target feature.

[0129] Optionally, the information aggregation unit is specifically used for:

[0130] Modal partitioning is performed on the nodes in the target heterogeneous graph to obtain a node set corresponding to at least two modes;

[0131] Information aggregation is performed on the nodes in the node sets corresponding to the at least two modalities to obtain at least two modal aggregation features.

[0132] Optionally, the fusion unit is specifically used for:

[0133] Map at least two modal aggregated features to at least two subspaces respectively to obtain the aggregated features corresponding to each subspace;

[0134] The target feature is obtained by linearly mapping the aggregated features corresponding to at least two subspaces.

[0135] Optionally, building module 203 includes:

[0136] The construction unit is used to construct a first heterogeneous graph based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0137] The noise removal unit is used to remove noise from the first heterogeneous map to obtain the target heterogeneous map.

[0138] Optionally, the noise removal unit is specifically used for:

[0139] Obtain the confidence level of each node in the first heterogeneous graph;

[0140] Nodes with confidence scores below the confidence threshold in the first heterogeneous graph are deleted to obtain the target heterogeneous graph.

[0141] Optionally, the building block is specifically used for:

[0142] Each pedestrian image in the set of pedestrian images to be identified and the corresponding feature information of each pedestrian image in the set of pedestrian images to be identified are used as nodes;

[0143] Get the similarity between any two nodes;

[0144] Connect two nodes whose similarity is greater than a similarity threshold using an edge.

[0145] The cross-modal pedestrian re-identification device provided in this embodiment of the invention can execute the cross-modal pedestrian re-identification method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0146] Example 3

[0147] Figure 5 A schematic diagram of an electronic device 30 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0148] like Figure 5As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32 or a random access memory (RAM) 33, communicatively connected to the at least one processor 31. The memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes based on the computer program stored in the ROM 32 or loaded from storage unit 38 into the RAM 33. The RAM 33 can also store various programs and data required for the operation of the electronic device 30. The processor 31, ROM 32, and RAM 33 are interconnected via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0149] Multiple components in electronic device 30 are connected to I / O interface 35, including: input unit 36, such as keyboard, mouse, etc.; output unit 37, such as various types of monitors, speakers, etc.; storage unit 38, such as disk, optical disk, etc.; and communication unit 39, such as network card, modem, wireless transceiver, etc. Communication unit 39 allows electronic device 30 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0150] Processor 31 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 31 performs the various methods and processes described above, such as cross-modal person re-identification methods:

[0151] Obtain a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images in at least two modalities;

[0152] Feature extraction is performed on the pedestrian images to be identified in the set of pedestrian images to be identified to obtain the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified;

[0153] A target heterogeneous graph is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified;

[0154] Determine target features based on the target heterogeneity graph;

[0155] Generate a target pedestrian image based on the target features.

[0156] In some embodiments, the cross-modal pedestrian re-identification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded into RAM 33 and executed by processor 31, one or more steps of the cross-modal pedestrian re-identification method described above may be performed. Alternatively, in other embodiments, processor 31 may be configured to perform the cross-modal pedestrian re-identification method by any other suitable means (e.g., by means of firmware).

[0157] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0159] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0161] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0162] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0163] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0164] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A cross-modal pedestrian re-identification method, characterized in that, include: Obtain a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images in at least two modalities; Feature extraction is performed on the pedestrian images to be identified in the set of pedestrian images to be identified to obtain the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified; A target heterogeneous graph is constructed based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified; the target heterogeneous graph is a model of the relationship between different entities constructed by each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified. Determine target features based on the target heterogeneity graph; A target pedestrian image is generated based on the target features. The target pedestrian image is then searched and matched in an image library to find an image with the same identity semantics as the target pedestrian image, thereby achieving cross-modal pedestrian re-identification. The step of determining the target features based on the target heterogeneity graph includes: Based on the target heterogeneous graph, information is aggregated from at least two modal images of pedestrians to be identified to obtain at least two modal aggregated features. The at least two modal aggregation features are fused to obtain the target feature.

2. The method according to claim 1, characterized in that, The step of aggregating information from at least two modalities of the pedestrian image to be identified based on the target heterogeneous graph to obtain at least two modal aggregated features includes: Modal partitioning is performed on the nodes in the target heterogeneous graph to obtain a node set corresponding to at least two modes; Information aggregation is performed on the nodes in the node sets corresponding to the at least two modalities to obtain at least two modal aggregation features.

3. The method according to claim 1, characterized in that, The process of fusing the at least two modal aggregation features to obtain the target feature includes: Map at least two modal aggregated features to at least two subspaces respectively to obtain the aggregated features corresponding to each subspace; The target feature is obtained by linearly mapping the aggregated features corresponding to at least two subspaces.

4. The method according to claim 1, characterized in that, The step of constructing a target heterogeneous map based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified includes: A first heterogeneous graph is constructed based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified; The first heterogeneous graph is subjected to noise removal to obtain the target heterogeneous graph.

5. The method according to claim 4, characterized in that, The step of removing noise from the first heterogeneous graph to obtain the target heterogeneous graph includes: Obtain the confidence level of each node in the first heterogeneous graph; Nodes with confidence scores below the confidence threshold in the first heterogeneous graph are deleted to obtain the target heterogeneous graph.

6. The method according to claim 4, characterized in that, The step of constructing a first heterogeneous graph based on each pedestrian image in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image in the set of pedestrian images to be identified includes: Each pedestrian image in the set of pedestrian images to be identified and the corresponding feature information of each pedestrian image in the set of pedestrian images to be identified are used as nodes; Get the similarity between any two nodes; Connect two nodes whose similarity is greater than a similarity threshold using an edge.

7. A cross-modal pedestrian re-identification device, characterized in that, include: The acquisition module is used to acquire a set of pedestrian images to be identified, wherein the set of pedestrian images to be identified includes pedestrian images to be identified in at least two modalities; The feature extraction module is used to extract features from the pedestrian images to be identified in the set of pedestrian images to be identified, and obtain the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified; A construction module is used to construct a target heterogeneous graph based on each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified; the target heterogeneous graph is a model of the relationship between different entities constructed by each pedestrian image to be identified in the set of pedestrian images to be identified and the feature information corresponding to each pedestrian image to be identified in the set of pedestrian images to be identified; The determination module is used to determine target features based on the target heterogeneity map; The generation module is used to generate a target pedestrian image based on the target features, and to search and match the target pedestrian image in the image library to find an image with the same identity semantics as the target pedestrian image, thereby realizing cross-modal pedestrian re-identification; The determining module includes: An information aggregation unit is used to aggregate information from at least two modalities of the pedestrian image to be identified based on the target heterogeneous graph, and obtain at least two modal aggregation features. A fusion unit is used to fuse the at least two modal aggregated features to obtain the target feature.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the cross-modal pedestrian re-identification method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the cross-modal pedestrian re-identification method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Heterogeneous face identification method based on deep convolutional neural network

    CN105608450A