Cross-domain pedestrian re-identification method based on feature memory distance
By initializing and training shallow and deep networks on the pre-trained pedestrian data set, the problem of high cost of data acquisition and manual labeling in the prior art is solved, and flexible and efficient pedestrian re-identification is achieved.
Patent Information
- Application Number
- CN202510485534.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-01
AI Technical Summary
The existing pedestrian re-identification method requires collecting data from scratch and labeling, which is costly and highly dependent on manual labeling.
The original network is generated using the pre-trained pedestrian data set, the shallow and deep networks are initialized, and iterative training is performed on the label pedestrian data set, the network parameters are updated through the memory module and the distance difference module, and finally the pedestrian re-identification is performed on the target pedestrian data set.
Reduces data costs and manual labeling dependencies, and achieves flexible and rapid pedestrian re-identification deployment.
Smart Images

Figure CN120412017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more particularly, to a cross-domain pedestrian re-identification method based on feature memory distance. Background Art
[0002] The pedestrian re-identification technology refers to matching pedestrians with the same identity in pedestrian pictures taken from different cameras and environments, which has great significance in modern security systems. Its typical applications include tracking escaped suspects in a large amount of surveillance data, looking for lost children, and so on.
[0003] Existing pedestrian re-identification methods need to collect data from scratch and label tags, which consume a large cost in practical applications, and the data collection also has a great dependence on accurate manual annotation. Summary of the Invention
[0004] In order to overcome the defects of high data cost and strong dependence on manual annotation in the above-mentioned prior art, the present invention provides a cross-domain pedestrian re-identification method based on feature memory distance.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] In a first aspect, a cross-domain pedestrian re-identification method based on feature memory distance includes:
[0007] Generating an original network on a pre-trained pedestrian dataset, and initializing an independent shallow network and a deep network with the network parameters of the original network;
[0008] Letting the shallow network and the deep network perform iterative training on a labeled pedestrian dataset, and respectively determining classification probability estimation information and relative distance deviation information of the output shallow features and deep features through a memorization module and a distance difference module, and synchronously updating network parameters;
[0009] Using a target pedestrian dataset as the input of the trained shallow network and deep network, determining a feature distance, and determining a pedestrian re-identification result based on the feature distance.
[0010] In a second aspect, a computer program product includes a computer program or computer-executable instructions, and when the computer program or computer-executable instructions are executed by a processor, the method described in the first aspect is implemented.
[0011] Compared with the prior art, the beneficial effect of the technical solution of the present invention is:
[0012] The present application discloses a cross-domain pedestrian re-identification method based on feature memory distance. The shallow network and the deep network are initialized by using the original network generated from a pre-trained pedestrian dataset, and then the shallow network and the deep network are trained on a labeled pedestrian dataset with improved corresponding labels, so that the shallow network and the deep network have targeted feature extraction capabilities, and finally pedestrian re-identification on a newly collected target pedestrian dataset is realized. Compared with the prior art, the present invention can effectively reduce the data cost and the dependence on manual annotation, has looser requirements for newly collected pedestrian data, and can realize a more flexible and rapid deployment of pedestrian re-identification. Description of the Drawings
[0013] Figure 1 It is a schematic flowchart of the cross-domain pedestrian re-identification method described in Embodiment 1 of the present application.
[0014] Figure 2 It is another schematic flowchart of the cross-domain pedestrian re-identification method described in Embodiment 1 of the present application. Detailed Embodiments
[0015] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices. The term "determine" broadly covers a variety of actions, which may include obtaining, calculating, computing, processing, deriving, researching, finding (e.g., looking up in a table, database or other data structure), ascertaining, and similar actions, may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and similar actions, may also include generating, creating, establishing and similar actions, as well as parsing, selecting, picking and similar actions, and so on. The relevant definitions of other terms will be given in the following description.
[0016] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediate element. In addition, in the following embodiments, "connection", if there is a transmission of electrical signals or data between the connected objects, should be understood as "electrical connection", "communication connection", etc.
[0017] The drawings are only for illustrative purposes and should not be construed as a limitation of the patent.
[0018] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which does not represent the size of the actual product;
[0019] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0020] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0021] Embodiment 1
[0022] This embodiment provides a cross-domain pedestrian re-identification method based on feature memory distance. Refer to Figure 1 , including:
[0023] S1. Generate an original network on a pre-trained pedestrian dataset, and initialize an independent shallow network and a deep network with the network parameters of the original network;
[0024] S2. Let the shallow network and the deep network perform iterative training on the labeled pedestrian dataset, and determine the classification probability estimation information and the relative distance deviation information of the output shallow features and deep features through a memorization module and a distance difference module respectively, and synchronously update the network parameters;
[0025] S3. Use the target pedestrian dataset as the input of the trained shallow network and deep network, determine the feature distance, and determine the pedestrian re-identification result based on the feature distance.
[0026] It should be noted that the labeled pedestrian dataset is a small amount of pedestrian data with complete annotation information (labels) for pedestrians, and the target pedestrian dataset is a rough (without annotation information or with only a small amount of annotation information) pedestrian dataset for target persons. In particular, the target pedestrian dataset can be a dataset associated with the "same" target person as the labeled pedestrian dataset, or a dataset associated with "different" persons from the labeled pedestrian dataset.
[0027] In some specific implementation processes, the target person associated with the target pedestrian dataset is different from the pedestrian associated with the labeled pedestrian dataset, which is usually more difficult but more technically valuable.
[0028] In this embodiment, the shallow network and the deep network are initialized with the original network generated from the publicly available pre-trained pedestrian dataset, and then the shallow network and the deep network are trained on the labeled pedestrian dataset with corresponding labels that have been improved, so that the shallow network and the deep network have targeted feature extraction capabilities, and finally pedestrian re-identification on the newly collected target pedestrian dataset is achieved, which can effectively reduce the data cost and the dependence on manual annotation, has looser requirements for the newly collected pedestrian data, and can achieve a more flexible and rapid deployment of pedestrian re-identification.
[0029] In some specific implementation processes, the target pedestrian dataset is used as the input of the trained shallow network and deep network for feature extraction, shallow features and deep features of the target pedestrian dataset are obtained, and the Euclidean distance is used to measure the shallow features and the deep features to determine the feature distance.
[0030] In some preferred embodiments, generating the original network on the pre-trained pedestrian dataset and initializing the independent shallow network and deep network with the network parameters of the original network includes:
[0031] Using the pre-trained pedestrian dataset as the input of the information extraction network N constructed based on the convolutional neural network to determine the extracted feature A, and using the extracted feature A to update the parameters of the information extraction network N through backpropagation;
[0032] Using the updated information extraction network as the original network, and using the network parameters of the original network as the initial values of the network parameters of the shallow network M and the deep network L.
[0033] In some specific implementation processes, the pre-trained pedestrian dataset is constructed based on the ImageNet dataset.
[0034] In some preferred embodiments, the memorization module includes a front part of the memorization module, a middle part of the memorization module, and a rear part of the memorization module. Among them, the front part of the memorization module is a memory feature projection layer with learnable parameters for mapping the memorization feature C, the middle part of the memorization module is used to determine the classification probability estimation information D, and the rear part of the memorization module is used to determine the classification memory feature Y; and,
[0035] The distance difference module includes a front part of the distance difference module, a middle part of the distance difference module, and a rear part of the distance difference module; among them, the front part of the distance difference module includes a distance estimation layer and a distance feature projection layer, the middle part of the distance difference module is used to determine the standard class distance H, and the rear part of the distance difference module is used to determine the relative distance deviation information J.
[0036] In some specific implementation processes, the memory feature projection layer uses a fully connected layer from 1024 channels to 1024 channels.
[0037] In some specific implementation processes, the distance estimation layer is a fully connected layer from 1024 channels to 1024 channels.
[0038] In some other specific implementation processes, the distance feature projection layer is a channel convolution layer from 2048 channels to 1024 channels.
[0039] In some alternative embodiments, referring to Figure 2 , the iterative training process of the shallow network includes:
[0040] Using the labeled pedestrian dataset as the input of the shallow network M to obtain the corresponding shallow features B;
[0041] Using the front part of the memorization module to determine the memorization features C based on the shallow features B;
[0042] Obtaining the memory clustering result X determined in the previous iteration process; or, when there is no memory clustering result X determined in the previous iteration process, using a clustering method to determine the memory clustering result X according to the shallow features B;
[0043] Obtaining the classification memory feature Y determined in the previous iteration process; or, when there is no classification memory feature Y determined in the previous iteration process, using a preset value as the classification memory feature Y;
[0044] Based on the memorization features C, the memory clustering result X, and the classification memory feature Y, determining the classification probability estimation information D through the middle part of the memorization module;
[0045] After normalizing and balancing the classification probability estimation information D with respect to the memorization features C, updating the classification memory feature Y of the current iteration through the rear part of the memorization module.
[0046] In some specific implementation processes, when there is no classification memory feature Y determined in the previous iteration process, the classification memory feature Y adopts all 0 values.
[0047] Furthermore, the expression of the classification probability estimation information D is:
[0048]
[0049] In the formula, represents calculating the similarity score between C and Y, and X(C) represents taking the category of C in the memory clustering result X to form the similarity score.
[0050] In some specific implementation processes, first calculate the vector inner product of the memorized feature C and the classified memory feature Y, then divide it by the modulus values of the two features, and subsequently take the ratio of the corresponding categories in the clustering result to form the similarity score.
[0051] Furthermore, the update process of the classified memory feature Y is as follows:
[0052] Y = (1 - λ)Y + λ[D(C)*C]
[0053] In the formula, λ represents the update amplitude; D(C) represents taking the category of C in the classification probability estimation information D; D(C)*C represents multiplying the category of C in the classification probability estimation information D and the item of C and retaining the result.
[0054] In some specific implementation processes, the update amplitude λ is set to 0.1.
[0055] In some alternative embodiments, the iterative training process of the deep network includes:
[0056] Using the labeled pedestrian dataset as the input of the deep network L to obtain the corresponding deep feature E;
[0057] Using the distance estimation layer to determine the relative distance F according to the deep feature;
[0058] Using the relative distance F and the memorized feature C as the input of the distance feature projection layer to determine the distance-transformed feature G;
[0059] Using the middle part of the distance difference module to determine the standard class distance H corresponding to the distance-transformed feature G according to the distance-transformed feature G and the memory clustering result X determined in the previous iteration process;
[0060] Using the rear part of the distance difference module to determine the relative distance deviation information J corresponding to the distance-transformed feature G according to the standard class distance H and the memory clustering result X.
[0061] In the above alternative embodiments, by adjusting the classification probability estimation information of the memorized feature and by calculating the relative distance deviation of the memorized feature to reduce the distance deviation of correct classification, the accuracy of re-identification is improved.
[0062] Furthermore, the expression of the standard class distance H is:
[0063]
[0064] In the formula, X(G) represents the class in the memory clustering result X that is the same as the distance-transformed feature G; represents taking the modulus operation; Represent other features that are of the same type but not exactly the same in the distance-based feature G.
[0065] In some specific implementation processes, during the calculation of the standard-class distance H at the rear of the distance difference module, other features that are of the same type but not exactly the same as the distance-based feature G Are 3 features sampled additionally in the same cluster.
[0066] Furthermore, the expression of the relative distance deviation information J is:
[0067]
[0068] In the formula, X(G) represents the class in the memory clustering result X that is the same as the distance-based feature G.
[0069] In some alternative embodiments, the updating of the network parameters includes:
[0070] Based on the classification probability estimation information D, the relative distance deviation information J, and the memory clustering result X, determine the error K, and its expression is:
[0071] K = μ(X - D)+(1 - μ)J
[0072] In the formula, μ represents the error ratio;
[0073] With the goal of minimizing the error K, based on the gradient descent method, update the network parameters of the shallow network M and the deep network L until the training ends.
[0074] It should be noted that when the error tends to converge or reaches the preset number of iteration rounds, the training can be ended, which is specifically set by those skilled in the art according to the actual situation.
[0075] In some specific implementation processes, when the relative change of the error K in two consecutive iterations is less than 1%, the training ends.
[0076] Embodiment 2
[0077] This embodiment provides a cross-domain pedestrian re-identification system based on feature memory distance, including:
[0078] An initialization module, used to generate an original network on a pre-trained pedestrian dataset, and initialize an independent shallow network and deep network with the network parameters of the original network;
[0079] A training module, used to make the shallow network and the deep network perform iterative training on a labeled pedestrian dataset, and respectively determine the classification probability estimation information and the relative distance deviation information through a memorization module and a distance difference module for the output shallow features and deep features, and synchronously update the network parameters;
[0080] A person re-identification module is used to load the trained shallow network and the deep network, determine a feature distance with a target person dataset as the input, and determine a person re-identification result based on the feature distance.
[0081] It can be understood that the device in this embodiment corresponds to the method in Embodiment 1 above. The optional items in Embodiment 1 above also apply to this embodiment, so they will not be repeated here.
[0082] Embodiment 3
[0083] This embodiment provides a computer-readable storage medium. At least one instruction, at least one program, a code set, or an instruction set is stored on the storage medium. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor, so that the processor executes some or all of the steps of the method provided in Embodiment 1 of this application.
[0084] It can be understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes but is not limited to various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0085] Exemplarily, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0086] Exemplarily, the read-only memory includes but is not limited to MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0087] Exemplarily, the random access memory includes but is not limited to DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0088] In some examples, a computer program product is provided, which can be specifically implemented in a way of hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied as the storage medium, and can also be embodied as a software product, such as an SDK (Software Development Kit, software development kit), etc.
[0089] As a non-limiting example, a computer program product is provided, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes some or all of the steps of the method described in the embodiments of the present application.
[0090] In some examples, a computer program is provided, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes to implement some or all of the steps of the method.
[0091] This embodiment also proposes an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, a code set or an instruction set. When the processor executes the at least one instruction, at least one program, the code set or the instruction set, some or all of the steps of the method described in Embodiment 1 are implemented.
[0092] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory and a communication interface; wherein, the processor generally controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or already processed by the processor and each module in the electronic device (including but not limited to image data, audio data, voice communication data and video communication data), and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or random access memory (RAM).
[0093] The processor may include one or more processing elements. Therefore, the processor may include one or more integrated circuits (ICs) configured to execute the functions of the processor. In addition, each integrated circuit may include circuits (such as a first circuit, a second circuit, and other circuits, etc.) configured to execute the functions of the processor.
[0094] Further, data transmission can be carried out among the processor, the communication interface, and the memory through a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together.
[0095] It can be understood that the optional items in the above-mentioned Embodiment 1 are equally applicable to this embodiment, so they will not be repeatedly described here.
[0096] The same or similar reference numerals correspond to the same or similar components;
[0097] The terms used to describe the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to this application;
[0098] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0099] In different specific implementations, the method or system described in this application can be implemented in software, hardware, or a combination thereof. In addition, the order of the steps of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0100] Obviously, the above embodiments of this application are merely examples given for clearly explaining this application, rather than limitations on the implementation manners of this application, and are not used to limit this application. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. Each separate structural / functional module or unit can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. The structures and functions of the discrete components can be implemented as a combined structure or component. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of this application shall be included within the protection scope of the claims of this application.
Claims
1. A cross-domain pedestrian re-identification method based on feature memory distance, characterized in that Including: Generate an original network on a pre-trained pedestrian dataset, and initialize an independent shallow network and a deep network with the network parameters of the original network; Let the shallow network and the deep network perform iterative training on the labeled pedestrian dataset, and respectively determine classification probability estimation information and relative distance deviation information for the output shallow features and deep features through a memorization module and a distance difference module, and synchronously update the network parameters; Use the target pedestrian dataset as the input of the trained shallow network and deep network, determine the feature distance, and determine the pedestrian re-identification result based on the feature distance.
2. The cross-domain pedestrian re-identification method based on feature memory distance according to claim 1, wherein The memorization module includes a front part of the memorization module, a middle part of the memorization module, and a rear part of the memorization module. Among them, the front part of the memorization module is a memory feature projection layer with learnable parameters for mapping the memorization feature C, the middle part of the memorization module is used to determine the classification probability estimation information D, and the rear part of the memorization module is used to determine the classification memory feature Y; and, The distance difference module includes a front part of the distance difference module, a middle part of the distance difference module, and a rear part of the distance difference module; among them, the front part of the distance difference module includes a distance estimation layer and a distance feature projection layer, the middle part of the distance difference module is used to determine the standard class distance H, and the rear part of the distance difference module is used to determine the relative distance deviation information J.
3. A cross-domain pedestrian re-identification method based on feature memory distance according to claim 2, characterized in that The iterative training process of the shallow network includes: Use the labeled pedestrian dataset as the input of the shallow network M to obtain the corresponding shallow feature B; Use the front part of the memorization module to determine the memorization feature C based on the shallow feature B; Obtain the memory clustering result X determined in the previous iteration process; or, when there is no memory clustering result X determined in the previous iteration process, use a clustering method to determine the memory clustering result X according to the shallow feature B; Obtain the classification memory feature Y determined in the previous iteration process; or, when there is no classification memory feature Y determined in the previous iteration process, use a preset value as the classification memory feature Y; Based on the memorization feature C, the memory clustering result X, and the classification memory feature Y, determine the classification probability estimation information D through the middle part of the memorization module; After performing normalization and balancing processing on the classification probability estimation information D for the memorization feature C, update the classification memory feature Y of the current iteration through the rear part of the memorization module.
4. A cross-domain pedestrian re-identification method based on feature memory distance according to claim 3, characterized in that The expression of the classification probability estimation information D is: In the formula, represents calculating the similarity score between C and Y, and X(C) represents the similarity score formed by taking the category of C in the memory clustering result X.
5. The cross-domain pedestrian re-identification method based on feature memory distance according to claim 3, wherein The update process of the classification memory feature Y is as follows: Y = (1 - λ)Y + λ[D(C)*C] In the formula, λ represents the update amplitude; D(C) represents taking the category of C in the classification probability estimation information D; D(C)*C represents taking the product of the category of C in the classification probability estimation information D and the item of C and retaining it.
6. A cross-domain pedestrian re-identification method based on feature memory distance according to claim 2, characterized in that The iterative training process of the deep network includes: Use the labeled pedestrian dataset as the input of the deep network L to obtain the corresponding deep feature E; Use the distance estimation layer to determine the relative distance F according to the deep feature; Use the relative distance F and the memorization feature C as the input of the distance feature projection layer to determine the distance feature G; Using the middle part of the distance difference module, determine the standard class distance H corresponding to the distance feature G according to the distance feature G and the memory clustering result X determined in the previous iteration process; Using the rear part of the distance difference module, determine the relative distance deviation information J corresponding to the distance feature G according to the standard class distance H and the memory clustering result X.
7. A cross-domain pedestrian re-identification method based on feature memory distance according to claim 6, characterized in that The expression of the standard class distance H is: Wherein, X(G) represents the class in the memory clustering result X that is the same as the distance-transformed feature G; represents taking the modulus operation; represents other features that are of the same class but not identical in the distance-transformed feature G.
8. A cross-domain pedestrian re-identification method based on feature memory distance according to claim 6, characterized in that The expression of the relative distance deviation information J is: In the formula, X(G) represents the class in the memory clustering result X that is the same as the distance feature G.
9. A cross-domain pedestrian re-identification method based on feature memory distance according to claim 2, wherein Updating the network parameters includes: Determine the error K according to the classification probability estimation information D, the relative distance deviation information J, and the memory clustering result X, and its expression is: K = μ(X - D)+(1 - μ)J In the formula, μ represents the error ratio; With the goal of minimizing the error K, based on the gradient descent method, update the network parameters of the shallow network M and the deep network L until the training ends.
10. A cross-domain pedestrian re-identification method based on feature memory distance according to any one of claims 1-9, characterized in that Generating the original network on the pre-trained pedestrian dataset and initializing the independent shallow network and deep network with the network parameters of the original network includes: Using the pre-trained pedestrian dataset as the input of the information extraction network N constructed based on the convolutional neural network, determine the extracted feature A, and use the extracted feature A to update the backpropagation parameters of the information extraction network N; Using the updated information extraction network as the original network, and using the network parameters of the original network as the initial values of the network parameters of the shallow network M and the deep network L.