Pedestrian re-identification method and device based on cross-camera prototype learning

By employing a cross-camera prototype learning method, pedestrian prototypes are learned and matched using in-camera and out-of-camera supervision information. By bringing the same prototypes closer together and pushing different prototypes further apart, the performance degradation of pedestrian re-identification models in independent camera scenarios is solved, and high-performance pedestrian re-identification is achieved.

CN117315712BActive Publication Date: 2026-06-02INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
Filing Date
2023-07-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In standalone camera scenarios, existing technologies ignore the prototype relationships between cameras, leading to a decline in the performance of pedestrian re-identification models due to factors such as camera background, lighting, and angle.

Method used

By employing a cross-camera prototype learning method, pedestrian prototypes are learned using labeled in-camera supervised information. The same pedestrian prototypes are matched across cameras, and the same prototypes are brought closer together while different prototypes are pushed further away. The model is optimized by introducing cross-camera prototype triplet loss and bringing-close loss.

Benefits of technology

The performance of the pedestrian re-identification model has been improved, eliminating the influence of factors such as camera background, lighting, and angle, and achieving high-performance pedestrian re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315712B_ABST
    Figure CN117315712B_ABST
Patent Text Reader

Abstract

The application provides a pedestrian re-identification method and device based on cross-camera prototype learning, and the method comprises the following steps: learning the pedestrian prototype of each camera by using the labeled camera internal supervision information, so as to perform the camera internal training of the pedestrian re-identification model; matching the same pedestrian prototypes between the cross cameras, and based on the matched pedestrian prototypes, narrowing the same pedestrian prototypes between different cameras and widening the different pedestrian prototypes, so as to perform the cross-camera training of the pedestrian re-identification model; and performing the prediction of the pedestrian re-identification based on the pedestrian re-identification model which has completed the camera internal training and the cross-camera training. The application can eliminate the influence of factors such as camera background, light, angle and the like on the pedestrian re-identification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent surveillance, and in particular relates to a pedestrian re-identification method and device based on cross-camera prototype learning. Background Technology

[0002] In recent years, with the continuous development of the field of intelligent surveillance, surveillance data has shown an explosive growth trend, and the massive amount of surveillance data has brought huge challenges to traditional manual surveillance analysis. As a result, intelligent surveillance systems based on computer vision technology have emerged. Among them, pedestrian re-identification technology, as a sub-problem in the field of image retrieval, aims to retrieve images of specific pedestrians from disjoint cameras. Existing supervised pedestrian re-identification work has achieved significant retrieval performance. However, in practical applications, the cost of data annotation increases dramatically with the increase in the number of cameras. This has led to a new scenario called Independent Camera Scene Pedestrian Re-identification (ICS). ICS assumes that pedestrian identities are labeled only within each camera, without cross-camera pedestrian identity association.

[0003] In pedestrian re-identification tasks within independent camera scenarios, two pioneering works, MTML and MATE, explored similar problems early on. They investigated from two perspectives: intra-camera training and inter-camera training. In the intra-camera training phase, they proposed a multi-camera, multi-task framework to learn discriminative features within each camera. In the inter-camera training phase, they proposed a multi-label alignment method to explore the identity relationships of pedestrians between cameras. However, these methods only explored the relationships between instances across cameras, neglecting the relationships between prototypes across cameras. This makes the distance between the same pedestrian prototypes in different cameras larger under the influence of factors such as camera background, lighting, and angle, inevitably leading to a decrease in the performance of the retrieval model. Summary of the Invention

[0004] To address the aforementioned issues, this invention discloses a pedestrian re-identification method and apparatus based on cross-camera prototype learning. This method can learn a high-performance retrieval model using only data labeled within an independent camera, even in scenarios lacking cross-camera annotations, thereby eliminating the influence of factors such as camera background, lighting, and angle on the pedestrian re-identification model.

[0005] The technical solution of the present invention includes:

[0006] A pedestrian re-identification method based on cross-camera prototype learning, the method comprising:

[0007] Using labeled in-camera supervision information, we learn pedestrian prototypes for each camera to train the pedestrian re-identification model in-camera.

[0008] Match the same pedestrian prototypes across cameras, and based on the matched pedestrian prototypes, bring the same pedestrian prototypes closer to different cameras and push the different pedestrian prototypes further away, so as to train the pedestrian re-identification model across cameras.

[0009] Pedestrian re-identification prediction is performed based on pedestrian re-identification models that have completed in-camera and cross-camera training.

[0010] Furthermore, the step of using labeled in-camera supervised information to learn the prototype features of pedestrians in each camera for in-camera training of the pedestrian re-identification model includes:

[0011] Calculate pedestrian image x i The global index j of the t-th pedestrian prototype within camera q. q ; where i, q, t, j q It is a natural number;

[0012] According to the pedestrian image x i The first feature f((x) within camera q i ) q and the global index j q Based on the corresponding pedestrian prototype features, calculate the cross-entropy loss L of the pedestrian re-identification model during in-camera training. intra_ID Wherein, the first feature f is the regularized feature;

[0013] For each batch of data, train the a-th image x of pedestrian t′. t′,a Calculate the image x t′,a The second feature h((x) within camera q t′,a ) q ), and combined with the feature h((x) t′,a ) q The second feature h((x) of the farthest positive sample t′,p ) q ) and the second feature h((x) of the nearest negative sample r,n ) q Construct the triplet loss L for mining difficult positive and negative instance-instance pairs. IT ;where x t′,p This indicates that the p-th image of the t′-th pedestrian is image x. t′,a The most difficult positive sample, x r,n Image x represents the nth image of the rth pedestrian. t′,a The hardest negative sample, where the second feature h represents the feature before regularization;

[0014] For each batch of data, train the a-th image x of pedestrian t′. t′,a Calculate the image x t′,aThe first feature f((x) within camera q t′,a ) q ), and based on the first feature f((x) t′,a ) q ) and the global index j q The distance to the corresponding pedestrian prototype feature, and the first feature f((x) t′,a ) q The minimum distance between the prototype features of the pedestrian and other pedestrians e within the camera q is used to construct the in-camera prototype loss L. IP ;

[0015] According to the cross-entropy loss L intra_ID The example triplet loss L IT and the prototype loss L in the camera IP The pedestrian re-identification model is trained and optimized within the camera.

[0016] Furthermore, the calculation of pedestrian image x i The global index j of the t-th pedestrian prototype within camera q. q ,include:

[0017] Get pedestrian images x i and the pedestrian image x i The local identity tag of the t-th pedestrian in the camera q.

[0018] Get the number M of pedestrian prototypes under each camera k before camera q. k ;

[0019] The number M of the pedestrian prototypes k Summing yields the sum B of all prototypes prior to camera q. q-1 ;

[0020] Calculate pedestrian image x i The global index of the t-th pedestrian prototype within camera q.

[0021] Furthermore, the matching of identical pedestrian prototypes across cameras includes:

[0022] Get the prototype set under camera q The set of prototypes under camera q′ Where q≠q′, Let d represent the prototype of the d-th pedestrian under camera q. M represents the prototype of the z-th pedestrian under camera q′. q M q′ Let q and q′ represent the total number of pedestrian prototypes under camera q and camera q′, respectively.

[0023] Calculate the pedestrian prototype With the prototype set G q′ Each of the pedestrian prototypes The similarity is calculated, and based on the obtained first similarity, the similarity between the camera q′ and the pedestrian prototype is obtained. Most similar pedestrian archetype

[0024] Calculate the pedestrian prototype With the prototype set G q Each of the pedestrian prototypes The similarity is calculated, and based on the obtained second similarity, the similarity between the camera q′ and the pedestrian prototype is obtained. Most similar pedestrian archetype

[0025] In the pedestrian prototype And the prototype of the pedestrian With pedestrian prototype If the similarity is greater than a set threshold, the pedestrian prototype is determined to be... With pedestrian prototype Match successful.

[0026] Furthermore, the step of bringing the same pedestrian prototypes closer together and pushing different pedestrian prototypes further apart, based on the matched pedestrian prototypes, to perform cross-camera training of the pedestrian re-identification model, includes:

[0027] Using cross-camera prototype triplet loss L CT It increases the distance between different pedestrian prototypes across cameras and decreases the distance between the same pedestrians across cameras;

[0028] Using a cross-camera prototype to narrow down the loss L CP Optimize the relationship between pedestrian prototypes representing the same pedestrian across cameras;

[0029] According to the cross-camera prototype triplet loss L CT And the cross-camera prototype's proximity loss L CP The pedestrian re-identification model is trained and optimized across cameras.

[0030] Furthermore, the cross-camera prototype triplet loss Where C represents the total number of cameras, m2 represents the interval of the second triplet, and F qq′ This indicates the prototype under camera q. The mapping is to the prototype matched in camera q′.

[0031] Furthermore, the cross-camera prototype zoom-in loss Where C represents the total number of cameras, F qq′ This indicates the prototype under camera q. The mapping is to the prototype matched in camera q′.

[0032] Furthermore, the backbone network of the pedestrian re-identification model is updated based on the exponential mean.

[0033] A pedestrian re-identification device based on cross-camera prototype learning, the device comprising:

[0034] The in-camera training module is used to learn the pedestrian prototype of each pedestrian in each camera using labeled in-camera supervision information, so as to train the pedestrian re-identification model in-camera; wherein, the backbone network of the pedestrian re-identification model is updated based on the exponential mean.

[0035] The cross-camera training module is used to match the same pedestrian prototypes across cameras, and based on the matched pedestrian prototypes, to bring the same pedestrian prototypes between different cameras closer and push different pedestrian prototypes further away, so as to train the pedestrian re-identification model across cameras.

[0036] The prediction module is used to predict pedestrian re-identification based on the pedestrian re-identification model that has completed in-camera training and cross-camera training.

[0037] A computer device, characterized in that the computer device comprises: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the pedestrian re-identification method based on cross-camera prototype learning as described above.

[0038] Unlike existing methods that only explore the relationship between pedestrian instances across cameras while ignoring the relationship between prototypes across cameras, this invention proposes a camera-independent pedestrian prototype alignment scheme to explicitly bring the same pedestrian prototypes closer together and push different pedestrian prototypes further apart under different cameras to improve the performance of the retrieval model. This scheme aims to eliminate the influence of factors such as camera background, lighting, and angle on the retrieval model. Attached Figure Description

[0039] Figure 1 The difference between this invention and previous methods.

[0040] Figure 2 Overall architecture diagram of a pedestrian re-identification method based on cross-camera prototype learning.

[0041] Figure 3 Pedestrian features visualized.

[0042] Figure 4 Compatibility analysis.

[0043] Figure 5 Similarity threshold analysis. Detailed Implementation

[0044] The present invention will now be described in detail with reference to the embodiments and accompanying drawings. The following embodiments only represent one possible implementation of the present invention and are not intended to limit the present invention.

[0045] This invention proposes a pedestrian re-identification method based on cross-camera prototype learning (CCPL). Figure 1 As can be seen from the differences between the present invention and previous methods, the previous methods were limited to exploring the relationship between pedestrian instances across cameras, while ignoring the relationship between prototypes across cameras. The present invention, however, further explores the relationship between pedestrian prototypes across cameras.

[0046] This invention mainly consists of two stages. In the intra-camera learning stage, labeled intra-camera supervised information is used to learn the prototype features of pedestrians in each camera through a multi-task learning strategy. In the inter-camera learning stage, this invention first uses a prototype matching mechanism to match identical pedestrian prototypes across cameras. Then, based on this, it explicitly brings identical pedestrian prototypes closer together and pushes different pedestrian prototypes further apart from those in different cameras. Finally, an inter-camera prototype convergence loss is introduced to further reduce the distance between identical pedestrian prototypes in different cameras.

[0047] (a) In-camera training.

[0048] The overall framework diagram of the present invention is as follows: Figure 2 As shown, the backbone network of this invention uses the exponential mean (EMA) method for updating, comprising two parts: an EMA convolutional network (EMA CNN) and a regular convolutional network (CNN). The EMA CNN utilizes the weights of the CNN for cumulative updates. During the in-camera training phase, this invention utilizes local pedestrian annotation information within the camera and employs a multi-task learning strategy to generate pedestrian prototypes for each camera, with each prototype representing a learnable parameter. During training, the feature extraction (CNN) for each camera shares weights. In this way, the learned prototypes possess implicit inter-camera discriminative capabilities, thereby facilitating subsequent cross-camera pedestrian prototype matching. Specifically, first, given a pedestrian image x... i Pedestrian Images x i The local identity tag of the t-th pedestrian in the camera q. (i, q, and t are all natural numbers), and then, this invention can be achieved through... Get image x iThe global index of the t-th pedestrian in the pedestrian prototype within camera q, where B is the sum of all prototypes before camera q (from the 1st camera to the (q–1)-th camera). q-1 It can be obtained through the formula M k represents the number of all pedestrian prototypes under the k-th camera (k < q). Then the objective of in-camera learning can be formulated as:

[0049]

[0050] where N represents the total number of training set images, and g[j q is the pedestrian prototype feature corresponding to the global index j q , which is a learnable network parameter. τ1 is the temperature coefficient used to control smoothness. M q represents the number of all pedestrian prototypes under camera q, and f((x i ) q ) is the feature of the t-th pedestrian in image x i after batch normalization (BN) within camera q, and L intra_ID is the in-camera pedestrian identity cross-entropy loss.

[0051] To improve the intra-class compactness under each camera, the present invention further introduces an instance triplet loss and an in-camera prototype loss, and further optimizes the in-camera training by mining hard positive-negative instance-instance pairs and hard positive-negative instance-prototype pairs. Specifically, the instance triplet loss L IT and the in-camera prototype loss L IP can be formulated as:

[0052]

[0053]

[0054] where P is the number of pedestrians sampled in each batch of data training, K represents the number of images sampled for each pedestrian, x t′,a represents the a-th image of the t'-th pedestrian in each batch of data training, x t′,p represents the p-th image of the t'-th pedestrian in each batch of data training, x r,n represents the n-th image of the r-th pedestrian in each batch of data training, represents which camera the a-th image of the t-th pedestrian belongs to, represents which camera the n-th image of the r-th pedestrian belongs to, the local identity label of the a-th image of the t-th pedestrian in camera q, N q represents the total number of local identity labels of the a-th image in camera q, Let m represent the e-th local identity label of the a-th image in camera q, and m1 be the interval of the first triplet loss. and These are the features before batch normalization (BN) of the anchor points, positive samples, and negative samples, where dist(,) represents the cosine distance between two features. The final loss function for in-camera training can be expressed as:

[0055] L Intra =L intra_ID +L IT +L IP (4)

[0056] (II) Camera-to-Camera Training:

[0057] During the inter-camera training phase, this invention first explores the relationships between prototypes from different cameras. This invention argues that for two prototypes from different cameras, the more similar they are, the more likely they represent the same pedestrian. Based on this, this invention introduces a prototype matching mechanism. Given the set of prototypes for camera q... The set of prototypes under camera q′ Where q and q′ represent camera labels (which are natural numbers), Let d represent the prototype of the d-th pedestrian under camera q. M represents the prototype of the z-th pedestrian under camera q′. q M q′ These represent the total number of prototypes under camera q and camera q′, respectively. This invention believes... It means that the same pedestrian is present if and only if and The prototypes are the most similar to each other under their respective cameras, and the similarity between them is greater than a certain threshold. Specifically, this invention calculates... and The similarity between them can be formalized into the following formula:

[0058]

[0059] in, express and The similarity between them. Then, the present invention can obtain the similarity between them under camera q′. Most similar prototype:

[0060]

[0061] z * Most similar prototype The subscript.

[0062] Similarly, prototype It can be done In G q The most similar prototype is calculated in the middle. G q and G q′ The matching between them can be obtained using the following formula:

[0063]

[0064] θ is the threshold, and sim(,) is the similarity between two features. The above formula guarantees that... The prototypes are the most similar to each other under their respective cameras, and the similarity between them is greater than a specified threshold.

[0065] Previous methods for independent camera scenarios explored the relationships between instances across different cameras by designing multi-label strategies, soft-label strategies, and consistency preservation learning. However, these methods all neglected the relationships between cross-camera prototypes. Therefore, this invention proposes a simple and effective camera-independent prototype alignment (CPA) method to optimize the relationships between prototypes across different cameras. This invention argues that the distance between different pedestrian prototypes across cameras should be pushed further apart, while the distance between the same pedestrians across cameras should be shortened. Specifically, the cross-camera prototype triplet loss proposed in this invention can be formalized as:

[0066]

[0067] Where C is the total number of cameras. m2 is the interval of the second triplet, F qq′ The prototype under camera q Mapped to the prototype matched in camera q′, T q,q′ It is the set of prototypes that are matched under camera q'.

[0068] Although the cross-camera prototype triplet loss L CT While it can be guaranteed that the distance between prototypes representing the same pedestrian across cameras is less than the distance between prototypes representing different pedestrians, the distance between prototypes representing the same pedestrian across cameras remains large due to factors such as camera viewpoint, lighting, and background. This affects the construction of a discriminative feature space. To further optimize the relationship between prototypes representing the same pedestrian across cameras, this invention introduces a cross-camera prototype convergence loss L. CP It can be formalized as:

[0069]

[0070] Combining all the above losses, the final optimization function of this invention can be expressed as:

[0071] L = L Intra +λ1LCT +λ2L CP (10)

[0072] Where λ1 represents the first weighting coefficient and λ2 represents the second weighting coefficient.

[0073] The present invention was tested on three basic datasets: Market1501, DukeMTMC-ReID, and MSMT17. Market1501 contains 1,501 people and 32,668 images from 6 cameras; the DukeMTMC-ReID dataset contains 1,404 pedestrians and 36,411 images from 8 cameras; and the MSMT17 dataset contains 4,101 pedestrians and 126,441 images from 15 cameras. The cumulative matching feature (CMC) and average precision across all classes (mAP) were used to evaluate the model of the present invention. A higher mAP indicates stronger retrieval performance.

[0074] Experimental results: In order to fully verify the method of the present invention, the present invention also compared (1) some supervised pedestrian re-identification methods and (2) all pedestrian re-identification methods in independent camera scenes.

[0075] Table 1 shows the protection effects of different methods on different models. From this invention, we can draw the following conclusions: Compared with other pedestrian re-identification methods in independent camera scenes, the proposed CCPL method achieves the best performance, especially on the challenging MSMT17 dataset, where it far surpasses other methods. Furthermore, compared with supervised pedestrian re-identification methods, the proposed CCPL still achieves competitive results, even exceeding the baseline supervised method PCB, demonstrating the effectiveness of this invention.

[0076]

[0077] Table 1

[0078] Ablation experiments: As shown in Table 2, and only the first stage of in-camera training was performed (using L... ID L IT L IP Compared to this, when the L in the second-stage CPA module is added to this invention... CT The accuracy of the time-based retrieval model was significantly improved. Similarly, during the inter-camera training phase, compared to using only the L module in the CPA module, the accuracy was significantly improved. CT In comparison, using the entire CPA module can further improve the performance of the retrieval model. The large distance between the same pedestrian prototypes under different cameras leads to significant differences in the features extracted by the final retrieval model for the same pedestrian under different cameras. Experimental results demonstrate that relying solely on L... CTThis is insufficient to solve the above problems, while the L proposed in this invention... CP The relationship between cross-camera prototypes can be further optimized.

[0079]

[0080] Table 2

[0081] Prototype Matching Mechanism Analysis: As shown in Table 3, this invention studies how the prototype matching mechanism affects the performance of the retrieval model and investigates four prototype matching strategies. Given camera c i The prototype below This invention can be found in camera c j The most similar prototype Two-way selection means, and They are the most similar prototypes under their respective cameras, represented by a one-way selection method. yes In camera c j The most similar prototype does not need to satisfy yes In camera c i The most similar prototype is identified. Experimental results demonstrate that the method described in this invention is robust to prototype matching mechanisms. Furthermore, this invention also reveals that the similarity threshold between two prototypes plays a crucial role. When combined with bidirectional selection, the retrieval model achieves optimal results.

[0082]

[0083] Table 3

[0084] Visualization Analysis: To further illustrate the role of the CPA module in this invention, the t-SNE plot is used to visualize the feature space distribution of the prototype and instances after learning through the CPA module. For example... Figure 3 As shown, this invention visualizes the prototype and instance features of 12 pedestrians from three cameras. It can be observed that before using CPA learning, the features of the same pedestrians under different cameras showed significant differences. However, after learning through the CPA module, this invention shows that the features of the same pedestrians under different cameras become more compact, verifying the effectiveness of the proposed CPA module.

[0085] Compatibility with other baseline methods: This invention attempts to combine the proposed CPA module with the recently proposed complex but high-performance CDL method, from... Figure 4As can be seen, the method of this invention can be easily combined with other methods and can further improve the performance of the CDL method. This proves the effectiveness of the CPA module proposed in this invention.

[0086] Prototype matching similarity analysis: This invention analyzes multiple parameters related to the threshold in the prototype matching mechanism, from... Figure 5 As can be seen, the model's performance significantly improves when the threshold increases from 0 to 0.2. The performance stabilizes when the threshold increases from 0.2 to 0.8. The performance begins to decline when the threshold increases from 0.8 to 0.9. The model of this invention is robust within the range [0.2, 0.8], exhibiting excellent performance.

[0087] In summary, after learning strong features for discrimination within each camera, this invention further explores the relationship between pedestrian prototypes across cameras. Specifically, it improves the performance of the pedestrian retrieval model by explicitly narrowing the distance between the same pedestrian prototypes under different cameras and widening the distance between different pedestrians under different cameras.

[0088] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and to implement it accordingly. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification; the scope of protection of the present invention is defined by the claims.

Claims

1. A pedestrian re-identification method based on cross-camera prototype learning, characterized in that, The method includes: Using labeled in-camera supervision information, we learn pedestrian prototypes for each camera to train the pedestrian re-identification model in-camera. Match the same pedestrian prototypes across cameras, and based on the matched pedestrian prototypes, bring the same pedestrian prototypes closer to different cameras and push the different pedestrian prototypes further away, so as to train the pedestrian re-identification model across cameras. Pedestrian re-identification prediction is performed based on pedestrian re-identification models that have completed in-camera and cross-camera training. The matching of identical pedestrian prototypes across cameras includes: Get camera The prototype set and camera The prototype set ;in, , Indicates camera Next The prototype of a pedestrian. Indicates camera Next The prototype of a pedestrian. , They represent cameras ,camera The total number of pedestrian prototypes; Calculate the pedestrian prototype With the prototype set Each pedestrian prototype The similarity is calculated, and based on the obtained first similarity, the camera is determined. Below is the pedestrian prototype Most similar pedestrian archetype ; Calculate the pedestrian prototype With the prototype set Each pedestrian prototype The similarity is calculated, and based on the obtained second similarity, the camera is determined. Below is the pedestrian prototype Most similar pedestrian archetype ; In the pedestrian prototype And it is consistent with the prototype of the pedestrian. With pedestrian prototype If the similarity is greater than a set threshold, the pedestrian prototype is determined to be... With pedestrian prototype Match successful; The process of training the pedestrian re-identification model across cameras by bringing identical pedestrian prototypes closer together and pushing different pedestrian prototypes further apart, based on the matched pedestrian prototypes, includes: Using cross-camera prototype triplet loss This increases the distance between different pedestrian prototypes across cameras and decreases the distance between the same pedestrians across cameras; wherein, the cross-camera prototype triplet loss , This indicates the total number of cameras. Indicates the interval of the second triplet. Indicates that the camera The prototype below Mapped to camera The prototype matched in; Use a cross-camera prototype to narrow down the loss. Optimize the relationship between pedestrian prototypes representing the same pedestrian across cameras; wherein, the cross-camera prototype convergence loss , This indicates the total number of cameras. Indicates that the camera The prototype below Mapped to camera The prototype matched in; Based on the cross-camera prototype triplet loss And the loss of the cross-camera prototype The pedestrian re-identification model is trained and optimized across cameras.

2. The method as described in claim 1, characterized in that, The method of using labeled in-camera supervision information to learn the prototype features of pedestrians in each camera for in-camera training of the pedestrian re-identification model includes: Calculate pedestrian images The first in A pedestrian was photographed. Global Index of the Insider Prototype ;in, , , , It is a natural number; Based on the pedestrian images In the camera The first feature inside and the global index Based on the corresponding pedestrian prototype features, calculate the cross-entropy loss of the pedestrian re-identification model during in-camera training. ; wherein, the first feature These are the features after regularization; For each batch of data, train pedestrians The Zhang Image Calculate the image In the camera Second feature inside and in combination with the aforementioned features The second feature of the farthest positive sample and the second feature of the nearest negative sample Constructing a triplet loss for mining difficult positive and negative instance-instance pairs. ;in, Indicates the first The first pedestrian The picture is a picture. The most difficult positive sample, Indicates the first The first pedestrian The picture is a picture. The hardest negative sample, the second feature Represents the features before regularization; for each batch of data trained on pedestrians The Zhang Image Calculate the image In the camera The first feature inside And based on the first feature With the global index The distance to the corresponding pedestrian prototype feature, and the first feature With the camera Other pedestrians The minimum distance to the corresponding pedestrian prototype features is used to construct the in-camera prototype loss. ; According to the cross-entropy loss The example triplet loss and the prototype loss within the camera The pedestrian re-identification model is trained and optimized within the camera.

3. The method as described in claim 2, characterized in that, The calculation of pedestrian images The first in A pedestrian was photographed. Global Index of the Insider Prototype ,include: Get pedestrian images and the pedestrian images The first in A pedestrian was photographed. The corresponding local identity tag ; Get camera Every previous camera Number of pedestrian prototypes ; The number of the pedestrian prototypes Summing is performed to obtain the camera's... The sum of all previous prototypes ; Calculate pedestrian images The first in A pedestrian was photographed. Global Index of the Insider Prototype .

4. The method according to any one of claims 1 to 3, characterized in that, The backbone network of the pedestrian re-identification model is updated based on the exponential average.

5. A pedestrian re-identification device based on cross-camera prototype learning, characterized in that, The device includes: The in-camera training module is used to learn the pedestrian prototype of each pedestrian in each camera using labeled in-camera supervision information, so as to train the pedestrian re-identification model in-camera; wherein, the backbone network of the pedestrian re-identification model is updated based on the exponential mean. The cross-camera training module is used to match the same pedestrian prototypes across cameras, and based on the matched pedestrian prototypes, to bring the same pedestrian prototypes between different cameras closer and push different pedestrian prototypes further away, so as to train the pedestrian re-identification model across cameras. The prediction module is used to predict pedestrian re-identification based on the pedestrian re-identification model that has completed in-camera training and cross-camera training. The matching of identical pedestrian prototypes across cameras includes: Get camera The prototype set and camera The prototype set ;in, , Indicates camera Next The prototype of a pedestrian. Indicates camera Next The prototype of a pedestrian. , They represent cameras ,camera The total number of pedestrian prototypes; Calculate the pedestrian prototype With the prototype set Each pedestrian prototype The similarity is calculated, and based on the obtained first similarity, the camera is determined. Below is the pedestrian prototype Most similar pedestrian archetype ; Calculate the pedestrian prototype With the prototype set Each pedestrian prototype The similarity is calculated, and based on the obtained second similarity, the camera is determined. Below is the pedestrian prototype Most similar pedestrian archetype ; In the pedestrian prototype And it is consistent with the prototype of the pedestrian. With pedestrian prototype If the similarity is greater than a set threshold, the pedestrian prototype is determined to be... With pedestrian prototype Match successful; The process of training the pedestrian re-identification model across cameras by bringing identical pedestrian prototypes closer together and pushing different pedestrian prototypes further apart, based on the matched pedestrian prototypes, includes: Using cross-camera prototype triplet loss This increases the distance between different pedestrian prototypes across cameras and decreases the distance between the same pedestrians across cameras; wherein, the cross-camera prototype triplet loss , This indicates the total number of cameras. Indicates the interval of the second triplet. Indicates that the camera The prototype below Mapped to camera The prototype matched in; Use a cross-camera prototype to narrow down the loss. Optimize the relationship between pedestrian prototypes representing the same pedestrian across cameras; wherein, the cross-camera prototype convergence loss , This indicates the total number of cameras. Indicates that the camera The prototype below Mapped to camera The prototype matched in; Based on the cross-camera prototype triplet loss And the loss of the cross-camera prototype The pedestrian re-identification model is trained and optimized across cameras.

6. A computer device, characterized in that, The computer device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the pedestrian re-identification method based on cross-camera prototype learning as described in any one of claims 1-4.