Unsupervised pedestrian searching method for original monitoring video

By using graph segmentation and multi-branch comparison learning methods in the original surveillance video, the problem of unknown pedestrian identity categories is solved, pseudo-labels are generated and feature encoder is optimized, which improves the accuracy of pedestrian search and is suitable for pedestrian feature consistency in multi-camera scenarios.

CN120339916AActive Publication Date: 2025-07-18LIAOCHENG UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510482521.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The prior art is difficult to accurately find specific pedestrian targets in original surveillance videos, especially when the pedestrian identity category is unknown and cannot be marked manually, and existing methods have application limitations and inefficiencies.

Method used

The pre-trained neural network is used to detect the pedestrian area, use the feature encoder to generate pedestrian features, and determine the identity category through the unsupervised clustering method of graph segmentation. Combined with multi-branch comparison learning, improve the distinction and similarity of pedestrian features, generate pseudo-labels, and optimize the parameters of the feature encoder.

Benefits of technology

It realizes that the accuracy of pedestrian search is improved without manually marking pedestrian identity categories, and is suitable for pedestrian characteristics consistency problems in multi-camera scenarios, and outputs pedestrian characteristics with good distinction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339916A_ABST
    Figure CN120339916A_ABST
Patent Text Reader

Abstract

The invention relates to an original monitoring video-oriented unsupervised pedestrian search method, which comprises the following steps of: detecting a pedestrian area in an original monitoring video by adopting a pre-trained neural network, and generating pedestrian features by adopting a feature encoder; in an original monitoring video, based on visual similarity and context similarity of pedestrians, searching for a first neighbor relation of pedestrian features and organizing a chain structure of the pedestrian features, determining spatial distribution of the pedestrian features, realizing unsupervised clustering according to segmented regions, and generating identity categories and pseudo labels of the pedestrians; in a forward propagation process, respectively calculating a contrast loss function of each branch from three branches of category distinguishing learning, single-camera contrast learning and multi-camera contrast learning, and respectively updating a clustering center memory bank and a multi-camera memory bank according to a momentum updating mechanism; and performing joint optimization on the contrast loss functions of the three branches, minimizing a total loss value through multiple iterations in a back propagation process, and optimizing parameters of the feature encoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and specifically to an unsupervised pedestrian search method for original surveillance videos. Background Art

[0002] Due to the exponential growth trend of video data collected by surveillance cameras, many problems still exist in the analysis of surveillance videos. On the one hand, the current processing of original surveillance videos mainly relies on "manual analysis" and is assisted by simple intelligent methods, which often leads to technical problems such as "the video exists but the target cannot be found", "the target can be found but it takes too long", and "there is a service but it is unreliable", thus causing the video screening work to often consume a large amount of time, manpower, and material resources. On the other hand, manual annotation of original surveillance videos faces problems such as a huge amount of data, high annotation time cost, slow annotation speed, and difficulty in ensuring annotation consistency. The workload is extremely heavy, and this work is often not feasible during the process of solving cases. In view of the above problems, how to accurately find specific pedestrian targets in original surveillance videos not only has important theoretical research significance in the field of computer vision but also has practical application value for maintaining social public order.

[0003] In recent years, pedestrian re-identification methods have been widely studied in the academic circles at home and abroad. This method pre-crops pedestrian images through bounding boxes, separates them from the surrounding background, and simultaneously annotates the identity categories of pedestrians. Then, the given pedestrian targets are associated and matched in the cropped pedestrian image library. Most pedestrian re-identification methods are based on pedestrian images with annotated identity categories and use supervised learning to train pedestrian re-identification models in image libraries such as CUHK, Market-150, DukeMTMC-reID, and MSMT17. However, there are neither prior pedestrian bounding boxes nor pre-determined pedestrian identity categories in the original surveillance videos, resulting in the inability of pedestrian re-identification methods using supervised learning to directly train pedestrian search models, and thus there are limitations in the application in the actual surveillance field. Through literature retrieval of the existing technologies, it is found that Patent CN114821661B provides an "unsupervised pedestrian re-identification method, device, equipment, and readable storage medium". In this method, a deep learning model is used to generate pedestrian features for the cropped pedestrian images, and after clustering, pedestrian features are regenerated for the pre-processed pedestrian images, and the deep learning model is trained in this cycle. However, this method trains on the cropped pedestrian images and requires the pre-setting of the number of pedestrian image categories, and cannot be applied to the original surveillance videos with unknown pedestrian bounding boxes and the number of pedestrian categories. Patent CN113065409B provides a "method for unsupervised pedestrian re-identification based on alignment constraints of camera distribution differences". This method belongs to the domain adaptation re-identification algorithm and also targets the cropped pedestrian images, mainly focusing on using the triplet loss function to solve the problem of the distribution consistency between the source domain and the target domain, and then using the source domain with annotated pedestrian categories to improve the accuracy of the target domain with unannotated pedestrian categories. However, there is no distinction between the source domain and the target domain in the directly collected original surveillance videos, and at the same time, this algorithm does not consider the spatio-temporal characteristics of the original surveillance videos, so there are certain limitations in the application in the actual surveillance field.

[0004] The pedestrian search method introduces pedestrian detection on the basis of pedestrian re-identification, narrowing the gap between pedestrian re-identification and actual surveillance applications. Compared with the pedestrian re-identification method based on cropped images, the pedestrian search method better meets the needs of actual surveillance. According to different integration methods of pedestrian detection, the existing pedestrian search methods can be divided into end-to-end methods and two-stage methods. Through the literature retrieval of the existing technologies, it is found that the end-to-end pedestrian search framework was first proposed in the literature "Joint Detection and Identification Feature Learning for Person Search" (CVPR 2017). In this framework, pedestrian detection and pedestrian re-identification are integrated in the Fast R-CNN neural network, and supervised training is carried out on the CHUK-SYSU and PRW image databases by using the Online Instance Matching Loss (OIM). Further retrieval reveals that the two-stage method was proposed in the literature "Person Search via A Mask-guided Two-stream CNN Model". In this method, pedestrian detection and pedestrian re-identification are processed separately through a mask-guided two-stream convolutional neural network, and mask fusion features are used to enhance the distinctiveness of pedestrian features. However, it still requires supervised training on the CHUK-SYSU and PRW image databases according to the labeled pedestrian identity categories. The literature "Person Search in Videos with One Portrait Through Visual and Temporal Links" (ECCV 2018) searches for specific persons through visual and temporal links in long videos. However, this method constructs a video dataset based on movie clips and uses manually annotated trajectory segments for supervised training. Patent CN110826424A provides a "pedestrian search method based on pedestrian re-identification-driven localization adjustment". This method realizes the supervision of the detection boxes output by the pedestrian detection network by the loss of the pedestrian re-identification network through constructing a pedestrian re-identification-driven localization adjustment model. However, this method still belongs to the supervised pedestrian search method.

[0005] Through research and analysis of the references in the field of person re-identification and person search, it is found that although unsupervised person re-identification methods have achieved certain development, these methods mainly use DBSCAN clustering to generate pseudo-labels for person features, and then use cross-entropy loss, label smoothing or non-parametric Softmax loss to train the person re-identification model. DBSCAN clustering divides the feature space based on density and classifies categories through the neighborhood radius and density threshold. However, it is sensitive to the selection of the above two parameters, and different parameters have a significant impact on the confidence of person pseudo-labels. At the same time, this method does not consider the spatio-temporal limitations of person features and is difficult to apply to the person feature space with uneven density in actual monitoring scenarios. At the same time, cross-entropy loss, label smoothing or non-parametric Softmax loss functions are difficult to solve the problem of person feature consistency in multiple surveillance camera scenarios. Compared with unsupervised person re-identification methods, most current person search methods still use supervised learning to train person search models in the CHUK-SYSU and PRW image databases with labeled person identity categories. However, it is not feasible to manually label the identity categories of persons in a large number of original surveillance videos, and the total number of person categories is unknown, which to a certain extent hinders the application of supervised person search methods in actual monitoring scenarios. Summary of the Invention

[0006] Aiming at the above deficiencies existing in the prior art, the present invention provides an unsupervised person search method for original surveillance videos. Aiming at the problem that the person identity categories in the original surveillance videos are unknown and difficult to label, after extracting person features from the original surveillance videos, an unsupervised clustering method based on graph segmentation is used to determine the person identity categories and generate corresponding pseudo-labels. On this basis, aiming at the problem of person feature consistency in multiple surveillance camera scenarios, a multi-branch contrast learning method is adopted, which, from three branches of category discriminative learning, single-camera contrast learning, and multi-camera contrast learning, while distinguishing different person features in a single-camera scenario, enhances the similarity of the same person features in a multi-camera scenario, thereby improving the accuracy of person search. The present invention is implemented through the following technical solutions.

[0007] An unsupervised person search method for original surveillance videos, characterized by comprising the following steps:

[0008] Use a pre-trained neural network to detect person regions in the original surveillance video, and use a feature encoder to generate person features;

[0009] The problem of clustering pedestrian features with an unknown total number of categories is regarded as a graph segmentation problem. Each node represents a pedestrian feature, the connections between nodes represent the similarity between pedestrian features, and each segmented area represents a clustering category. In the original surveillance video, multiple pedestrians appearing in a single camera at the same moment belong to different identity categories. However, at different moments, the same pedestrian may appear in different cameras. Based on this, the weighted total similarity is calculated based on the visual similarity and context similarity of pedestrian features, the first-nearest neighbor relationship of pedestrian features is found and the chain structure of pedestrian features is organized, the spatial distribution of pedestrian features is determined, and unsupervised clustering is achieved according to the segmented areas. The identity category of pedestrians is labeled according to the clustering category, and pseudo-labels of pedestrians are generated according to the clustering centers.

[0010] During the forward propagation process, from three branches of category discriminative learning, single-camera contrast learning, and multi-camera contrast learning, the contrast loss functions of each branch are calculated respectively, and the clustering center memory bank and multi-camera memory bank are updated respectively according to the momentum update mechanism.

[0011] The contrast loss functions of the three branches are jointly optimized, and the total loss value is minimized through multiple iterations during the backpropagation process, thereby optimizing the parameters of the feature encoder.

[0012] Furthermore, the use of a pre-trained neural network to detect pedestrian regions in the original surveillance video and the use of a feature encoder to generate pedestrian features means that: the original surveillance video consists of multiple video segments and their corresponding camera labels. In each video segment, a pre-trained neural network is used to detect pedestrian bounding boxes to determine pedestrian regions; a feature encoder is used to generate corresponding pedestrian features for each pedestrian region respectively.

[0013] Furthermore, the use of a pre-trained neural network to detect pedestrian regions in the original surveillance video and the use of a feature encoder to generate pedestrian features specifically include the following steps:

[0014] 1) For the original surveillance video V = {(v i , s j )}, v i (1 ≤ i ≤ M) is the i-th video segment, M is the number of video segments, s j (1 ≤ j ≤ N) is the j-th camera label corresponding to v i , indicating that the video segment v i is captured by the surveillance camera s j ; the video segment v i is processed by a pre-trained neural network to detect the pedestrian region r i , and all pedestrian regions in the original surveillance video V are represented as R = {(r i , s j)};

[0015] 2) For the pedestrian area r i , generate the pedestrian feature d θ through the feature encoder f i =f θ (r i ), and all the pedestrian features in the original surveillance video V are represented as D={(d i , s j )}={(f θ (r i ), s j )}.

[0016] Further, based on the visual similarity and context similarity of pedestrian features, calculate the weighted total similarity, find the first-nearest neighbor relationship of pedestrian features and organize the chain structure of pedestrian features, determine the spatial distribution of pedestrian features and achieve unsupervised clustering according to the segmentation area; label the identity category of pedestrians according to the clustering category, and generate pseudo-labels for pedestrians according to the clustering center, specifically including the following steps:

[0017] 1) In all pedestrian features D, calculate the visual similarity P(m,n)=sim(d i,m , d j,n )=sim(f θ (r i,m ), f θ (r j,n ))), where d i,m represents the m-th pedestrian feature in the i-th video segment v i , and d j,n represents the n-th pedestrian feature in the j-th video segment v j ; calculate the context similarity of the above pedestrian features where N i is the number of pedestrian features in the video segment v i , and N j is the number of pedestrian features in the video segment v j ; the weighted total similarity of two pedestrian features is calculated as T(m,n)=λP(m,n)+(1 - λ)Q(m,n), where λ is the trade-off coefficient between the two similarities; sort the total similarity T(m,n) from largest to smallest, and find the first-nearest neighbor relationship of the pedestrian feature d i,m

[0018] 2) According to the first-nearest neighbor relationship of pedestrian features, construct a symmetric sparse matrix A(m,n,t)∈{0,1} for graph segmentation. At time t, the matrix where ​Indicates that the first neighbor of the m-th pedestrian feature is the pedestrian feature n, Indicates that the first neighbor of the n-th pedestrian feature is the pedestrian feature m, Indicates that the first neighbors of the m-th pedestrian feature and the n-th pedestrian feature are the same, t m and t n respectively represent the moments when the m-th pedestrian feature and the n-th pedestrian feature appear; if A(m,n,t) = 1, it indicates that there is a connection between the m-th pedestrian feature and the n-th pedestrian feature; based on the chain structure between pedestrian features, connected components between pedestrian features are constructed; the segmentation regions of pedestrian features are determined through these connected components, and each segmentation region represents a clustering category;

[0019] 3) After determining the clustering categories among all pedestrian features D, label the identity categories of pedestrians through the clustering categories, and the clustering centers are represented as C = {c i}(1 ≤ i ≤ K), and generate pseudo-labels for pedestrians according to the clustering centers.

[0020] Furthermore, in the forward propagation process, from three branches of class discriminative learning, single-camera contrast learning, and multi-camera contrast learning, calculate the contrast loss functions of each branch respectively, and update the clustering center memory bank and the multi-camera memory bank according to the momentum update mechanism, which means: in the class discriminative learning branch, calculate the class contrast loss function to ensure the distinctiveness between different pedestrian features; in the single-camera contrast learning branch, calculate the single-view contrast loss function to distinguish the distinctiveness of different pedestrian features under the same view; in the multi-camera contrast learning branch, calculate the multi-view contrast loss function to improve the similarity of the same pedestrian features under different views; after calculating the loss functions of the above three branches, update the clustering center memory bank and the multi-camera memory bank according to the momentum update mechanism.

[0021] Furthermore, the steps of calculating the contrast loss functions of each branch respectively from three branches of class discriminative learning, single-camera contrast learning, and multi-camera contrast learning, and updating the clustering center memory bank and the multi-camera memory bank according to the momentum update mechanism in the forward propagation process include:

[0022] 1) In the class discriminative learning branch, calculate the class contrast loss function where q represents the query pedestrian feature in the training set, c + represents the clustering center corresponding to the query pedestrian feature q, represents the i-th clustering center in the t-th iteration, τ c represents the temperature coefficient; according to the momentum update mechanism, update the clustering center memory bank M c , and calculate as where q hardrepresents the query pedestrian feature with the minimum visual similarity to its cluster center in the training set in the training set, and λ c represents the momentum update coefficient of the cluster center memory bank;

[0023] 2) In the single-camera contrast learning branch, calculate the single-view contrast loss function where p + represents the relevant cluster center corresponding to q′ hard and q′ hard represents the query pedestrian feature with the minimum visual similarity to p + in the single-camera training set, represents an irrelevant cluster center that is relatively similar to the query pedestrian feature q, and K - represents the number of irrelevant cluster centers, and τ c represents the temperature coefficient;

[0024] 3) In the multi-camera contrast learning branch, calculate the multi-view contrast loss function where represents the i-th cluster center with camera label s j in the t-th iteration, represents the relevant cluster center with a different camera label corresponding to q easy and q easy represents the query pedestrian feature with the maximum visual similarity to in the multi-camera training set, and τ c represents the temperature coefficient; the multi-camera memory bank M p is composed of N single-camera memory banks, and the j-th surveillance camera corresponds to the single-camera memory bank M p,j ; according to the momentum update mechanism, update the single-camera memory bank M p,j and the calculation is where λ p represents the momentum update coefficient of the single-camera memory bank.

[0025] Furthermore, the joint optimization of the contrast loss functions of the three branches, and minimizing the total loss value through multiple iterations in the backpropagation process to optimize the parameters of the feature encoder means: summing up the contrast loss functions of the three branches of class discriminative learning, single-camera contrast learning, and multi-camera contrast learning, and solving the minimum value of the total loss function through multiple iterations in the backpropagation process to achieve the purpose of optimizing the parameters of the feature encoder.

[0026] Furthermore, the steps of the joint optimization of the contrast loss functions of the three branches, and minimizing the total loss value through multiple iterations in the backpropagation process to optimize the parameters of the feature encoder include:

[0027] 1) For the category contrast loss function $L$ cluster , the single - view contrast loss function $L$ Intra and the multi - view contrast loss function $L$ Inter , the total loss function $L$ is calculated as $L = L$ cluster + $L$ Intra + $L$ Inter ;

[0028] 2) During the back - propagation process, after $t$ iterations of training, the minimum value of the total loss function $L$ is solved to optimize the parameters $\theta$ in the feature encoder $f$ θ (·), so as to output discriminative pedestrian features.

[0029] The beneficial effects of the present invention are as follows: First, the present invention uses the idea of graph segmentation to perform unsupervised clustering on all pedestrian features extracted from the original surveillance video, determines the identity categories of pedestrians through the clustering categories and generates pseudo - labels, solving the problem that the pedestrian categories in the original surveillance video are unknown and difficult to label. Then, the unsupervised clustering and multi - branch contrast learning are integrated. On the basis of using category - discriminative learning to ensure the discriminability of different pedestrian characteristics, single - camera contrast learning is used to enhance the discriminability of different pedestrian features under the same view, and at the same time, multi - camera contrast learning is used to improve the similarity of the same pedestrian features under different views. Finally, the feature encoder is trained and optimized based on the loss function in the multi - branch contrast learning, and discriminative pedestrian features can be output. Compared with the prior art, the present invention can be directly applied to the original surveillance video, train the pedestrian search model without manually labeling the pedestrian identity categories, and ensure the accuracy of pedestrian search. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is the flow chart of the present invention.

[0031] Figure 2 is the flow chart for determining pedestrian categories by the unsupervised clustering method based on graph segmentation.

[0032] Figure 3 is the flow chart for training the category - discriminative learning branch.

[0033] Figure 4 is the flow chart for training the single - camera contrast learning branch and the multi - camera contrast learning branch.

[0034] Figure 5 is the comparison of the search accuracy between the unsupervised clustering method based on graph segmentation and the DBSCAN clustering method.

[0035] Figure 6 is the comparison of the search accuracy between the method of the present invention and the collaborative contrast refinement method and the context - guided feature learning method. DETAILED DESCRIPTION OF THE INVENTION

[0036] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.

[0037] This embodiment adopts an unsupervised pedestrian search method for the original surveillance video, and the specific implementation steps are as follows:

[0038] 1. First, use the pre-trained YOLO neural network to detect pedestrian regions in the original surveillance video, and use a feature encoder to generate pedestrian features.

[0039] (1) In the original surveillance video V = {(v i , s j )}, first use the pre-trained YOLO neural network to detect the pedestrian region r i in the i-th video segment v i , and at the same time label the corresponding camera tag s j . After processing M video segments, all pedestrian regions in the original surveillance video V can be expressed as R = {(r i , s j )};

[0040] (2) Use the Residual Neural Network ResNet50 as the feature encoder f θ (·), and generate the pedestrian feature d i = f i (r θ ) for the pedestrian region r i . Furthermore, all pedestrian features can be expressed as D = {(d i , s j )} = {(f θ (r i ), s j ).

[0041] 2. Then use an unsupervised clustering method based on graph segmentation to cluster all pedestrian features, determine the identity category of pedestrians according to the clustering categories, and generate corresponding pseudo-labels;

[0042] (1) Among all pedestrian features D, first calculate the visual similarity P(m, n) = sim(d i,m , d j,n ) = sim(f θ (r i,m ), f θ (r j,n)( ), and then calculate the context similarity of two pedestrian features based on visual similarity Finally, calculate the weighted total similarity of two pedestrian features \(T(m,n)=\lambda P(m,n)+(1 - \lambda)Q(m,n)\), where the trade-off coefficient \(\lambda = 0.7\). Sort the total similarity \(T(m,n)\) from largest to smallest to find the first nearest neighbor relationship of pedestrian feature \(d\) i,m of the first nearest neighbor relationship

[0043] (2) According to the first nearest neighbor relationship of pedestrian feature \(d\) i,m of the first nearest neighbor relationship Construct a symmetric sparse matrix \(A(m,n,t)\in\{0,1\}\) to perform graph segmentation on the feature space where pedestrian feature \(D\) is located. If \(A(m,n,t) = 1\), it means there is a connection between the \(m\)-th pedestrian feature and the \(n\)-th pedestrian feature. Based on the chain structure between pedestrian features, construct the connected components between pedestrian features, and then form the segmentation regions of pedestrian features. Each segmentation region represents a clustering category. By counting the number of clustering categories, the total number of pedestrian categories in the original surveillance video can be determined. Generate pseudo-labels for pedestrians according to the cluster centers \(C = \{c\}\) i \((1\leq i\leq K)\).

[0044] 3. Then, in the forward propagation process, calculate the contrast loss functions of each branch respectively from three branches: category discriminative learning, single-camera contrast learning, and multi-camera contrast learning, and update the cluster center memory bank and the multi-camera memory bank respectively according to the momentum update mechanism;

[0045] (1) In the category discriminative learning branch, for the query pedestrian feature \(q\) and the cluster center \(C\) in the training set, calculate the category contrast loss function where the temperature coefficient \(\tau\) c \(= 0.05\). Correspondingly, the cluster center memory bank \(M\) c is initialized by the cluster center \(C\). At the \(t\)-th iteration, it is updated by where the momentum update coefficient \(\lambda\) of the cluster center memory bank c \(= 0.1\);

[0046] (2) In the single-camera contrast learning branch, based on the camera labels, extract the single-camera training set and the cluster center, and calculate the single-view contrast loss function In the multi-camera contrast learning branch, based on the multi-camera training set and the cluster centers corresponding to each single camera, calculate the multi-view contrast loss function where the temperature coefficient \(\tau\) c \(= 0.05\). The multi-camera memory bank \(M\) p contains \(N\) single-camera memory banks, and the \(j\)-th surveillance camera corresponds to the single-camera memory bank \(M\) p,jAfter initialization corresponding to the cluster centers, at the t-th iteration, update is performed through where the momentum update coefficient λ of the single-camera memory bank p = 0.1.

[0047] 4. Finally, jointly optimize the contrastive loss functions of the three branches, and minimize the total loss value through multiple iterations during backpropagation, thereby optimizing the parameters of the feature encoder.

[0048] (1) Based on the class contrastive loss function L cluster , the single-view contrastive loss function L Intra and the multi-view contrastive loss function L Inter , calculate the total loss function L = L cluster + L Intra + L Inter ;

[0049] (2) During the backpropagation process, after the t-th iteration, optimize the parameters in the ResNet50 neural network by minimizing the total loss function L, so as to output discriminative pedestrian features and ensure the accuracy of pedestrian search.

[0050] The simulation experiment of the method of the present invention is as follows:

[0051] In this experiment, a set of video surveillance systems was constructed to collect the original surveillance videos. This system consists of 9 surveillance cameras, which collect surveillance videos in different scenarios respectively. 20,000 original surveillance video segments were selected from them. To test the performance of pedestrian search, 15,000 original surveillance video segments were used as the training data set, and the remaining 5,000 video segments were used as the test data set. At the same time, the mean Average Precision (mAP) was selected as the metric for measuring the performance of pedestrian search.

[0052] In Figure 5In this paper, the proposed unsupervised clustering method based on graph segmentation is compared with the DBSCAN method in terms of search accuracy. When changing the number of original surveillance video segments in the training dataset, the search accuracy of the unsupervised clustering method based on graph segmentation is higher than that of the DBSAC clustering method. This is because the DBSCAN clustering algorithm divides the feature space based on density and defines the density threshold depending on the neighborhood radius and the minimum number of points in the neighborhood. The selection of the two parameters directly affects the clustering effect and the confidence of pedestrian pseudo-labels, and thus affects the subsequent pedestrian search accuracy. The unsupervised clustering method based on graph segmentation divides the feature space through a hierarchical structure and merges or splits clustering categories by finding the nearest neighbors. Compared with the DBSCAN clustering method, it has less dependence on parameter selection and introduces the constraint conditions of pedestrian spatio-temporal features, so as to generate pedestrian pseudo-labels with higher confidence and improve the accuracy of pedestrian search.

[0053] In Figure 6 In this paper, when changing the number of original surveillance video segments in the training dataset, this method is compared with the collaborative contrastive refining method proposed in the literature "Collaborative Contrastive Refining for Weakly Supervised Person Search" (IEEE Transactions on Image Processing, 2023) and the context-guided feature learning method proposed in the literature "Exploring Visual Context for Weakly Supervised Person Search" (AAAI 2022). It can be seen from this that the mean average precision of this method is the highest, while the mean average precision of the context-guided feature learning method is the lowest. This is because the context-guided feature learning method focuses on pedestrian context detection, memory context, and scene context, but does not consider the spatio-temporal constraints in the original surveillance video and the correlation between surveillance cameras. In contrast, the collaborative contrastive refining method improves the discriminative ability of pedestrian features, but this method treats a single pedestrian as a noise sample, which reduces the similarity of pedestrian features in a multi-camera scenario to a certain extent. This method not only considers the discriminability of different pedestrian features in a single-camera scenario, but also involves the similarity of the same pedestrian features in a multi-camera scenario, thus enhancing the expressive ability of pedestrian features and being conducive to improving the accuracy of pedestrian search.

Claims

1. An unsupervised pedestrian search method for original surveillance videos, characterized in that, Including the following steps: Using a pre-trained neural network to detect pedestrian regions in the original surveillance video and using a feature encoder to generate pedestrian features; Regarding the clustering problem of pedestrian features with an unknown total number of categories as a graph segmentation problem, where each node represents a pedestrian feature, the connections between nodes represent the similarity between pedestrian features, and each segmented region represents a clustering category; In the original surveillance video, multiple pedestrians appearing in a single camera at the same moment belong to different identity categories. However, at different moments, the same pedestrian may appear in different cameras. Based on this, calculate the weighted total similarity based on the visual similarity and context similarity of pedestrian features, find the first-nearest neighbor relationship of pedestrian features and organize the chain structure of pedestrian features, determine the spatial distribution of pedestrian features, and achieve unsupervised clustering according to the segmented regions; Label the identity categories of pedestrians according to the clustering categories and generate pseudo-labels for pedestrians according to the clustering centers; During the forward propagation process, calculate the contrast loss functions of each branch from three branches: category discriminative learning, single-camera contrast learning, and multi-camera contrast learning, and update the clustering center memory bank and multi-camera memory bank respectively according to the momentum update mechanism; Jointly optimize the contrast loss functions of the three branches and minimize the total loss value through multiple iterations during the backpropagation process to optimize the parameters of the feature encoder.

2. The unsupervised pedestrian search method for the original surveillance video according to claim 1, characterized in that The statement of using a pre-trained neural network to detect pedestrian regions in the original surveillance video and using a feature encoder to generate pedestrian features means that the original surveillance video consists of multiple video segments and their corresponding camera labels. In each video segment, use a pre-trained neural network to detect pedestrian bounding boxes and determine pedestrian regions; Use a feature encoder to generate corresponding pedestrian features for each pedestrian region respectively.

3. The unsupervised pedestrian search method for original surveillance videos according to claim 2, characterized in that The statement of using a pre-trained neural network to detect pedestrian regions in the original surveillance video and using a feature encoder to generate pedestrian features specifically includes the following steps: 1) For the original surveillance video V = {(v i , s j )}, v i (1 ≤ i ≤ M) is the i-th video segment, M is the number of video segments, s j (1 ≤ j ≤ N) is the j-th camera tag corresponding to v i , indicating that the video segment v i is collected by the surveillance camera s j ; The video segment v i is processed by a pre-trained neural network to detect the pedestrian region r i . All pedestrian regions in the original surveillance video V are represented as R = {(r i , s j )}; 2) For the pedestrian area r i , generate the pedestrian feature d θ through the feature encoder f i = f θ (r i ). All the pedestrian features in the original surveillance video V are represented as D = {(d i , s j )} = {(f θ (r i ), s j )}.

4. The unsupervised pedestrian search method for original surveillance videos according to claim 3, characterized in that Based on the visual similarity and context similarity of pedestrian features, calculate the weighted total similarity, find the first-nearest neighbor relationship of pedestrian features and organize the chain structure of pedestrian features, determine the spatial distribution of pedestrian features, and achieve unsupervised clustering according to the segmented regions; label the identity categories of pedestrians according to the clustering categories and generate pseudo-labels for pedestrians according to the clustering centers, specifically including the following steps: 1) Among all pedestrian features D, calculate the visual similarity P(m,n) = sim(d i,m , d j,n ) = sim(f θ (r i,m ), f θ (r j,n ))), where d i,m represents the m-th pedestrian feature in the i-th video segment v i , and d j,n represents the n-th pedestrian feature in the j-th video segment v j . Calculate the context similarity of the above pedestrian features where N i is the number of pedestrian features in video segment v i and N j is the number of pedestrian features in video segment v j The total similarity of the weighted two pedestrian features is calculated as T(m,n) = λP(m,n) + (1 - λ)Q(m,n), where λ is the trade-off coefficient between the two similarities; sort the total similarity T(m,n) from largest to smallest, and find the first nearest neighbor relationship of pedestrian feature d i,m ​ 2) Construct a symmetric sparse matrix \(A(m,n,t)\in\{0,1\}\) according to the first - nearest - neighbor relationship of pedestrian features for graph segmentation. At time \(t\), the matrix where indicates that the first - nearest - neighbor of the \(m\) - th pedestrian feature is the \(n\) - th pedestrian feature, indicates that the first - nearest - neighbor of the \(n\) - th pedestrian feature is the \(m\) - th pedestrian feature, indicates that the first - nearest - neighbors of the \(m\) - th pedestrian feature and the \(n\) - th pedestrian feature are the same, \(t\) m and \(t\) n respectively represent the times when the \(m\) - th pedestrian feature and the \(n\) - th pedestrian feature appear. If \(A(m,n,t)=1\), it means there is a connection between the \(m\) - th pedestrian feature and the \(n\) - th pedestrian feature; Based on the chain structure between pedestrian features, construct the connected components between pedestrian features; Determine the segmented regions of pedestrian features through these connected components, and each segmented region represents a clustering category; 3) After determining the clustering categories among all pedestrian features D, label the identity categories of pedestrians through the clustering categories, where the cluster centers are represented as C = {c i}(1 ≤ i ≤ K), and generate pseudo-labels for pedestrians based on the cluster centers.

5. The unsupervised pedestrian search method for original surveillance videos according to claim 1, characterized in that, The statement of during the forward propagation process, calculate the contrast loss functions of each branch from three branches: category discriminative learning, single-camera contrast learning, and multi-camera contrast learning, and update the clustering center memory bank and multi-camera memory bank respectively according to the momentum update mechanism means that in the category discriminative learning branch, calculate the category contrast loss function to ensure the discriminability between different pedestrian features; In the single-camera contrast learning branch, calculate the single-view contrast loss function to distinguish the discriminability of different pedestrian features under the same view; In the multi-camera contrast learning branch, calculate the multi-view contrast loss function to improve the similarity of the same pedestrian features under different views; after calculating the loss functions of the above three branches, update the clustering center memory bank and the multi-camera memory bank respectively according to the momentum update mechanism.

6. The unsupervised pedestrian search method for original surveillance videos according to claim 1, characterized in that The steps of calculating the contrast loss functions of each branch respectively from the three branches of class discriminative learning, single-camera contrast learning and multi-camera contrast learning and updating the clustering center memory bank and the multi-camera memory bank respectively according to the momentum update mechanism in the forward propagation process include: 1) In the class discriminative learning branch, calculate the class contrast loss function where q represents the query pedestrian feature in the training set, c + represents the cluster center corresponding to the query pedestrian feature q, represents the i-th cluster center in the t-th iteration, τ c represents the temperature coefficient; according to the momentum update mechanism, update the cluster center memory bank M c , and it is calculated as where q hard represents the query pedestrian feature in the training set with the minimum visual similarity to its cluster center , λ c represents the momentum update coefficient of the cluster center memory bank; 2) In the single-camera contrastive learning branch, calculate the single-view contrastive loss function where p + represents the relevant cluster center corresponding to q′ hard and q′ hard represents the query pedestrian feature with the minimum visual similarity to p + in the single-camera training set, represents an irrelevant cluster center relatively similar to the query pedestrian feature q, K - represents the number of irrelevant cluster centers, and τ c represents the temperature coefficient; 3) In the multi-camera contrastive learning branch, calculate the multi-view contrastive loss function where represents the i-th cluster center with camera label s in the t-th iteration j and represents the relevant cluster center with a different camera label corresponding to q, where q easy represents the query pedestrian feature with the maximum visual similarity to easy in the multi-camera training set, and τ represents the temperature coefficient; the multi-camera memory bank M c is composed of N single-camera memory banks, and the j-th surveillance camera corresponds to the single-camera memory bank M p ; according to the momentum update mechanism, update the single-camera memory bank M p,j and it is calculated as p,j where λ represents the momentum update coefficient of the single-camera memory bank p ​ 7. The unsupervised pedestrian search method for original surveillance videos according to claim 1, wherein The joint optimization of the contrast loss functions of the three branches and the minimization of the total loss value through multiple iterations in the backpropagation process to optimize the parameters of the feature encoder means: summing up the contrast loss functions of the three branches of class discriminative learning, single-camera contrast learning and multi-camera contrast learning, and solving the minimum value of the total loss function through multiple iterations in the backpropagation process to achieve the purpose of optimizing the parameters of the feature encoder.

8. The unsupervised pedestrian search method for original surveillance videos according to claim 1, characterized in that The steps of jointly optimizing the contrast loss functions of the three branches and minimizing the total loss value through multiple iterations in the backpropagation process to optimize the parameters of the feature encoder include: 1) For the class contrast loss function L cluster , the single-view contrast loss function L Intra and the multi-view contrast loss function L Inter , the total loss function L is calculated as L = L cluster + L Intra + L Inter ; 2) During the backpropagation process, the minimum value of the total loss function L is solved after t iterations of training to optimize the feature encoder f θ for the parameters θ in (·), thereby outputting discriminative pedestrian features.

Citation Information

Patent Citations

  • Pedestrian search method based on pedestrian re-identification driving positioning adjustment

    CN110826424A

  • An unsupervised person re-identification method based on camera distribution difference alignment constraint

    CN113065409B

  • Cross-camera pedestrian re-identification method with supervision in camera based on comparative learning

    CN112784772A

  • Pedestrian re-identification method and system based on momentum network and comparative learning

    CN114724075A

  • Unsupervised pedestrian re-identification method based on Multiform and outlier sample re-distribution

    CN115601791A