Semi-supervised pedestrian re-identification method based on multistage clustering

Through the combination of multi-level clustering and the teacher-student mutual learning framework, the problem of difficulty in obtaining tags in pedestrian re-identification is solved, and the efficient use of unsupervised data is achieved, and the robustness and recognition performance of the model are improved.

CN120388391APending Publication Date: 2025-07-29SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510270374.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing pedestrian re-identification method is difficult to obtain labels in large-scale data scenarios, and manual labeling is expensive, and when using unsupervised data, it is easy to degrade model performance due to noise interference, making it difficult to maintain identification effect in multiple scenarios.

Method used

The multi-level clustering strategy is used to process unsupervised data, and pseudo-labels are generated through single-track, single-video and cross-video clustering, and distillation training of the teacher-student mutual learning framework is combined with supervised data to improve the robustness and recognition performance of the model.

Benefits of technology

Reliance on manual annotation is reduced, the model's robustness and recognition performance in diverse scenarios is improved, the utilization value of unsupervised data is enhanced, and the generalization ability of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388391A_ABST
    Figure CN120388391A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised pedestrian re-identification method based on multistage clustering. The method comprises the following steps: acquiring video data; and inputting the video data into a pedestrian re-identification model to obtain a pedestrian identification result. Wherein the pedestrian re-identification model is obtained according to the following steps: training a feature extraction network by using a labeled training set; for unlabeled unsupervised data, extracting corresponding output features by using the trained feature extraction network; performing single-track clustering, single-video clustering and cross-video clustering on the output features to obtain label information, and further constructing an automatic labeling pedestrian training set; carrying out distillation training on a teacher-student mutual learning framework based on the automatic marking pedestrian training set and supervised data, wherein the teacher-student mutual learning framework comprises a teacher model and a student model; and taking the student model subjected to distillation training as the pedestrian re-identification model. According to the invention, the data annotation cost is reduced, and the generalization ability of the pedestrian recognition model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more particularly, to a semi-supervised pedestrian re-identification method based on multi-level clustering. Background Art

[0002] Pedestrian re-identification is a technology that uses computer vision technology to determine whether a specific pedestrian exists in an image or video sequence. In order to efficiently utilize limited supervised data (i.e., data manually labeled), pedestrian re-identification methods often focus on improving the network depth and complexity. This strategy will significantly increase the computational and storage overhead, and when the model is deployed to different application scenarios, the recognition effect will significantly decline. For each scenario, it is necessary to first collect sufficient data and then train a new domain adaptation model each time, thus restricting the wide application of the pedestrian re-identification model.

[0003] To improve the adaptability of the model in multiple scenarios, it is necessary to label unsupervised data. However, manual labeling is not only time-consuming and laborious, but also costly, making it difficult to promote on a large scale. Therefore, how to generate appropriate labels for unsupervised data (i.e., data not manually labeled) to further promote the training of pedestrian re-identification models based on large-scale data has become an important research topic with application potential.

[0004] Existing unsupervised data annotation schemes are mainly divided into two types of methods. When the first type of method processes unsupervised data, positive samples are generated by data augmentation for the same image, and the remaining images are regarded as negative samples. This type of method does not consider the identity invariance of the same pedestrian in different images, that is, it ignores the identity correlation between different images, because it is difficult to make full use of more unsupervised data to further improve the representation ability of the model.

[0005] The second type of method uses a tracking algorithm to process a single video, and then regards all images included in each trajectory in the single video as positive samples, so that multiple images can be associated according to the tracking ID, and the remaining different trajectories and images across videos are regarded as negative samples. However, this tracking algorithm is difficult to achieve complete accuracy, and there are still some images in the trajectory that do not belong to the same pedestrian. This part of the error will have a negative impact on the incremental training of the model, and it also does not focus on the situation where the same pedestrian appears in different videos. Therefore, this type of method is limited to the identity association of a single video and is difficult to establish the identity association of images across videos.

[0006] In summary, the current pedestrian re-identification task faces the problem of difficult label acquisition, especially in the scenario of large-scale data, where manual labeling is costly and time-consuming. And when the existing methods use unsupervised data, the model performance is easily degraded due to noise interference, thus affecting the recognition effect. Summary of the Invention

[0007] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide a semi-supervised pedestrian re-identification method based on multi-level clustering. The method includes the following steps:

[0008] Obtain video data;

[0009] Input the video data into a pedestrian re-identification model to obtain a pedestrian recognition result;

[0010] Among them, the pedestrian re-identification model is obtained according to the following steps:

[0011] Fully supervise the training of the feature extraction network using the labeled training set;

[0012] For the unlabeled unsupervised data, use the trained feature extraction network to extract the corresponding output features;

[0013] For the output features, obtain label information through single-trajectory clustering, single-video clustering, and cross-video clustering, and then construct an automatically labeled pedestrian training set;

[0014] Taking a set loss criterion as the optimization objective, perform distillation training on the teacher-student mutual learning framework based on the automatically labeled pedestrian training set and the supervised data. The teacher-student mutual learning framework includes a teacher model and a student model;

[0015] Use the student model after distillation training as the pedestrian re-identification model.

[0016] Compared with the prior art, the advantages of the present invention are that the provided semi-supervised pedestrian re-identification method based on multi-level clustering processes unsupervised data by constructing a multi-level clustering strategy of single-trajectory, single-video, and cross-video, enhances the diversity of images with the same identity, and then effectively uses supervised and unsupervised data to perform distillation training on the pedestrian re-identification model based on the teacher-student mutual learning framework. By combining supervised and unsupervised data for training, the potential value of unsupervised data is fully exploited, thereby improving the robustness and recognition performance of the model. The present invention not only reduces the dependence on manual annotation but also significantly improves the robustness and recognition performance of the model in diverse scenarios, providing an efficient and feasible solution path for semi-supervised data-driven pedestrian re-identification.

[0017] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. Description of the Drawings

[0018] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.

[0019] Figure 1 is a flowchart of a semi-supervised pedestrian re-identification method based on multi-level clustering according to an embodiment of the present invention;

[0020] Figure 2 is a schematic diagram of the process of a semi-supervised pedestrian re-identification method based on multi-level clustering according to an embodiment of the present invention;

[0021] Figure 3 is a flowchart of performing basic clustering according to an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of a teacher-student mutual learning framework according to an embodiment of the present invention. Detailed implementation manners

[0023] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.

[0024] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention, its application, or its use.

[0025] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and devices should be regarded as part of the specification.

[0026] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0027] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0028] Generally speaking, the present invention processes unsupervised data through single-trajectory, single-video, and cross-video multi-level clustering strategies, which is conducive to automatically obtaining diverse image subsets of the same identity, and then combining supervised data with unsupervised data, and further improving the model generalization ability through distillation training based on a teacher-student mutual learning framework.

[0029] Combined with Figure 1 and Figure 2 as shown, the provided semi-supervised pedestrian re-identification method based on multi-level clustering includes the following steps:

[0030] Step S1, training a feature extraction network using supervised data to obtain a trained pedestrian re-identification feature extraction network.

[0031] For example, a supervised data is used to train the feature extraction network, and the optimization loss is the cross-entropy loss function and the triplet loss function, obtaining a trained feature extraction network (or a fully supervised person re-identification model). The training set used for the full supervision training of the feature extraction network can be a public manually annotated training set or a self-annotated training set.

[0032] The feature extraction network can adopt various types of neural network structures, such as Swin Transformer, Vision Transformer, or ResNet50, etc.

[0033] In one embodiment, the cross-entropy loss function is expressed as:

[0034]

[0035] where N represents the number of images in the training batch, m represents the number of categories in the training set, the subscript i represents the current batch image index, the subscript j represents the person category, represents the parameters of the identity classification layer, and f represents the pedestrian feature.

[0036] In one embodiment, the triplet loss works by inputting a triplet composed of an anchor sample a, a positive sample p, and a negative sample n, where p and a belong to the same category, while n and a are of different categories. After the model extracts features from the input samples, positive sample feature pairs {f a , f p} and negative sample feature pairs {f a , f n} are obtained. The goal of the triplet loss is to optimize the model so that the distance of the positive sample feature pair is less than that of the negative sample feature pair and maintain a set minimum margin, thereby achieving the aggregation of features within the category and the separation across categories. For example, the triplet loss function is expressed as:

[0037]

[0038] where ρ1 is the set margin value, y represents the category corresponding to the input sample. For example, y a is the category corresponding to the input sample a, y p is the category corresponding to the input sample p, and y n is the category corresponding to the input sample n.

[0039] The overall loss function used for the full supervision training of the feature extraction network can be the weighted sum of the cross-entropy loss function and the triplet loss function. The weighting coefficient can be determined according to actual needs or simulations. For example, the weighting coefficient is set to 1.

[0040] Step S2. For unsupervised data, use the trained feature extraction network to extract features and obtain unsupervised data features.

[0041] Use the trained feature extraction network to perform feature extraction on the unsupervised data to output the features of the unsupervised data.

[0042] For example, extract unlabeled video frames from video streams from multiple cameras, initialize them using a pedestrian detection and tracking algorithm, obtain a pedestrian trajectory dataset, and input it as unsupervised data into the trained feature extraction network to output unsupervised data features.

[0043] Step S3. Perform multi-level clustering on the unsupervised data features to obtain label information, and then construct it into an automatic annotation training set. The multi-level clustering includes single-trajectory clustering, single-video clustering, and cross-video clustering.

[0044] To understand multi-level clustering, first introduce the Base Cluster process. See Figure 3 As shown, the Base Cluster includes the following steps:

[0045] S11. Set the distance threshold parameter K and prepare the feature set.

[0046] S12. Treat each feature as a cluster, calculate the cosine distance between each pair according to the following formula. Assuming there are N points, an N×N distance matrix can be obtained.

[0047]

[0048] where Cosine Distance represents the cosine distance, X·Y is the dot product of two feature vectors X and Y, and ‖X‖ and ‖Y‖ are the Euclidean norms of vectors X and Y respectively.

[0049] S13. Find the minimum value in the distance matrix, and cluster the two clusters (single feature or class) corresponding to this value into the same class, that is, obtain a new cluster.

[0050] S14. Recalculate the distance between the new cluster and all clusters.

[0051] S15. Repeat steps S13 and S14 until the minimum value in the distance matrix is greater than the distance threshold K and then stop.

[0052] It should be noted that when performing the above iterative process, the distance between the two farthest features between clusters can be used as the inter-cluster distance.

[0053] According to the above basic clustering process, multi-level clustering such as single-trajectory clustering, single-video clustering, and cross-video clustering can be sequentially performed on the unsupervised data features to obtain corresponding pseudo-labels.

[0054] In one embodiment, single-trajectory clustering includes the following steps:

[0055] S21, Use a tracking algorithm to obtain tracking trajectories from a single video segment, and each trajectory serves as a feature subset.

[0056] Existing technologies such as DeepSort and ByteTrack can be used for the tracking algorithm.

[0057] S22, Apply the basic clustering to each feature subset separately to ensure the identity consistency of the features included in the feature subset, and temporarily store the clusters with only a single feature filtered out in the single-trajectory filtered feature set E1.

[0058] In one embodiment, single-video clustering includes the following steps:

[0059] S31, Calculate the central feature of the features included in each feature subset according to the following formula:

[0060]

[0061] where c is the central feature vector, x i is the feature vector of the i-th sample, and N is the number of samples in the feature subset.

[0062] S32, Combine these central features and the set E1 to obtain the single-video feature set s.

[0063] S33, After applying the basic clustering to each video feature set s i ∈{s1, s2, s3,..., s n} separately, discard the clusters with only a single feature in each video feature set to obtain a new video feature set k i ∈{k1, k2, k3,..., k n}.

[0064] Single-trajectory clustering and single-video clustering mainly achieve the reallocation of abnormal samples in a single trajectory and the merging of similar trajectories in a single video, reducing the impact caused by errors generated by the tracking algorithm, thereby ensuring the identity consistency and identity discriminability of the pedestrian feature subsets in each video.

[0065] In one embodiment, cross-video clustering includes the following steps:

[0066] S41, Combine the feature sets of all videos into an overall feature set K.

[0067] S42. Construct a sliding window W with a custom length to process the feature set K, apply basic clustering to the range covered by W, and assign link relationships to the features f i and f j (f i , f j ∈K).

[0068] S43. After the sliding window W traverses the entire feature set, merge the features with link relationships, thereby completing the identity association and differentiation of pedestrian images across videos.

[0069] Cross-video clustering establishes identity associations for different trajectories of the same pedestrian in multiple videos, ensuring instance diversity of the same identity.

[0070] In summary, pseudo-labels of unsupervised data can be obtained through multi-level clustering, and then an automatically annotated training set can be constructed.

[0071] Step S4. Based on the automatically annotated training set combined with supervised data, perform distillation training on the teacher-student mutual learning framework until the set distillation loss criterion is met.

[0072] Utilize the automatically annotated training set combined with supervised data (such as a publicly available manually annotated training set) to further improve the generalization ability of the feature extraction network and improve the recognition efficiency by performing knowledge distillation training on the teacher-student interactive learning framework.

[0073] The main idea of knowledge distillation is to construct a teacher-student mutual learning framework. The student model is supervised by the teacher model. By minimizing the distillation loss function, the class prediction probability distribution of the student model gradually approaches that of the teacher model, thereby transferring knowledge from the teacher model to a smaller and more efficient student model.

[0074] For example, use the parameters of the trained feature extraction network to initialize the student model and the teacher model, construct a teacher-student mutual learning framework. The supervised data is optimized using the cross-entropy loss function, the triplet loss function, and the distillation loss, while the unsupervised data is optimized using the distillation loss function based on automatically generated labels. Finally, use the student model obtained through distillation training as a pedestrian re-identification model for target pedestrian recognition in actual scenarios, also known as a semi-supervised pedestrian re-identification model.

[0075] In one embodiment, the teacher-student mutual learning framework mainly includes a teacher model and a student model. The teacher model contains an identity prediction module, while the student model contains two identity prediction modules. Since the teacher model is a feature extraction network trained on supervised data in step S1 and has strong identity recognition ability, one identity prediction module is set for the classification task, and the goal is to provide stable features to guide the student model to learn. In addition, considering that the supervised data has real labels while the automatically labeled labels of the unsupervised data may be noisy, the student model is equipped with two independent identity prediction modules, which can reduce the negative impact of the automatically labeled labels on the distillation training. By learning the distributions of the real labels and the automatically labeled labels respectively, the student model can better adapt to the automatically labeled data, thus performing more stably in different scenarios and further improving the generalization ability of the student model.

[0076] The output dimensions of the above three identity prediction modules are the same. In each training iteration, the parameters of the student model are updated by the conventional gradient descent method, while the parameters of the teacher model are updated according to the following momentum formula. This means that the parameter update of the teacher model is a smoothed average based on the parameters of the student model. Specifically, the parameters of the teacher model are updated by the following formula:

[0077] θ teacher = λ·θ teacher +(1 - λ)·θ student (5)

[0078] where the λ coefficient controls the update speed of the parameters of the teacher model, θ teacher is the parameter of the teacher model, and θ student is the parameter of the student model. This momentum update method makes the parameter change of the teacher model smoother and avoids the drastic fluctuations of the parameters of the teacher model caused by the unstable training of the student model.

[0079] In one embodiment, the calculation formula of the distillation loss is as follows:

[0080]

[0081]

[0082] where d represents the output of the identity prediction module (the calculation methods of the three identity prediction modules are the same, and the subscript k represents the k-th element of the feature vector), τ represents the hyperparameter controlling the sharpness of the identity prediction distribution, N represents the number of samples used for training, K represents the total number of output dimensions, p(t * ) corresponds to the output of the teacher model, p(s * ) corresponds to the output of the student model, and p(*) is a unified symbol representing the output of the identity prediction module. For example, p(ti,j ) represents the output of the teacher model, p(s i,j ) represents the output of the identity prediction module corresponding to the supervised data in the student model, i is the training sample index, and j is the output dimension index.

[0083] After the unsupervised data undergoes multi-level clustering, the pedestrian images have corresponding labels, so that the labels automatically generated by multi-level clustering can be introduced into the calculation of the distillation loss. The unsupervised data distillation loss based on the multi-level clustering labels is expressed as:

[0084]

[0085] Among them, T 2 is used to compensate for the influence of gradient scaling during backpropagation to maintain the numerical stability of the loss. V represents the number of automatically generated labels, B represents the number of samples with consistent multi-level clustering labels in the current training batch, and K represents the total number of output dimensions. represents the output of the teacher model, represents the output of the identity prediction module corresponding to the automatically labeled data in the student model, m is the index of the automatically generated label, h is the training batch index, and j is the output dimension index.

[0086] In the distillation training, minimizing the overall distillation loss function is used as the optimization objective to optimize the parameters of the student model. For example, the overall distillation loss function is the distillation loss and the unsupervised data distillation loss based on the multi-level clustering labels The weighted sum of, and the weighting coefficient can be determined according to actual needs or simulations. For example, the weighting coefficient is set to 1.

[0087] Through multi-level clustering, each pedestrian image in the unsupervised data is assigned a relatively accurate label. Although these labels are automatically generated, they still contain the structural similarities between the images. Introducing these clustering labels into the distillation loss allows the student model to simultaneously learn the differences between classes and the details within the same class during the training process, thereby improving the learning effect and generalization ability of the model.

[0088] In summary, during the distillation process, the teacher model can usually provide certain supervision signals, and using the labels generated by clustering can further strengthen the student model's understanding and discrimination of different classes, simulating the label information in supervised learning, but without the need for a large amount of manually labeled data.

[0089] Step S5, using the student model obtained through distillation training as a semi-supervised pedestrian re-identification model for actual pedestrian identification.

[0090] After the above training, the optimized student model is used as a semi-supervised pedestrian re-identification model for actual target pedestrian recognition. For example, the model application process is as follows: Obtain video data of a predetermined scenario; for the video data, calculate the similarity based on the pedestrian features extracted by the semi-supervised pedestrian re-identification model to obtain the pedestrian re-identification result.

[0091] It should be noted that the training process of the model or network of the present invention can be carried out offline on a server or in the cloud. Embedding the trained model into an electronic device can achieve real-time target pedestrian recognition. The electronic device can be a terminal device or a server. The terminal device includes any terminal device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a point-of-sale (POS) terminal, an in-vehicle computer, a smart wearable device (such as a smart watch, virtual reality glasses, virtual reality helmets, etc.). The server includes, but is not limited to, an application server or a web server, and can be an independent server, a cluster server, or a cloud server, etc.

[0092] To further verify the effect of the present invention, experimental verification was carried out. First, full-supervised training was performed using the publicly labeled datasets Market-1501, DukeMTMC, CUHK03, and MSMT17. Then, semi-supervised training based on the teacher-student mutual learning framework was performed using the publicly labeled datasets and the collected unsupervised data. Finally, inference testing was carried out on the publicly available ENTIReID pedestrian re-identification test dataset, which contains 2,730 query IDs and a total of 13,415 images. The experiment uses two evaluation criteria, Rank-1 and mean average precision (mAP), to measure the performance of the model. Both metrics are numbers between 0 and 1, and the larger the value, the higher the accuracy of pedestrian re-identification. The experimental results are shown in Table 1.

[0093] Table 1: Experimental results on the ENTIReID dataset (%)

[0094]

[0095] As can be seen from Table 1, the teacher-student mutual learning framework can use unsupervised data to improve the performance of the model. After performing multi-level clustering processing on the unsupervised data, with the same amount of data, the teacher-student mutual learning framework brings greater performance improvement to the model.

[0096] In summary, the present invention realizes the automatic annotation of unsupervised data through a multi-level clustering method of single trajectory, single video, and cross-video. On the premise of ensuring the accuracy of automatically generated labels, the instance diversity of the same labels is improved. In addition, by combining supervised and automatically annotated data, the generalization ability of the pedestrian re-identification model is further improved using the teacher-student mutual learning framework, and a smaller and more efficient student model is used for pedestrian recognition, improving the recognition efficiency.

[0097] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.

[0098] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0099] The computer-readable program instructions described herein may be downloaded to each computing / processing device from a computer-readable storage medium or may be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0100] The computer program instructions for carrying out the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages - such as Smalltalk, C++, Python, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0101] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0102] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0103] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As will be apparent to those of ordinary skill in the art, implementations in hardware, in software, and in a combination of software and hardware are all equivalent.

[0105] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A semi-supervised pedestrian re-identification method based on multi-level clustering, comprising the following steps: Obtain video data; Input the video data into a pedestrian re-identification model to obtain pedestrian recognition results; Wherein, the pedestrian re-identification model is obtained according to the following steps: Perform full-supervised training on a feature extraction network using a labeled training set; For unlabeled unsupervised data, use the trained feature extraction network to extract corresponding output features; For the output features, obtain label information through single-trajectory clustering, single-video clustering, and cross-video clustering, and then construct an automatically labeled pedestrian training set; Using a set loss criterion as the optimization objective, perform distillation training on a teacher-student mutual learning framework based on the automatically labeled pedestrian training set and supervised data, and the teacher-student mutual learning framework includes a teacher model and a student model; Use the student model after distillation training as the pedestrian re-identification model.

2. The method according to claim 1, characterized in that, The teacher model and the learning model are constructed based on the feature extraction network, and the training set used for distillation training of the teacher-student mutual learning framework includes the labeled training set and the automatically labeled pedestrian training set.

3. The method according to claim 1, wherein The overall loss function for training the feature extraction network is the cross-entropy loss function and the triplet loss function which are weighted sums, respectively set as: Among them, N represents the number of images in a training batch, m represents the number of categories in the training set, the subscript i represents the current batch image index, the subscript j represents the pedestrian category, v represents the parameters of the identity classification layer, f represents the pedestrian feature, a is the anchor sample, p is the positive sample, n is the negative sample, y represents the category corresponding to the input sample, ρ1 is the set margin value, p and a belong to the same category, n and a are different categories, {f a , f p} is the positive sample feature pair, {f a , f n} is the negative sample feature pair.

4. The method according to claim 1, wherein The teacher model includes an identity prediction module, and the student model includes two identity prediction modules, which are respectively used to process the automatically labeled pedestrian training set and the supervised data.

5. The method according to claim 4, wherein The overall loss function used for distillation training of the teacher-student mutual learning framework includes the distillation loss function of supervised data and the distillation loss function of unsupervised data which are respectively set as: Among them, T 2 is the influence factor used to compensate for the gradient scaling during backpropagation, V represents the number of automatically generated labels, B represents the number of samples with consistent multi-level clustering labels in the current training batch, K represents the total number of output dimensions, τ represents the hyperparameter controlling the sharpness of the identity prediction distribution, N represents the number of samples used for training, K represents the total number of output dimensions, represents the output of the teacher model, represents the output of the identity prediction module corresponding to the automatically labeled pedestrian training set in the student model, j is the output dimension index, p(t i,j ) represents the output of the teacher model, p(s i,j ) represents the output of the identity prediction module corresponding to the supervised data in the student model, i is the training sample index, j is the output dimension index, h is the training batch index, and m is the index of the automatically generated label.

6. The method according to claim 1, characterized in that, The single-trajectory clustering includes: Obtain tracking trajectories from a single video segment, and regard each trajectory as a feature subset; Perform basic clustering on each of the feature subsets separately to ensure the identity consistency of the features included in the feature subset, and store the clusters filtered out with only a single feature in a single-trajectory filtered feature set E1; The single-video clustering includes: Calculate the central features of the features included in each feature subset according to the following formula: where c is the central feature vector, x i is the feature vector of the i-th sample, and N is the number of samples in the feature subset; Combine these central features and the set E1 to obtain a single-video feature set s; For each video feature set s i ∈ {s1, s2, s3, …, s n}, after separately applying the described basic clustering, discard the clusters with only a single feature in each video feature set, obtaining a new video feature set k i ∈ {k1, k2, k3, …, k n}; The cross-video clustering includes the following steps: Combine the feature sets of all videos into an overall feature set K; Construct a sliding window W with a custom length to process the overall feature set K, apply the basic clustering to the range covered by W, and assign a linking relationship to the features f i and f j where f i , f j ∈K; When the sliding window W traverses the overall feature set, merge the features with a link relationship, and then complete the identity association and differentiation of cross-video pedestrian images.

7. The method according to claim 6, characterized in that, The basic clustering includes: Set a distance threshold parameter and prepare a feature set; Regard each feature as a cluster, and calculate the cosine distance between each pair according to the following formula: Where, Cosine Distance represents the cosine distance, X·Y is the dot product of two feature vectors X and Y, ‖X‖ and ‖Y‖ are the Euclidean norms of vectors X and Y respectively, and for N points, an N×N distance matrix is obtained; Find the minimum value in the distance matrix, and cluster the two clusters corresponding to this value into the same class to obtain a new cluster; Recalculate the distances between the new cluster and all clusters until the minimum value in the distance matrix is greater than the distance threshold and then stop.

8. The method according to claim 1, wherein During the distillation training of the teacher-student mutual learning framework, the parameters of the teacher model are updated according to the following formula: θ teacher = λ·θ teacher + (1 - λ)·θ student where λ is a coefficient that controls the update speed of the teacher model's parameters, and θ teacher are the parameters of the teacher model, and θ student are the parameters of the student model.

9. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.