Pedestrian re-identification method, system, medium and device based on top corner shooting
The pedestrian re-identification method based on top-angle shooting uses a vertical top-view camera to acquire pedestrian trajectories and features. Combined with neural networks and clustering algorithms, it solves the problems of occlusion and unidirectional images in pedestrian re-identification, and improves recognition accuracy and robustness.
Patent Information
- Application Number
- CN202310150855.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-02-21
AI Technical Summary
In existing pedestrian re-identification technologies, the inability to accurately determine pedestrian trajectories is due to situations where the probe is obstructed when capturing images of the target and can only capture images of a single direction of the human body.
A top-angle-based shooting method is adopted to acquire pedestrian trajectories through a vertical top-view camera, collect pedestrian images from multiple angles, extract pedestrian features using neural network training, optimize the model using cross-entropy loss function and ternary pair loss function, and obtain pedestrian re-identification results by combining hierarchical density clustering method.
It improves the accuracy and robustness of pedestrian re-identification, alleviates the identification problems caused by occlusion, can obtain pedestrian features from multiple directions, and enhances the adaptability to complex environments.
Smart Images

Figure CN116189236B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of image processing and computer vision in the application of artificial intelligence, and relates to a pedestrian re-identification method, in particular to a pedestrian re-identification method based on top-angle shooting, a system, a medium and an apparatus. BACKGROUND
[0002] In recent years, pedestrian re-identification (Person Re-Identification, also known as ReID) has gradually gained attention in various scenarios in digital cities. The pedestrian re-identification model aims to associate pictures of the same pedestrian captured under different probes. Currently, the common ReID scheme is based on oblique probe, but due to the complexity of the production environment, the installation of the device is often limited, and the human body image captured at the device end has problems. The specific problems are as follows:
[0003] (1) There are various occlusion phenomena in the process of target capture by the probe, and it is difficult to appear without occlusion in the case of dense pedestrians;
[0004] (2) Generally, only a single direction of the human body (such as the front, back, or side) can be captured, and the information amount of the pedestrian trajectory may not be able to support the ReID algorithm to make correct judgments.
[0005] Therefore, in the current pedestrian re-identification technology, due to the occlusion of the target collected by the probe, and the problem that the trajectory of the pedestrian cannot be accurately judged because only a single direction of the human body image can be obtained when the probe is captured. SUMMARY
[0006] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a pedestrian re-identification method based on top-angle shooting, a system, a medium and an apparatus, which is used to solve the problem that in the process of implementing the pedestrian re-identification technology in the prior art, the trajectory of the pedestrian cannot be accurately judged because of the occlusion of the target collected by the probe, and only a single direction of the human body image can be obtained when the probe is captured.
[0007] To achieve the above-mentioned purposes and other related purposes, in a first aspect, the present application provides a pedestrian re-identification method based on top-angle shooting, comprising the following steps: acquiring a pedestrian trajectory of a pedestrian under a vertical top-view camera, and collecting a pedestrian image based on the pedestrian trajectory; acquiring a pedestrian direction under each pedestrian trajectory based on the pedestrian image; acquiring a pedestrian feature under each pedestrian trajectory based on the pedestrian image; and acquiring a pedestrian re-identification result based on the pedestrian feature and the pedestrian direction under each pedestrian trajectory.
[0008] In the present application, the pedestrian trajectory under the vertical top-view camera is first acquired, and the pedestrian image is collected based on the pedestrian trajectory; and the obtained pedestrian trajectory is predicted according to the pedestrian trajectory information, the snapshot information, and the corresponding position in the trajectory to predict the direction of the pedestrian in the snapshot; at the same time, the neural network is trained based on the collected pedestrian image to extract the pedestrian features under each pedestrian trajectory; the predicted pedestrian direction is combined with the pedestrian features to finally obtain the pedestrian re-identification result and obtain all the trajectories of the pedestrian in the area. The present application has a wide coverage angle and less occlusion, solves the problem of occlusion when the target is collected by the probe, and solves the problem that the image of only a single direction of the human body can be obtained when the probe takes a snapshot, thereby enabling accurate identification of the pedestrian identity.
[0009] In an implementation form of the first aspect, acquiring the pedestrian trajectory under the vertical top-view camera and collecting the pedestrian image based on the pedestrian trajectory comprises the following steps: acquiring the pedestrian video collected by the vertical top-view camera; acquiring the pedestrian trajectory in the pedestrian video based on a tracking algorithm; and collecting the pedestrian image based on the pedestrian trajectory.
[0010] In the present implementation, the pedestrian video is first collected by the vertical top-view camera; then the pedestrian trajectory in the pedestrian video is acquired by using a tracking algorithm; and then the pedestrian image is extracted according to the pedestrian trajectory. The pedestrian trajectory is used to predict the direction of the pedestrian in the snapshot.
[0011] In an implementation form of the first aspect, acquiring the pedestrian direction under each pedestrian trajectory based on the pedestrian image comprises the following steps: taking the projection region centered on the vertical top-view camera as a preset region, acquiring the coordinates of each point on the pedestrian trajectory in the preset region; and predicting the pedestrian direction in the pedestrian image based on the pedestrian image and the coordinates of each point on the pedestrian trajectory.
[0012] In the present implementation, the projection region centered on the vertical top-view camera is taken as the preset region, and the coordinates of the positions of each point on the pedestrian trajectory are given; and the direction of the pedestrian at each coordinate position is predicted according to the coordinate position and the shooting angle.
[0013] In an implementation form of the first aspect, acquiring the pedestrian feature under each pedestrian trajectory based on the pedestrian image comprises the following steps: pre-processing the pedestrian image under each pedestrian trajectory; and inputting the pre-processed pedestrian image into a trained convolutional neural network for pedestrian feature recognition to acquire the pedestrian feature of the pedestrian image.
[0014] In the present implementation, the collected pedestrian image is pre-processed to achieve a uniform size; and the pedestrian image of the uniform size is subjected to model training to achieve the purpose of pedestrian feature extraction.
[0015] In an implementation form of the first aspect, the loss function adopted by the convolutional neural network model comprises a cross-entropy loss function, a triplet loss function and a human direction specialized cross-entropy loss function.
[0016] In the implementation form, in the process of extracting the pedestrian features, the cross-entropy loss function and the triplet loss function are adopted to optimize the training model and achieve the purpose of recognizing pedestrian features across directions; and then the human direction specialized cross-entropy loss function is used to distinguish different directions of pedestrians on the basis of pedestrian classification, so as to improve the perception ability of the network model to the anisotropy of pedestrians.
[0017] In an implementation form of the first aspect, the obtaining of the pedestrian re-identification result based on the pedestrian features and the pedestrian directions under the respective pedestrian trajectories comprises the following steps: constructing a same-direction adjacency matrix based on the pedestrian features and the pedestrian directions under the respective pedestrian trajectories; clustering the pedestrian trajectories and the pedestrian directions in the same-direction adjacency matrix to obtain clustered clusters; constructing a different-direction adjacency matrix based on the clustered clusters; and clustering the clusters in the different-direction adjacency matrix to obtain a final pedestrian re-identification result.
[0018] In the implementation form, on the basis of completing feature reasoning, the pedestrian features of different directions on different pedestrian trajectories are clustered by a hierarchical combined density method to obtain a final result of pedestrian matching.
[0019] In an implementation form of the first aspect, the formula for constructing the same-direction adjacency matrix is:
[0020]
[0021] wherein, A represents an adjacency matrix of pedestrian i and pedestrian j; xi and xj represent pedestrian features of pedestrian i and pedestrian j respectively; d represents a direction of a pedestrian trajectory, and belongs to an intersection of a pedestrian trajectory i direction and a pedestrian trajectory j direction; n represents a number of pedestrian directions d; represents pedestrian features of the pedestrian trajectory i in the d direction, represents pedestrian features of the pedestrian trajectory j in the d direction.
[0022] In an implementation form of the first aspect, the formula for constructing the different-direction adjacency matrix is:
[0023]
[0024] wherein, c i and c j both represent feature sets of clustered cluster i and cluster j; n represents; d i represents a direction set of the cluster i; k represents a kth direction of the cluster i; d jThe direction set of cluster j is represented; l represents the lth direction of cluster j; n represents the number of different combinations of directions in all direction pairs of cluster i and cluster j.
[0025] In a second aspect, the application provides a pedestrian re-identification system based on top angle shooting, comprising: a collection processing module, which acquires pedestrian trajectories of pedestrians under a vertical top view camera and collects pedestrian images based on the pedestrian trajectories; a direction reasoning module, which is used to acquire pedestrian directions under each pedestrian trajectory based on the pedestrian images; a feature reasoning module, which is used to acquire pedestrian features under each pedestrian trajectory based on the pedestrian images; and a feature clustering module, which is used to acquire a pedestrian re-identification result based on the pedestrian features and the pedestrian directions under each pedestrian trajectory.
[0026] In the application, pedestrian trajectories of pedestrians under a vertical top view camera are acquired, and pedestrian images are collected based on the pedestrian trajectories; the direction of a pedestrian in a snapshot is predicted by a direction reasoning module according to pedestrian trajectory information in the pedestrian trajectories, snapshot information, and a position corresponding to the snapshot in the trajectory; meanwhile, pedestrian features are extracted from the pedestrian images by a feature reasoning module; and finally, the pedestrian directions of the pedestrian features are clustered by a feature clustering module to obtain all trajectories of the pedestrians in a region. It can collect pedestrian images in multiple directions from a lens and can alleviate the recognition problems caused by occlusion, thereby improving the accuracy of pedestrian re-identification.
[0027] The application covers a wide angle and less occlusion, solves the problem of occlusion when a probe collects a target, and the problem that a probe can only obtain an image of a single direction of a human body when shooting, which leads to an inaccurate judgment of a pedestrian trajectory
[0028] In a third aspect, the application provides a pedestrian re-identification device based on top angle shooting, comprising: a processor and a memory. The memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the pedestrian re-identification device based on top angle shooting executes the pedestrian re-identification method based on top angle shooting.
[0029] As described above, the pedestrian re-identification method, system, medium and device based on top angle shooting of the application have the following advantages
[0030] Advantages:
[0031] The application utilizes the characteristic that a single track under a vertical overhead camera can capture multi-angle pedestrian pictures, classifies the pedestrian pictures into multiple directions, trains only the ReID model for matching pedestrians in the same direction, finally synthesizes the adjacency matrix calculated under the features of each angle to obtain the final adjacency matrix for clustering. The application proposes a reliable scheme based on ReID of a vertical overhead camera, provides effective algorithm support for the implementation of ReID of a vertical overhead camera, and meanwhile, the overhead ReID has greater flexibility in a specific scene, can capture more human feature information, and has higher robustness to environmental problems such as occlusion. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 An implementation schematic diagram of the pedestrian re-identification method based on overhead angle shooting of the application in an application scenario is shown.
[0033] Figure 2 A flowchart of S11 in the pedestrian re-identification method based on overhead angle shooting of the application in an embodiment is shown.
[0034] Figure 3 A flowchart of S11 in the pedestrian re-identification method based on overhead angle shooting of the application is shown.
[0035] Figure 4A A flowchart of S12 in the pedestrian re-identification method based on overhead angle shooting of the application is shown.
[0036] Figure 4B A schematic diagram of pedestrian track formation under a vertical overhead camera in the pedestrian re-identification method based on overhead angle shooting of the application in an embodiment is shown.
[0037] Figure 5 A flowchart of S13 in the pedestrian re-identification method based on overhead angle shooting of the application is shown.
[0038] Figure 6 A flowchart of S14 in the pedestrian re-identification method based on overhead angle shooting of the application is shown.
[0039] Figure 7 A schematic diagram of the principle structure of the pedestrian re-identification system based on overhead angle shooting of the application in an embodiment is shown.
[0040] Figure 8 A schematic diagram of the principle structure of the pedestrian re-identification device based on overhead angle shooting of the application in an embodiment is shown.
[0041] Element number explanation
[0042] 71 acquisition processing module
[0043] 72 direction reasoning module
[0044] 73 feature inference module
[0045] 74 feature clustering module
[0046] 81 processor
[0047] 82 memory
[0048] S11-S14 steps DETAILED DESCRIPTION
[0049] The present application is herein described, by way of example only, with reference to embodiments thereof. It is to be understood that variations and modifications will be apparent to those skilled in the art and that the scope of the present application encompasses all such obvious variations and modifications. Where, in the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application, it will be apparent to those skilled in the art that the present application can be practiced without such specific details. In other instances, well-known methods have not been described in detail in order not to unnecessarily obscure the present application.
[0050] It is to be understood that the drawings shown below are only schematic and that the actual implementation of the present application can vary as a consequence of, for example, the desired use, embedded technologies, as well as the technical possibilities available. It is also to be understood that the drawings are only schematic and that, for clarity, not every component can be depicted with exact dimensions.
[0051] The pedestrian re-identification method based on top-view angle shooting provided in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application.
[0052] Please refer to Figure 1 and Figure 2 , respectively showing the implementation schematic diagram of the pedestrian re-identification method based on top-view angle shooting of the present application in an application scenario and showing the flow schematic diagram of the pedestrian re-identification method based on top-view angle shooting of the present application in an embodiment. As shown in Figure 1 and Figure 2 , the present embodiment provides a pedestrian re-identification method based on top-view angle shooting.
[0053] The pedestrian re-identification method based on top-view angle shooting specifically includes the following steps:
[0054] S11, obtaining a pedestrian trajectory of a pedestrian under a vertical top-view camera, and collecting a pedestrian image based on the pedestrian trajectory. Please refer to Figure 3 , showing the flow schematic diagram of S11 in the pedestrian re-identification method based on top-view angle shooting of the present application. As shown in Figure 3 , the S11 includes the following steps:
[0055] S111, obtaining a pedestrian video collected by the vertical overhead camera.
[0056] A vertical overhead camera is installed at the top of the target area, and the shooting angle of the vertical overhead camera is vertically distributed with the ground.
[0057] A pedestrian video is collected by the vertical overhead camera. The vertical overhead camera uses a video collection device that can shoot the target area.
[0058] The pedestrian video can be directly obtained by a monitoring device (such as a vertical overhead camera) installed in the target area, and the pedestrian video in the target area is obtained.
[0059] S112, obtaining a pedestrian trajectory in the pedestrian video based on a tracking algorithm.
[0060] After obtaining the pedestrian video, a monitoring picture is extracted from the pedestrian video, and finally a pedestrian trajectory is obtained. The pedestrian trajectory in the pedestrian video is obtained based on a tracking algorithm. The pedestrian trajectory includes: trajectory information of the pedestrian, snapshot information, and corresponding position of the snapshot in the trajectory, etc. The tracking algorithm can be implemented by DeepSort, ByteTrack, etc.
[0061] S113, uniformly collecting pedestrian images based on the pedestrian trajectory.
[0062] Pedestrian images are extracted from the pedestrian trajectory at a uniform sampling rate. The information described by the pedestrian images includes: picture ID, image position information, ID of the uploaded vertical overhead camera, image shooting time, image description information, etc. The pedestrian images are RGB images. Among them, the extracted pedestrian images include the whole body of the pedestrian.
[0063] S12, obtaining a pedestrian direction under each pedestrian trajectory based on the pedestrian image.
[0064] Please refer to Figure 4A and Figure 4B , respectively showing the flowchart of S12 in the pedestrian re-identification method based on the top angle shooting of the present application and the pedestrian trajectory formation diagram under the vertical overhead camera in an embodiment of the pedestrian re-identification method based on the top angle shooting of the present application. As shown in Figure 4A and Figure 4B , the S12 comprises the following steps:
[0065] S121, the projection area centered on the vertical overhead camera is a preset area, and the coordinates of each point on the pedestrian trajectory in the preset area are obtained.
[0066] In this embodiment, the projection area of the vertical top-view camera is taken as the center and the coordinate position of each point on the pedestrian trajectory in the preset area is obtained.
[0067] Specifically, let's take a pedestrian trajectory as an example. Figure 4B As shown, in the image captured by the vertical top-view camera, the camera is positioned at the exact center of the frame. A pedestrian is visible in the frame, and the arc represents the pedestrian's walking trajectory. A coordinate axis is established, centered on the vertical top-view camera, with its coordinates set to (0, 0). The x-axis is a line parallel to the long sides AB and CD of the camera's frame and passing through the camera. The y-axis is a line parallel to the short sides AC and BD of the camera's frame and passing through the camera. Any point on the pedestrian's trajectory is represented as (x, y). These coordinates are relative coordinates, and their values range from [-1, 1]. When x = 1, it represents any point on side BD; when x = -1, it represents any point on side AC. When a point (x, y) in the pedestrian's trajectory forms an angle with the line connecting it to the vertical top-view camera, and this angle is set to α, a circular projection area is formed when the image is projected downwards onto the ground at an angle β, centered on the vertical top-view camera. This projection area is the preset area of the vertical top-view camera. Therefore, the corresponding coordinates are present in both the probe image captured by the vertical top-view camera and within this area.
[0068] S122, Based on the coordinates of each point in the pedestrian image and the pedestrian trajectory, predict the direction of the pedestrian in the pedestrian image.
[0069] Please continue reading. Figure 4B In this embodiment, the direction in the pedestrian image is predicted based on relevant information such as the pedestrian trajectory and the coordinates of each point on the pedestrian trajectory in the image captured by the vertical top-view camera. The pedestrian images captured are categorized by orientation as follows: front, back, left side, right side, and top. Specifically, front, back, left side, and right side refer to the orientation corresponding to images captured by the pedestrian at a point other than the center of the vertical top-view camera; top refers to the orientation corresponding to images captured near the center point of the camera (i.e., within a preset area). Unlike oblique-illuminated cameras, vertical top-view cameras can naturally obtain the direction information of a pedestrian at a certain point by observing the pedestrian's normal walking trajectory.
[0070] Specifically, as shown in Table 1, once the location of the pedestrian, the angle α between the pedestrian's location and the vertical top-view camera, and the tilt angle β between the vertical top-view camera and the ground are determined, the pedestrian's orientation can be predicted.
[0071] Table 1: Relationship of Pedestrian Trajectory Position Parameters under Vertical Top-View Camera
[0072] x y α β Direction [-1,1] [-1,1] [0°,180°] [0°,25°] Top surface [-1,1] [-1,1] [0°,45°] [25°,90°) Front surface [-1,1] [-1,1] [135°,180°] [25°,90°) Rear surface [-1,0] [-1,0] [45°,90°] [25°,90°) Right side surface [0,1] [-1,0] [90°,135°] [25°,90°) Right side surface [-1,0] [-1,0] [90°,135°] [25°,90°) Left side surface [0,1] [-1,0] [45°,90°] [25°,90°) Left side surface [-1,0] [0,1] [90°,135°] [25°,90°) Right side surface [0,1] [0,1] [45°,90°] [25°,90°) Right side surface [-1,0] [0,1] [45°,90°] [25°,90°) Left side surface [0,1] [0,1] [90°,135°] [25°,90°) Left side surface
[0073] For example, when the coordinates (x, y) of the pedestrian are (0.5, 1), the angle β of the vertical top-view camera is 10°, and the angle α between the line connecting the coordinates (0.5, 1) of the pedestrian and the vertical top-view camera and the vertical top-view camera is 120°, according to the parameter relationship in the table, it can be predicted that the direction of the pedestrian at this point is the top surface; when the coordinates (x, y) of the pedestrian are (-0.5, 0), the angle β of the vertical top-view camera is 40°, and the angle α between the line connecting the coordinates (-0.5, 0) of the pedestrian and the vertical top-view camera and the vertical top-view camera is 110°, according to the parameter relationship in the table, it can be predicted that the direction of the pedestrian at this point is the left side surface. Similarly, the directions of all points in the picture under the vertical top-view camera can be predicted.
[0074] S13, acquiring pedestrian features under each pedestrian trajectory based on the pedestrian image.
[0075] Referring to Figure 5 , a flowchart of S13 in the pedestrian re-identification method based on top-view angle shooting is shown. As Figure 5 indicated, S13 includes the following steps:
[0076] S131, pre-processing the pedestrian image under each pedestrian trajectory.
[0077] Image pre-processing is performed based on the acquired pedestrian image.
[0078] In this embodiment, there are several captured images for each pedestrian trajectory, that is, there are several pedestrian images for each pedestrian trajectory. Each pedestrian image is scaled to obtain a pedestrian image of uniform size.
[0079] Specifically, the width and height of the acquired pedestrian image are scaled to obtain a scaled pedestrian image. The size of the scaled pedestrian image is (N, 3, H, W); wherein N is the number of images in a batch processing, 3 is the number of channels, H is the height of the scaled pedestrian image, and W is the width of the scaled pedestrian image.
[0080] S132, inputting the pre-processed pedestrian image into a trained convolutional neural network for pedestrian feature recognition to acquire pedestrian features of the pedestrian image.
[0081] In this embodiment, the pedestrian features of the pedestrian image are extracted by training the scaled pedestrian image through the convolutional neural network model. ResNet50 is selected as an example for illustration in this embodiment.
[0082] Specifically, Resnet50 is a neural network structure of deep learning. Among them, Resnet.Stage comes from the module of Resnet50 itself, and ResNet is divided into 5 stage stages (i.e., Stage 0, Stage 1, Stage 2, Stage 3, and Stage 4). The structure is to take the scaled pedestrian image as the input object of the Resnet50 structure, process it through the five stages of Stage 0-4, realize the purpose of pedestrian feature extraction of the pedestrian image, and finally obtain the pedestrian feature with a size of (N, Fdim). Wherein, Fdim is the dimension of the pedestrian feature of the pedestrian image.
[0083] Specifically, the scaled pedestrian image (i.e., the scaled image) obtained in the S131 step is input into the Resnet50 structure, and finally the pedestrian feature of the scaled image is obtained through the processes of convolution layer, batch normalization layer and linear rectifier function, maximum pooling layer, and average pooling layer. As shown in Table 2:
[0084] Table 2: Resnet50 extracts pedestrian feature implementation structure
[0085] Conv3x3-bn-relu Maxpool 7*7 ResNet50 layer 1 ResNet50 layer 2 ResNet50 layer 3 ResNet50 layer 4 Avgpool FC(2048, IDnum)
[0086] In the above structure, Conv3x3-bn-relu represents a convolution layer with a convolution kernel size of 3x3; a batch normalization layer BN layer and a linear rectifier function ReLU; MaxPool 7x7 represents a maximum pooling layer with a kernel size of 7x7 and a stride of 7; ResNet50layer is a component in the ResNet50 network; AvgPool is a global average pooling layer with an output dimension of 2048; FC is a fully connected layer, and finally a feature inference model is obtained, and the output is the number of pedestrian IDs.
[0087] In this embodiment, the basic structure of using a neural network model for feature inference can use any mainstream ReID model, such as ViT, SeResNet, etc. for processing.
[0088] The learning of the neural network refers to the process of automatically obtaining the optimal weight parameters from the training data, that is, the process of self-training of the neural network. In order to optimize the neural network parameters, a loss function method is used. That is, in the process of neural network training, through the Cross Entropy loss and the Triplet loss, the model can recognize the pedestrian features in different directions.
[0089] The Cross Entropy loss function is:
[0090]
[0091] where p(k|x) represents the predicted probability of the pedestrian image x belonging to the pedestrian k after the model; q(k|x) is the one hot form of the true label of the pedestrian image x.
[0092] In the cross-entropy loss function, q(k|x) is 1 only when k is equal to the true label of x, and 0 in other cases.
[0093] The triplet loss function is:
[0094] L(A, P, N) = ||f(A) - f(P)| 2 -||f(A) - f(N)| 2 + α
[0095] Where A is an anchor point, which can be any feature; P is the sample farthest from the anchor point, but the label is the same as A; N is the sample closest to the anchor point, but the label is different from A; α is a hyperparameter greater than 0 to prevent the value of the loss from being negative and to enlarge the optimization space.
[0096] After the above cross-entropy loss and triplet loss, the ability of the model to recognize pedestrian features across directions is improved. However, the problem of insufficient recognition ability of the feature inference model for directionally different features has not been overcome. Among them, directionally different features refer to some features that are unique to pedestrians in different directions, such as different clothes logos on the front / back of the pedestrian, and side features such as diagonal shoulder bags. These are factors that affect the accuracy of pedestrian re-identification. Therefore, on the basis of processing by the cross-entropy loss function and the triplet loss function, the model is further processed by a human direction-specific cross-entropy loss function.
[0097] The human direction-specific cross-entropy loss function is:
[0098]
[0099]
[0100] Where ε is a hyperparameter representing the degree of softening; n represents the number of labels; ydir(k|x) represents the softening value of the pedestrian picture x under the label k; pdir(k|x) represents the probability of the model predicting that the pedestrian picture x belongs to the label k.
[0101] The human body direction specialization cross-entropy loss function is in the form of a smooth cross-entropy loss, but the label is extended from the ID of the pedestrian to the ID of the pedestrian and the direction; n is the number of labels, which corresponds to the number of directions 5 multiplied by the number of pedestrian IDs, that is: the loss function distinguishes the different directions of the pedestrians on the basis of pedestrian classification, so as to improve the anisotropic perception ability of the network to pedestrians. At the same time, the human body direction specialization cross-entropy loss function does not need to obtain direction information through manual labeling.
[0102] S14, obtaining a pedestrian re-identification result based on the pedestrian features and the pedestrian directions under each pedestrian trajectory.
[0103] Referring to Figure 6 , a flowchart of S14 in the pedestrian re-identification method based on top-angle shooting of the present application is shown. As Figure 6 indicated, the S14 includes the following steps:
[0104] S141, constructing a same-direction adjacency matrix based on the pedestrian features and the pedestrian directions under each pedestrian trajectory.
[0105] In step S13, the pedestrian features are extracted from the pedestrian image. At the same time, each pedestrian trajectory has multiple features of different directions, but due to the randomness of the movement route of the pedestrian in the vertical top-view camera, not all trajectories can cover all directions. Therefore, in this embodiment, a hierarchical combined density clustering method is used to output the final result of pedestrian matching.
[0106] Specifically, in this embodiment, a same-direction adjacency matrix is constructed based on the pedestrian features and the pedestrian directions under each pedestrian trajectory.
[0107] The same-direction adjacency matrix is as follows:
[0108]
[0109] Wherein, A represents the adjacency matrix of pedestrian i and pedestrian j; x i and x j represent the pedestrian features of pedestrian i and pedestrian j respectively; d represents the direction of the pedestrian trajectory, which belongs to the intersection of the direction of pedestrian trajectory i and the direction of pedestrian trajectory j; n represents the number of pedestrian directions d; x i d represents the pedestrian feature of pedestrian trajectory i in the d direction, x j d represents the pedestrian feature of pedestrian trajectory j in the d direction.
[0110] After the same-direction adjacency matrix is constructed, the average value of the feature pairs with common directions in the pedestrian features in the same-direction adjacency matrix is calculated by the Euclidean distance formula, as the distance of the feature pairs with common directions.
[0111] The calculation formula of the Euclidean distance or cosine distance used in this embodiment is as follows:
[0112] Euclidean distance:
[0113]
[0114] Cosine distance:
[0115]
[0116] wherein Dist represents a pedestrian feature distance calculation function, X represents a pedestrian feature vector 1; Y represents a pedestrian feature vector 2; x i represents a specific value in the pedestrian feature vector X; y i represents a specific value in the pedestrian feature vector Y.
[0117] The Euclidean distance between each pair of pedestrian features in the pedestrian features of all common directions is calculated, and the average value is obtained as the distance between the pedestrian trajectory i and the pedestrian trajectory j in the direction. Similarly, if the two do not have a common direction, the similarity of the two is considered to be 0.
[0118] S142, the pedestrian trajectories and pedestrian directions in the same direction adjacency matrix are clustered to obtain the clustered clusters.
[0119] The pedestrian features in the same direction of all trajectories in step S141 are calculated in turn. After completion, all trajectories are clustered to obtain the clustered clusters.
[0120] In this embodiment, the DBSCAN density clustering method is taken as an example for illustration. First, the search radius is set to τ same , and the minimum number of sample points in the neighborhood is P same All trajectories are clustered to obtain the clustered clusters.
[0121] Specifically, first, a plurality of pedestrian features obtained from all pedestrian trajectories are taken as a pedestrian feature data set, and a data p is randomly selected from the pedestrian feature data set; the data is taken as the center, the set search radius τ same is taken as the preset neighborhood radius, and the neighborhood region of the data is set.
[0122] When the number of data points in the neighborhood region of the data is greater than the preset point threshold, the data is defined as a core point; when the number of data points in the neighborhood region of the data is less than the preset point threshold, the data is defined as a boundary point; and the remaining data is defined as a noise point.
[0123] The selected data object point p is the center point, the neighborhood radius is τ same , and the minimum number of sample points in the neighborhood is Psame is a preset value of the neighborhood radius, and the minimum point number is a preset value of the point number threshold; and entering a next operation.
[0124] In this embodiment, it is assumed that the total number of data in the pedestrian feature dataset is 100, the neighborhood radius is preset to be 3, and the minimum point number is preset to be 5. When the number of data points in the neighborhood range with the data object point p as the center and the radius of 3 is greater than 5, the data object point p is a core point of the data object p; otherwise, it is a non-core point of the data object p. All data in the detection dataset are processed in sequence to complete the judgment of all data in the detection dataset.
[0125] If the distance between the core points is less than the preset neighborhood radius, the two core points with the distance less than the neighborhood radius are classified into the same dense region.
[0126] All core points obtained by the judgment in the above steps are sequentially calculated; the relationship between the distance between different core points and the neighborhood radius is calculated. When the distance between the core points is less than the neighborhood radius, it can be determined that the core points are in the same dense region; when the distance between the core points is greater than or equal to the neighborhood radius, it can be determined that the core points are not in the same dense region.
[0127] In this embodiment, it is assumed that there are different core points m and n. When the distance between the core point m and the core point n is less than 3, the two core points m and n are considered to be in the same dense region; when the distance between the core point m and the core point n is greater than or equal to 3, the two core points m and n are considered not to be in the same dense region.
[0128] The above steps are executed in a loop until all data in the pedestrian feature dataset are selected to form a plurality of different data dense regions, and each data dense region is a cluster. Therefore, a plurality of clustered clusters can be obtained.
[0129] In this embodiment, the input of DBSCAN is a plurality of 2048-dimensional pedestrian features, and the preset neighborhood radius τ of DBSCAN same and the minimum point number threshold P same will be flexibly valued according to different sites, and will be adjusted according to the actual effect, and finally the cluster of density-connected dense regions is output.
[0130] S143, based on the clustered cluster, constructing an asymmetric adjacency matrix.
[0131] Based on the clustered cluster obtained in step S142, an asymmetric adjacency matrix is constructed.
[0132] The asymmetric adjacency matrix is as follows:
[0133]
[0134] wherein c i and c j represent the feature set of cluster i and cluster j after clustering; d i represents the direction set of cluster i; k represents the kth direction of cluster i; d j represents the direction set of cluster j; l represents the lth direction of cluster j; n represents the number of different combinations of directions between cluster i and cluster j.
[0135] After the construction of the anti-adjacency matrix, the average value of the feature pairs of different directions in the co-adjacency matrix is calculated by the Euclidean distance formula, as the distance of different direction feature pairs.
[0136] The Euclidean distance between all different direction clusters is calculated, and the average value is obtained as the distance between cluster c i and cluster c j . Otherwise, the similarity between clusters is 0.
[0137] S144, clustering the clusters in the anti-adjacency matrix to obtain the final pedestrian re-identification result.
[0138] According to the calculation of all clusters in step S143, the final pedestrian re-identification result is obtained by clustering all clusters.
[0139] In this embodiment, the DBSCAN density clustering method is used again for clustering. The search radius is set to τ diff , and the minimum number of sample points in the neighborhood is P diff . All clusters are clustered to obtain the result of pedestrian re-identification.
[0140] Specifically, first, a data q is randomly selected from all clusters; the data is taken as the center, the set search radius P diff is set as the preset neighborhood radius, and the neighborhood area of the data is set.
[0141] The selected data object point q is the center point, the neighborhood radius is P diff , and the minimum number of sample points P diff in the neighborhood is the point threshold; the neighborhood radius is preset, and the neighborhood radius preset value is taken as the neighborhood area; the minimum number of points is preset, and the minimum number of points preset value is taken as the point threshold; the next step of operation is entered.
[0142] All core points determined in the above steps are sequentially calculated; the relationship between the distance between different core points and the neighborhood radius is calculated, when the distance between the core points is less than the neighborhood radius, it can be determined that the core points are in the same dense area; when the distance between the core points is greater than or equal to the neighborhood radius, it can be determined that the core points are not in the same dense area.
[0143] The above steps are repeatedly executed until all clusters are selected to obtain a pedestrian re-identification result, thereby realizing the process of determining the identity of the pedestrian.
[0144] The clustering algorithm in the above steps can also be replaced by other classical clustering methods or ReID related clustering algorithms. The feature distance calculation method can use Euclidean distance or cosine distance, or use the ReRank method in the ReID field. At the same time, common search constraints such as time and space can also be added to improve the clustering effect.
[0145] The pedestrian re-identification method based on top angle shooting adopted in the present application has a wide coverage angle, can pass a person under a vertical top view camera, and can present images in multiple directions such as front, back, top and side; at the same time, the occlusion between pedestrians is greatly alleviated in the case that the angle between the normal line of the probe and the ground and the head of the pedestrian is small.
[0146] The protection scope of the pedestrian re-identification method based on top angle shooting described in the embodiments of the present application is not limited to the order of steps listed in the embodiments, and any scheme realized by adding, replacing or replacing steps of the prior art according to the principle of the present application is included in the protection scope of the present application.
[0147] The embodiments additionally provide a computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the pedestrian re-identification method based on top angle shooting. Figure 1 The pedestrian re-identification method based on top angle shooting.
[0148] At any possible technical detail combination level, the present application can be a system, a method and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon, for causing a processor to implement various aspects of the present application.
[0149] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0150] The computer readable program here described can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device. Computer readable program instructions for carrying out operations of the present application can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.
[0151] The embodiment of the present application also provides a pedestrian re-identification system based on top-angle shooting. The pedestrian re-identification system based on top-angle shooting can implement the pedestrian re-identification method based on top-angle shooting described in the present application. However, the implementation device of the pedestrian re-identification method based on top-angle shooting described in the present application includes but is not limited to the structure of the pedestrian re-identification system based on top-angle shooting listed in the embodiment. Any modification and replacement of the existing technology according to the principle of the present application are included in the protection scope of the present application.
[0152] The pedestrian re-identification system based on top-angle shooting provided by the embodiment will be described in detail below in combination with the drawings.
[0153] Please refer to Figure 7 , which shows the principle structure schematic diagram of the pedestrian re-identification system based on top-angle shooting in an embodiment of the present application. As shown inFigure 7 As shown, the pedestrian re-identification system based on top corner shooting includes a collection processing module 71, a direction reasoning module 72, a feature reasoning module 73, and a feature clustering module 74.
[0154] The collection processing module 71 is configured to acquire a pedestrian image collected by a vertical top-view camera, and acquire at least one pedestrian trajectory based on the pedestrian image.
[0155] A vertical top-view camera is installed at the top of the target area, so that the shooting angle of the vertical top-view camera is vertically distributed with the ground.
[0156] A pedestrian video is collected by the vertical top-view camera. The vertical top-view camera is a video collection device capable of shooting the target area.
[0157] The pedestrian video can be directly shot by a monitoring device (such as a vertical top-view camera) installed in the target area to obtain a pedestrian video in the target area.
[0158] After the pedestrian video is acquired, a monitoring picture is extracted from the pedestrian video, and finally a pedestrian trajectory is obtained. The pedestrian trajectory of the pedestrian video is acquired based on a tracking algorithm. The pedestrian trajectory includes trajectory information of the pedestrian, snapshot information, and a corresponding position in the trajectory where the snapshot is taken, etc.
[0159] The pedestrian image is uniformly collected based on the pedestrian trajectory.
[0160] Specifically, the pedestrian image is extracted from the pedestrian trajectory at a uniform sampling rate. The information described by the pedestrian image includes a picture ID, image position information, an ID of an uploaded vertical top-view camera, image shooting time, image description information, etc. The pedestrian image is an RGB image. The extracted pedestrian image includes a full body of the pedestrian.
[0161] The direction reasoning module 72 is connected to the collection processing module 71, and is configured to acquire a pedestrian direction under each pedestrian trajectory based on the pedestrian image.
[0162] A projection area centered on the vertical top-view camera is a preset area, and coordinates of each point on the pedestrian trajectory in the preset area are acquired.
[0163] In this embodiment, the projection area of the vertical top-view camera is taken as the preset area, and then the coordinate positions of each point on the pedestrian trajectory in the preset area are acquired.
[0164] Specifically, taking a pedestrian trajectory as an example. In the picture taken by the vertical top-view camera, the vertical top-view camera is located at the center of the picture, there is a pedestrian in the picture, and the arc line represents the walking trajectory of the pedestrian. At this time, the coordinate axis is set with the vertical top-view camera as the center, and the coordinates are set as (0, 0); the line parallel to the long edges AB and CD of the picture of the vertical top-view camera and passing through the probe is taken as the horizontal axis x-axis, and the line parallel to the short edges AC and BD of the picture of the vertical top-view camera and passing through the probe is taken as the vertical axis y-axis; any point on the pedestrian trajectory is (x, y). When a point (x, y) in the pedestrian trajectory and the vertical top-view camera form an included angle, the included angle is set as a; when the vertical top-view camera is projected downward at an angle of β to the ground, a circular projection area can be formed, which is the preset area of the vertical top-view camera. Therefore, in the picture taken by the vertical top-view camera and in the area, there are corresponding coordinates.
[0165] In this embodiment, the direction of the pedestrian image is predicted according to the related information such as the pedestrian trajectory and the coordinates of each point of the pedestrian trajectory in the picture taken by the vertical top-view camera. The pedestrian image captured by the pedestrian is classified into the following categories of direction, such as front, back, left side, right side and top. Among them, the front, back, left side and right side are the directions corresponding to the pictures captured by the vertical top-view camera at non-central points of the pedestrian; and the top is the direction corresponding to the picture captured by the vertical top-view camera near the center of the probe (i.e. in the preset area). Unlike the oblique probe, the vertical top-view camera can naturally obtain the direction information of the pedestrian at a certain point through the trajectory of the pedestrian walking normally.
[0166] The feature inference module 73 is configured to obtain pedestrian features under each pedestrian trajectory based on the pedestrian image.
[0167] Image preprocessing is performed based on the obtained pedestrian image.
[0168] In this embodiment, there are several captured images for each pedestrian trajectory, that is, there are several pedestrian images for each pedestrian trajectory. Each pedestrian image is scaled to obtain a pedestrian image of a uniform size.
[0169] Specifically, the width and height of the obtained pedestrian image are scaled to a uniform size to obtain a scaled pedestrian image. The size of the scaled pedestrian image is (N, 3, H, W); wherein N is the number of images in a batch processing, 3 is the number of channels, H is the height of the scaled pedestrian image, and W is the width of the scaled pedestrian image.
[0170] The pedestrian features of the pedestrian image are extracted by training the scaled pedestrian image through the convolutional neural network model.
[0171] Specifically, the scaled pedestrian image (i.e., a scaled image) obtained in the step S131 is input into the neural network structure, and is processed through a convolution layer, a batch normalization layer and a linear rectifier function, a maximum pooling layer, and an average pooling layer, and finally, pedestrian features of the scaled image are obtained.
[0172] The learning of the neural network refers to a process of automatically obtaining optimal weight parameters from training data, that is, a process of self-training of the neural network. In order to optimize the neural network parameters, a loss function method is used. That is, in the neural network training process, through cross-entropy loss and ternary pair loss, the model can recognize pedestrian features across directions.
[0173] Through the cross-entropy loss and the ternary pair loss, the ability of the model to recognize pedestrian features across directions is improved. However, the problem of insufficient recognition ability of directionally different features of the feature inference model has not been overcome. Among them, directionally different features refer to some features that are unique to pedestrians in different directions, such as different clothes logos on the front / back of the pedestrian and the diagonal shoulder bag feature displayed on the side. These are factors that affect the accuracy of pedestrian re-identification. Therefore, on the basis of processing by the cross-entropy loss function and the ternary pair loss function, the model is further processed by a human direction-specific cross-entropy loss function.
[0174] The human direction-specific cross-entropy loss function is in the form of a Smooth cross-entropy loss, but the label is extended from the ID of the pedestrian to the ID of the pedestrian and the direction; n is the number of labels, corresponding to the number of directions 5 multiplied by the number of pedestrian IDs, that is, the loss function distinguishes different directions of pedestrians on the basis of pedestrian classification, in order to improve the perception ability of the network to the anisotropy of pedestrians. At the same time, the human direction-specific cross-entropy loss function does not need to obtain direction information through manual labeling.
[0175] The feature clustering module 74 is configured to obtain a pedestrian re-identification result based on the pedestrian features and the pedestrian directions under each pedestrian trajectory.
[0176] Pedestrian features are extracted from pedestrian images. At the same time, each pedestrian trajectory has multiple features in different directions, but due to the randomness of the movement route of pedestrians in the vertical top view camera, not all trajectories can cover all directions. Therefore, in this embodiment, a hierarchical clustering method combining density is used to output the final result of pedestrian matching.
[0177] Specifically, in this embodiment, a same-direction adjacency matrix is constructed based on the pedestrian features and the pedestrian directions under each pedestrian trajectory. After the same-direction adjacency matrix is constructed, the average value of each pair of features with the same direction in the pedestrian features in the same-direction adjacency matrix is calculated by using the Euclidean distance formula, as the distance of the pair of features with the same direction. The Euclidean distance between each pair of pedestrian features with the same direction in all pedestrian features is calculated, and the average value is obtained as the distance between the pedestrian trajectory i and the pedestrian trajectory j in the direction. Similarly, if the two do not have the same direction, the similarity between the two is considered to be 0.
[0178] The pedestrian features with the same direction in all trajectories are sequentially calculated. After the calculation is completed, clustering is performed on all trajectories to obtain the clustered clusters.
[0179] In this embodiment, the DBSCAN density clustering method is taken as an example for illustration. The search radius is first set to τ same , and the minimum number of sample points in the neighborhood is P same . All trajectories are clustered to obtain the clustered clusters.
[0180] In this embodiment, the input of DBSCAN is multiple 2048-dimensional pedestrian features, and the parameters of DBSCAN, the neighborhood radius Eps and the minimum point threshold MinPts, are flexibly valued according to different sites. Finally, the clusters of the density-connected dense regions are output.
[0181] Based on the clustered clusters, a different-direction adjacency matrix is constructed.
[0182] After the different-direction adjacency matrix is constructed, the average value of each pair of features with different directions in the different-direction adjacency matrix is calculated by using the Euclidean distance formula, as the distance of the pair of features with different directions.
[0183] The Euclidean distance between each pair of clusters with different directions is calculated, and the average value is obtained as the distance between the cluster c i and the cluster c j . Otherwise, the similarity between the cluster and the cluster is 0.
[0184] All clusters are sequentially calculated. After the calculation is completed, clustering is performed on all clusters to obtain the final pedestrian re-identification result.
[0185] In this embodiment, the DBSCAN density clustering method is again used for clustering. The search radius is first set to τ diff , and the minimum number of sample points in the neighborhood is P diff . All clusters are clustered to obtain the pedestrian re-identification result. The above steps are repeatedly executed until all clusters are selected to obtain complete pedestrian trajectories.
[0186] It should be noted that the division of the above system into various modules is only a logical functional division, and in actual implementation, all or part of them can be integrated into one physical entity, or can be physically separated. These modules can all be implemented in the form of software called by a processing element; they can all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the x module can be a separately established processing element, or it can be integrated into a chip of the above system, in addition, it can also be stored in the form of program code in the memory of the above system, and the function of the above x module is called and executed by a processing element of the above system. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described herein can be an integrated circuit with signal processing capability. In the implementation process, the steps of the above method or the above various modules can be completed by the integrated logic circuit of the hardware in the processor element or the instructions in the form of software.
[0187] The above modules can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. For example, when a certain module above is implemented in the form of program code called by a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together to implement in the form of system on a chip (SOC).
[0188] Please refer to Figure 8 , which shows the principle structure schematic diagram of the pedestrian re-identification device based on top angle shooting in an embodiment of the present application. As shown in Figure 8 , the present embodiment provides a pedestrian re-identification device based on top angle shooting, which comprises a processor 81 and a memory 82; the memory 82 is used to store computer programs; the processor 81 is connected with the memory 82, and is used to execute the computer programs stored in the memory 82, so that the pedestrian re-identification device based on top angle shooting executes each step of the pedestrian re-identification method based on top angle shooting as described above.
[0189] Preferably, the memory can comprise a Random Access Memory (RAM) and can also comprise a non-volatile memory, such as at least one disk memory.
[0190] The processor described above can be a general processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0191] To sum up, the pedestrian re-identification method, system, medium and device based on top angle shooting provided by the application have the following beneficial effects:
[0192] The application utilizes the characteristic that a single track under a vertical top-view camera can capture multi-angle pedestrian pictures, classifies the pedestrian pictures into multiple directions, trains only the ReID model for matching pedestrians in the same direction, finally calculates the adjacency matrix of the features in each angle, and obtains the final adjacency matrix for clustering. The application provides a reliable scheme based on the ReID of the vertical top-view camera, provides effective algorithm support for the implementation of the ReID of the vertical top-view camera, and meanwhile, the top-view ReID has greater flexibility in a specific scene, can capture more human feature information, and has higher robustness to environmental problems such as occlusion.
[0193] The above embodiments only exemplarily illustrate the principles and effects of the application, and are not used to limit the application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the application. Therefore, all equivalent modifications or changes completed by those skilled in the art without departing from the spirit and technical thought disclosed by the application should be covered by the claims of the application.
Claims
1. A pedestrian re-identification method based on top-angle photography, characterized in that, Includes the following steps: Obtain the pedestrian trajectory under the vertical top-view camera, and acquire pedestrian images based on the pedestrian trajectory; Based on the pedestrian images, obtain the pedestrian direction under each pedestrian trajectory; Based on the pedestrian images, obtain pedestrian features under each pedestrian trajectory; Based on the pedestrian features and directions under each pedestrian trajectory, the pedestrian re-identification results are obtained; The process of obtaining pedestrian re-identification results includes the following steps: Step 1: Based on the pedestrian features and directions under each pedestrian trajectory, construct a common-direction adjacency matrix; the formula for constructing the common-direction adjacency matrix is: Where A represents the adjacency matrix of pedestrian i and pedestrian j; x i and x j Let i and j represent the pedestrian characteristics of pedestrian i and pedestrian j respectively; d represents the direction of the pedestrian trajectory, which is the intersection of the directions of pedestrian trajectory i and pedestrian trajectory j; n represents the number of pedestrian directions d. Let i represent the pedestrian characteristics along the d direction of the pedestrian trajectory i. This represents the pedestrian characteristics along the d direction of pedestrian trajectory j; Step 2: Cluster the pedestrian trajectories and pedestrian directions in the same-direction adjacency matrix to obtain the clustered clusters; Step 3: Based on the clustered clusters, construct an off-direction adjacency matrix; the formula for constructing the off-direction adjacency matrix is: Among them, c i and c j Each represents the feature set of cluster i and cluster j after clustering; d i Denotes the set of directions of cluster i; k denotes the k-th direction of cluster i; d j Let represent the set of directions of cluster j; l represents the l-th direction of cluster j; n represents the number of different combinations of directions among all direction pairs between cluster i and cluster j; Step 4: Cluster the clusters in the anisotropic adjacency matrix to obtain the final pedestrian re-identification result.
2. The pedestrian re-identification method based on top-angle photography according to claim 1, characterized in that, Acquiring pedestrian trajectories from a vertical top-view camera and acquiring pedestrian images based on those trajectories includes the following steps: Acquire pedestrian videos captured by the vertical top-view camera; Pedestrian trajectories are obtained from pedestrian videos based on tracking algorithms; Pedestrian images are collected uniformly based on the pedestrian trajectory.
3. The pedestrian re-identification method based on top-angle photography according to claim 1, characterized in that, Obtaining the pedestrian direction based on the pedestrian images and the pedestrian trajectories includes the following steps: Using the projection area centered on the vertical top-view camera as the preset area, the coordinates of each point on the pedestrian trajectory in the preset area are obtained; Based on the coordinates of each point in the pedestrian image and the pedestrian trajectory, the direction of the pedestrian in the pedestrian image is predicted.
4. The pedestrian re-identification method based on top-angle photography according to claim 1, characterized in that, Obtaining pedestrian features based on the pedestrian images and their respective trajectories includes the following steps: Preprocess the pedestrian images under each pedestrian trajectory; The preprocessed pedestrian image is input into a trained convolutional neural network for pedestrian feature recognition to obtain the pedestrian features of the pedestrian image.
5. The pedestrian re-identification method based on top-angle photography according to claim 4, characterized in that, The loss function used in the convolutional neural network model; This includes: cross-entropy loss function, triplet loss function, and human orientation-specific cross-entropy loss function.
6. A pedestrian re-identification system based on top-angle photography, characterized in that, include: The acquisition and processing module is used to acquire the pedestrian trajectory under the vertical top-view camera and acquire pedestrian images based on the pedestrian trajectory; The direction reasoning module is used to obtain the pedestrian direction under each pedestrian trajectory based on the pedestrian image; The feature reasoning module is used to obtain pedestrian features under each pedestrian trajectory based on the pedestrian image; The feature clustering module is used to obtain pedestrian re-identification results based on pedestrian features and pedestrian directions under each pedestrian trajectory; The process of obtaining pedestrian re-identification results includes the following steps: Step 1: Based on the pedestrian features and directions under each pedestrian trajectory, construct a common-direction adjacency matrix; the formula for constructing the common-direction adjacency matrix is: Where A represents the adjacency matrix of pedestrian i and pedestrian j; x i and x j Let i and j represent the pedestrian characteristics of pedestrian i and pedestrian j respectively; d represents the direction of the pedestrian trajectory, which is the intersection of the directions of pedestrian trajectory i and pedestrian trajectory j; n represents the number of pedestrian directions d. Let i represent the pedestrian characteristics along the d direction of the pedestrian trajectory i. This represents the pedestrian characteristics along the d direction of pedestrian trajectory j; Step 2: Cluster the pedestrian trajectories and pedestrian directions in the same-direction adjacency matrix to obtain the clustered clusters; Step 3: Based on the clustered clusters, construct an off-direction adjacency matrix; the formula for constructing the off-direction adjacency matrix is: Among them, c i and c j Each represents the feature set of cluster i and cluster j after clustering; d i Denotes the set of directions of cluster i; k denotes the k-th direction of cluster i; d j Let represent the set of directions of cluster j; l represents the l-th direction of cluster j; n represents the number of different combinations of directions among all direction pairs between cluster i and cluster j; Step 4: Cluster the clusters in the anisotropic adjacency matrix to obtain the final pedestrian re-identification result.
7. A pedestrian re-identification device based on top-angle photography, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is connected to the memory and is used to execute the computer program stored in the memory so that the pedestrian re-identification device based on top-angle photography performs the pedestrian re-identification method based on top-angle photography according to any one of claims 1 to 5.
Citation Information
Patent Citations
Pedestrian re-identification method based on top view image, storage medium and electronic equipment
CN113033350A
Cross-border head tracking method and device based on pedestrian re-identification
CN114359348A