Cross-border re-identification method and device based on track matching
By adopting a trajectory matching method in cross-border re-identification technology, a comprehensive trajectory of feature coordinates is generated and global identification is allocated and updated, the problem of target recognition errors in the prior art with similar appearances or large perspectives is solved, and the accuracy and robustness of the recognition are improved.
Patent Information
- Application Number
- CN202411667064.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-21
AI Technical Summary
When existing cross-border re-identification technologies face different targets with high appearance similarity, the same target with large differences in different perspectives, and the same target with large differences caused by occlusion, they can easily lead to incorrect re-identification results.
Using a trajectory matching method, a comprehensive feature coordinate trajectory to be matched in the video stream is generated, and compared with the global identification of other video devices to determine the re-identification confidence. When the confidence level reaches the threshold, global identification is allocated and updated to ensure the typicality and accuracy of the feature trajectory.
It improves the accuracy and robustness of cross-border re-identification, reduces false recognition due to similar appearances or differences in perspectives, and enhances the reliability of the system.
Smart Images

Figure CN120071207A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a cross-border re-identification method and device based on trajectory matching. Background Art
[0002] Cross-border re-identification refers to a computer vision technology that understands the same target captured by different cameras as the same target from a computer perspective. It obtains the target of interest in the image through a deep convolutional neural network and uses image feature matching and other technologies to distinguish the identities of these targets of interest, and assigns the same ID to the same target. This is a newly emerging research direction in the field of computer vision, mainly involving identifying and tracking targets across multiple cameras or multiple geographical locations. This technology has broad application prospects in fields such as security monitoring, cross-border tracking, and drone / vehicle clusters.
[0003] Currently, cross-border re-identification is usually achieved through image feature matching. This method uses a convolutional neural network to extract the feature vectors of the target. When the Euclidean distance and cosine distance of two feature vectors meet the threshold requirements, it is considered that these two feature vectors are extracted from the same target. This method can achieve relatively ideal results under specific preconditions. However, when performing cross-border re-identification on different targets with high appearance similarity, the same target with large differences in different perspectives, and the same target with large differences caused by occlusion, since image feature matching only compares the similarity between images, it is inevitable that the computer will understand similar different targets as the same target and dissimilar same targets as different targets, thus obtaining incorrect re-identification results. Summary of the Invention
[0004] Provide a cross-border re-identification method and device based on trajectory matching to solve the above problems or at least partially solve the above problems.
[0005] In a first aspect, a cross-border re-identification method based on trajectory matching is disclosed. The method includes:
[0006] Step S1: Obtain the to-be-matched feature trajectory α corresponding to the video stream obtained by the i-th video device in , generate the to-be-matched coordinate trajectory β corresponding to the to-be-matched feature trajectory α in , and generate the to-be-matched feature coordinate comprehensive trajectory γ based on the to-be-matched feature trajectory α in and the to-be-matched coordinate trajectory β in ; where n is the n-th target newly detected in the video stream obtained by the i-th video device; in ; in ;
[0007] Step S2: Use the to-be-matched feature coordinate comprehensive trajectory γ inCompare with the comprehensive trajectory of the characteristic coordinates of the target obtained by other road video devices and already generated with a global identifier one by one to determine each re-identification confidence level ω in,jm , where each video stream corresponds to one or more targets; j represents the j-th road video device, and m is the m-th target obtained by analyzing the j-th video stream. Each road video device corresponds to a video stream;
[0008] Step S3: When there is one or more re-identification confidence levels ω in,jm greater than or equal to the preset threshold, obtain the maximum value of the re-identification confidence level ω in,jm , and denote this maximum value as ω in,j'm' , obtain the road number j' corresponding to the video device corresponding to ω in,j'm' and the global identifier corresponding to the m'-th target obtained by analyzing the corresponding video stream, and use this global identifier as the global identifier corresponding to the to-be-matched characteristic trajectory γ in ;
[0009] When all re-identification confidence levels ω in,jm are less than the preset threshold:
[0010] When the to-be-matched comprehensive trajectory of characteristic coordinates γ in is a continuation trajectory of the first comprehensive trajectory of characteristic coordinates, and the global identifier of the first comprehensive trajectory of characteristic coordinates corresponds to other comprehensive trajectories of characteristic coordinates except the first comprehensive trajectory of characteristic coordinates, assign new global identifiers to the to-be-matched comprehensive trajectory of characteristic coordinates γ in and the first comprehensive trajectory of characteristic coordinates; The continuation trajectory refers to: the comprehensive trajectory of characteristic coordinates generated by the same target that generated the first comprehensive trajectory of characteristic coordinates at a later time after the time when the first comprehensive trajectory of characteristic coordinates was generated;
[0011] When the to-be-matched comprehensive trajectory of characteristic coordinates γ in is a continuation trajectory of the first comprehensive trajectory of characteristic coordinates, and the global identifier of the first comprehensive trajectory of characteristic coordinates only corresponds to the first comprehensive trajectory of characteristic coordinates, assign the global identifier of the first comprehensive trajectory of characteristic coordinates to the to-be-matched comprehensive trajectory of characteristic coordinates γ in ;
[0012] When the to-be-matched comprehensive trajectory of characteristic coordinates γ in is not a continuation trajectory of any comprehensive trajectory of characteristic coordinates, assign a new global identifier to the to-be-matched comprehensive trajectory of characteristic coordinates.
[0013] Preferably, in step S1, obtaining the to-be-matched characteristic trajectory α in corresponding to the video stream obtained by the i-th road video device includes:
[0014] Step S11: Perform object detection on the video stream of the i-th video device to determine newly detected objects;
[0015] Step S12: For each newly detected object n, perform the following operations:
[0016] Perform RelD feature extraction on the video stream of the i-th video device through a feature extraction network model to obtain the visual feature vector of object n
[0017] The visual feature vector is matched with the feature trajectories of the first to the n - 1st objects that have been determined for the i-th video device;
[0018] If there is an object among the first to the n - 1st objects that have been determined for the i-th video device that matches object n, then record the local identifier of object n in the i-th video device as the local identifier of the object that matches object n in the i-th video device; obtain the feature trajectory trace′ corresponding to the object that matches object n in , and the visual feature vector is merged into the tail of the feature trajectory trace’ in to obtain the updated feature trajectory trace’ in ; the updated feature trajectory is used as the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in ;
[0019] If there is no object among the first to the n - 1st objects that have been determined for the i-th video device that matches object n, then verify the visual feature vector ;
[0020] The verification method is: Obtain several frames after the i-th video device detects object n. If the objects in at least three frames match object n, then the visual feature vector passes the verification; otherwise, the verification fails;
[0021] After the verification passes, obtain the visual feature vectors corresponding to the video frames that match object n respectively; arrange the visual feature vectors corresponding to the video frames that match object n respectively in chronological order to form the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in .
[0022] Preferably, matching the visual feature vector with the feature trajectories of the first to the n - 1st objects that have been determined for the i-th video device includes:
[0023] Calculate the visual feature vector The cosine distance, Euclidean distance, and intersection over union (IoU) with the feature trajectories of the first to the (n - 1)th targets already determined by the ith video device. When the visual feature vector meets the conditions that the cosine distance with a certain target among the first to the (n - 1)th targets already determined by the ith video device is greater than the first threshold, the Euclidean distance is less than the second threshold, and the IoU is greater than the third threshold, determine that the visual feature vector matches the target.
[0024] Preferably, in step S1: Generate a coordinate trajectory β to be matched corresponding to the feature trajectory α to be matched in , and based on the feature trajectory α to be matched in and the coordinate trajectory β to be matched in , generate a comprehensive feature coordinate trajectory γ to be matched in , including: in
[0025] Step S13: For each visual feature vector in the feature trajectory α to be matched in , determine the three-dimensional spatial coordinates (X W , Y W , Z W ) of the target corresponding to the visual feature vector in the world coordinate system, where:
[0026] (X W , Y W , Z W ) = R(X C , Y C , Z C ) + T
[0027]
[0028] where X W , Y W , Z W are the coordinates of the target corresponding to the visual feature vector on the X-axis, Y-axis, and Z-axis in the world coordinate system respectively, R is the rotation matrix of the camera coordinate system relative to the world coordinate system, X C , Y C, Z C are the coordinates of the target in the camera coordinate system respectively, T is the translation matrix of the camera coordinate system relative to the world coordinate system, u 0 , v 0 , f x , f y are all variables related to the internal parameters of the video device, f x , f y are the normalized focal lengths on the u-axis and v-axis respectively, u 0 and v0 are all the intersections of the video device and the image plane, f x = f / dX, f y = f / dY, where f is the focal length of the video device, and dX and dY respectively represent the sizes of unit pixels on the u-axis and v-axis of the pixel coordinate system; i1 is the coordinate value of the target on the u-axis in the pixel coordinate system, j1 is the coordinate value of the target on the v-axis in the pixel coordinate system, and D(i1, j1) is the depth value of the point (i1, j1);
[0029] Step S14: Denote the coordinate trajectory to be matched corresponding to the nth feature trajectory α in the ith video stream as β in in , where is the time - coordinate pair saved in β in ,
[0030] t inl is the time point when the target is detected, β inl =(X Winl , Y Winl , Z Winl ), β inl is the three - dimensional space coordinate of the target obtained by solution in the world coordinate system, and X Winl , Y Winl , Z Winl are respectively the coordinate values of the target on each coordinate axis in the world coordinate system; Combine the obtained feature trajectory α in and the coordinate trajectory β in to obtain the combined trajectory γ of the feature coordinates to be matched in =(α in , β in ).
[0031] Preferably, in step S2: Compare the combined trajectory γ of the feature coordinates to be matched with the combined trajectories of the feature coordinates of the targets obtained by other video devices and having generated global identifiers one by one to determine each re - identification confidence level ω in , including: in,jm For the combined trajectory of the feature coordinates of each target obtained by other video devices and having generated global identifiers, perform the following operations:
[0032]
[0033] Step S21: Randomly extract 3 feature vectors from the feature trajectory γ in the combined trajectory of the feature coordinates of the target obtained by other video devices and having generated global identifiers jm where x1, x2, x3 are the numbers of the feature vectors randomly extracted from γ jm
[0034] Randomly extract 3 feature vectors from the comprehensive trajectory γ of the to-be-matched feature coordinates in The corresponding α in where y1, y2, and y3 are the numbers of the feature vectors randomly extracted from α respectively; in ;
[0035] Calculate the similarity of the feature trajectories
[0036] Step S22: Determine whether the time interval corresponding to γ jm overlaps with the time interval corresponding to γ in . If so, go to Step S23; otherwise, go to Step S25;
[0037] Step S23: Take the overlapping part of the time interval corresponding to γ jm and the time interval corresponding to γ in as the comprehensive sub-trajectory of feature coordinates. Denote the comprehensive sub-trajectory of feature coordinates corresponding to γ jm and the comprehensive sub-trajectory of feature coordinates corresponding to γ in as {β jm(z1) , β jm(z2) , …, β jm(zo)} and {β in(z1) , β in(z2) , …, β in(zo)};
[0038] Determine the Euclidean distance between the two comprehensive sub-trajectories of feature coordinates:
[0039]
[0040] Use the hyperbolic tangent function tanh to transform the value range of the Euclidean distance between the two comprehensive sub-trajectories of feature coordinates into the interval [0, 1], denoted as
[0041] Calculate {β jm(z2) - β jm(z1) , β jm(z3) - β jm(z2) , …, β jm(zo) - β jm(z(o-1))};
[0042] Calculate {β in(z2) - β in(z1) , β in(z3) - β in(z2) , …, β in(zo) - β in(z(o-1))};
[0043] Obtain the similarity of the coordinate trajectories, denoted as
[0044]
[0045] Among them, β jm(z1) ,β jm(z2) ,…,β jm(zo) They are γ jm The corresponding characteristic coordinates are the position coordinates at each moment in the comprehensive sub-trajectory, β in(z1) ,β in(z2) ,…,β in(zo) They are γ in The corresponding characteristic coordinates are the position coordinates at each moment in the comprehensive sub-trajectory, is the Euclidean distance between the two feature coordinate integrated sub-trajectories, r is the cumulative variable, β jm(zr) The x, y, and z axis coordinate values, β in(zr) The x, y, and z axis coordinate values;
[0046] Step S24: Determine the re-identification confidence ω in,jm ,in:
[0047]
[0048] η, μ, and λ are weight parameters respectively;
[0049] The processing of the comprehensive trajectory of characteristic coordinates of the target obtained by the other video devices and having generated a global identifier is completed;
[0050] Step S25: Determine the re-identification confidence ω in,jm ,in:
[0051]
[0052] The processing of the comprehensive trajectory of characteristic coordinates of the target obtained by the other video devices and having generated a global identification is completed.
[0053] In a second aspect, a cross-border re-identification device based on trajectory matching is disclosed, the device comprising:
[0054] Acquisition module: configured to obtain the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in , generate the feature trajectory α to be matched in The corresponding coordinate trajectory to be matched β in , based on the feature trajectory to be matched α in And the coordinate trajectory to be matched β in , generate the comprehensive trajectory of the feature coordinates to be matched γ in; where n is the nth target newly detected in the video stream obtained by the ith video device;
[0055] Matching module: configured to comprehensively compare the trajectory γ of the feature coordinates to be matched in with the comprehensive trajectories of the feature coordinates of the targets obtained by other video devices and having generated global identifiers one by one, and determine various re-identification confidence levels ω in,jm , where each video stream corresponds to one or more targets; j represents the jth video device, m is the mth target obtained by analyzing the jth video stream, and each video device corresponds to a video stream;
[0056] Identification module: configured to, when there is one or more re-identification confidence levels ω in,jm greater than or equal to the preset threshold, obtain the maximum value of the re-identification confidence level ω in,jm , and this maximum value is denoted as ω in,j'm' , obtain the road number j' corresponding to the video device corresponding to ω in,j'm' and the global identifier corresponding to the m'th target obtained by analyzing the corresponding video stream, and use this global identifier as the global identifier corresponding to the trajectory γ of the feature coordinates to be matched in ;
[0057] When all the re-identification confidence levels ω in,jm are less than the preset threshold:
[0058] When the comprehensive trajectory γ of the feature coordinates to be matched in is a continuation trajectory of the first comprehensive trajectory of feature coordinates, and the global identifier of the first comprehensive trajectory of feature coordinates corresponds to other comprehensive trajectories of feature coordinates except the first comprehensive trajectory of feature coordinates, allocate new global identifiers for the comprehensive trajectory γ of the feature coordinates to be matched in and the first comprehensive trajectory of feature coordinates; the continuation trajectory refers to: the comprehensive trajectory of feature coordinates generated by the same target that generated the first comprehensive trajectory of feature coordinates at a later time after the time when the first comprehensive trajectory of feature coordinates was generated;
[0059] When the comprehensive trajectory γ of the feature coordinates to be matched in is a continuation trajectory of the first comprehensive trajectory of feature coordinates, and the global identifier of the first comprehensive trajectory of feature coordinates only corresponds to the first comprehensive trajectory of feature coordinates, allocate the global identifier of the first comprehensive trajectory of feature coordinates for the comprehensive trajectory γ of the feature coordinates to be matched in ;
[0060] When the comprehensive trajectory γ of the feature coordinates to be matched in is not a continuation trajectory of any comprehensive trajectory of feature coordinates, allocate a new global identifier for the comprehensive trajectory γ of the feature coordinates to be matched.
[0061] In a third aspect, an electronic device is disclosed, the electronic device comprising:
[0062] at least one processor; and
[0063] a memory communicatively connected to the at least one processor; wherein,
[0064] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0065] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is disclosed, the computer instructions being used to cause the computer to execute the method as described above.
[0066] The present invention has the following technical effects:
[0067] The present invention performs quality assessment on the target screenshots detected by the target detection algorithm, and uses the target screenshots with qualified quality assessment for ReID feature extraction and comparison. On the basis of ensuring the detection rate, it prevents the features extracted from the occluded or atypically state target screenshots from participating in feature comparison or being stored in the feature trajectory, ensuring the typicality of the feature trajectory. Therefore, it can greatly improve the problem of poor robustness of the original re-identification algorithm, where the feature vectors change greatly due to occlusion or atypically state, resulting in re-identification failure (the same target cannot be de-duplicated due to the existence of atypically state), and the reliability is greatly improved.
[0068] The present invention solves the three-dimensional space coordinates of the target by adding a depth sensor, and uses the similarity of the coordinate trajectory as a weight. At the same time, it considers the cosine similarity of the target feature trajectory, the balanced Euclidean distance of the target coordinate trajectory, and the average value of the cosine similarity of the target coordinate trajectory motion vector, so that the success rate of re-identification increases significantly. Since this method takes into account the spatial coordinates and motion rules, it provides an additional basis for distinguishing different targets with similar appearances, thus solving the problem that the original re-identification algorithm cannot distinguish different targets with high similarity; in addition, the weighting of the coordinate trajectory can compensate for the problem that the appearance of the same target varies in different video streams due to reasons such as angle and illumination, making it easier to de-duplicate and fuse the same target with low appearance similarity. This method simultaneously solves the problem of poor re-identification effect of the original re-identification algorithm in the face of different targets with similar appearances, the same target with different appearances, etc., and greatly improves the robustness of the re-identification algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a schematic diagram of the architecture of the cross-border re-identification method based on trajectory matching;
[0070] Figure 2 Schematic diagram of the target trajectory in the video stream of a single-channel video device;
[0071] Figure 3 Schematic diagram of the conversion relationship of the four major visual coordinate systems;
[0072] Figure 4 Schematic diagram of the process of another cross-border re-identification method based on trajectory matching. Detailed implementation mode
[0073] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0074] As Figure 1 shown, the present invention provides a cross-border re-identification method based on trajectory matching, and the method includes:
[0075] Step S1: Obtain the to-be-matched feature trajectory α corresponding to the video stream obtained by the i-th video device in , generate the to-be-matched coordinate trajectory β corresponding to the to-be-matched feature trajectory α in , and based on the to-be-matched feature trajectory α in and the to-be-matched coordinate trajectory β in , generate the to-be-matched feature coordinate comprehensive trajectory γ in ; where n is the n-th target newly detected in the video stream obtained by the i-th video device; in
[0076] Step S2: Compare the to-be-matched feature coordinate comprehensive trajectory γ in with the feature coordinate comprehensive trajectories of the targets obtained by other video devices and having generated global identifiers one by one, and determine each re-identification confidence level ω in,jm , where each video stream corresponds to one or more targets; j represents the j-th video device, m is the m-th target obtained by analyzing the j-th video stream, and each video device corresponds to a video stream;
[0077] Step S3: When there is one or more re-identification confidence levels ω in,jm greater than or equal to the preset threshold, obtain the maximum value of the re-identification confidence level ω in,jm , and this maximum value is denoted as ω in,j'm' , obtain the road number j' of the video device corresponding to ω in,j'm' and the global identifier corresponding to the m'-th target obtained by analyzing the corresponding video stream, and use this global identifier as the global identifier corresponding to the to-be-matched feature trajectory γ in ;
[0078] When all the re-identification confidence levels ω in,jm are less than the preset threshold:
[0079] When the comprehensive trajectory γ of the feature coordinates to be matched in is a continuation trajectory of the first comprehensive trajectory of feature coordinates, and the global identifier of the first comprehensive trajectory of feature coordinates corresponds to other comprehensive trajectories of feature coordinates other than the first comprehensive trajectory of feature coordinates, for the comprehensive trajectory γ of the feature coordinates to be matched in and the first comprehensive trajectory of feature coordinates, assign a new global identifier; the continuation trajectory refers to: the comprehensive trajectory of feature coordinates generated by the same target that generated the first comprehensive trajectory of feature coordinates at a later time after the time when the first comprehensive trajectory of feature coordinates is generated;
[0080] When the comprehensive trajectory γ of the feature coordinates to be matched in is a continuation trajectory of the first comprehensive trajectory of feature coordinates, and the global identifier of the first comprehensive trajectory of feature coordinates only corresponds to the first comprehensive trajectory of feature coordinates, for the comprehensive trajectory γ of the feature coordinates to be matched in assign the global identifier of the first comprehensive trajectory of feature coordinates;
[0081] When the comprehensive trajectory γ of the feature coordinates to be matched in is not a continuation trajectory of any comprehensive trajectory of feature coordinates, assign a new global identifier to the comprehensive trajectory of the feature coordinates to be matched.
[0082] As Figure 2 shown, the step S1 of obtaining the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in includes:
[0083] Step S11: Perform object detection on the video stream of the i-th video device to determine newly detected objects;
[0084] Step S12: For each newly detected object n, perform the following operations:
[0085] Perform RelD feature extraction on the video stream of the i-th video device through the feature extraction network model to obtain the visual feature vector of object n
[0086] Match the visual feature vector with the feature trajectories of the first to the n-1th objects that have been determined in the i-th video device;
[0087] If there is an object in the first to the n-1th objects that have been determined in the i-th video device that matches object n, record the local identifier of object n in the i-th video device as the local identifier of the object that matches object n in the i-th video device; obtain the feature trajectory trace′ corresponding to the object that matches object n in and merge the visual feature vector into the feature trajectory trace′in At the tail of, the updated feature trajectory trace' is obtained in ; The updated feature trajectory is used as the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in ;
[0088] If there is no target among the first to the n-1st targets already determined by the i-th video device that matches the target n, then the visual feature vector is verified;
[0089] The verification method is: Obtain several frames after the i-th video device detects the target n. If the targets in at least three frames match the target n, then the visual feature vector passes the verification; otherwise, the verification fails;
[0090] After the verification passes, obtain the visual feature vectors corresponding to the video frames that match the target n respectively; Arrange the visual feature vectors corresponding to the video frames that match the target n respectively in chronological order to form the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in 。
[0091] In the present invention, for a single video device, when performing RelD feature extraction, the extracted is a feature vector, and the feature trajectory is a set of feature vectors. When it is just recognized and extracted, it is not known which trajectory it belongs to. Therefore, it needs to be matched with all the existing trajectories of this path. This is a process of classifying the feature vector into a certain feature trajectory or generating a new feature trajectory for it.
[0092] Furthermore, matching the visual feature vector with the feature trajectories of the first to the n-1st targets already determined by the i-th video device includes:
[0093] Calculate the cosine distance, Euclidean distance and intersection over union of the visual feature vector and the feature trajectories of the first to the n-1st targets already determined by the i-th video device. When the visual feature vector meets the conditions that the cosine distance is greater than the first threshold, the Euclidean distance is less than the second threshold and the intersection over union is greater than the third threshold with a certain target among the first to the n-1st targets already determined by the i-th video device, it is determined that the visual feature vector matches this target.
[0094] In the present invention, object detection is performed on the video stream obtained from the i-th video device. Specifically, for the video stream of a single-channel video captured, an object recognition algorithm is used to detect the objects in the video stream. For example, the Yolov5 object detection algorithm is adopted. Since the Yolov5 object detection algorithm is a relatively mature open-source object detection algorithm, it will not be introduced in detail in the present invention and can be replaced according to needs. Through the object detection algorithm, screenshots of the objects, detection frames, and information on the object categories can be output.
[0095] Performing object detection on the video stream obtained from the i-th video device further includes: using the Yolov5 object detection algorithm for object detection, using a binary classification network to detect the results output by the Yolov5 object detection algorithm, retaining the results output by the Yolov5 object detection algorithm with qualified detection results, and filtering out the results output by the Yolov5 object detection algorithm with unqualified detection results.
[0096] The binary classification network model is a convolutional neural network model, and the loss function of the binary classification network model is:
[0097] Loss = ∑L(y, p) = ∑ -y·log(p) - (1 - y)·log(1 - p)
[0098] Among them, L(y, p) is the cross-entropy loss function of the training sample, y is the true label, taking values of 0 or 1, p is the predicted probability, taking values between 0 and 1, and L(y, p) represents the cross-entropy loss between the true label and the predicted probability.
[0099] In the present invention, the similarity of the appearance features of the object is an important basis for measuring whether two object screenshots come from the same object. However, when the object is occluded or in an atypical state (such as a person crawling), its appearance features will change greatly. In such cases, the features of the object are stored in the feature trajectory, which will cause the representation of the object by the feature trajectory to lose its typicality, resulting in a serious decline in the robustness of the cross-border re-identification algorithm. Object samples that are not occluded and in a typical state are defined as qualified samples, and object samples that are occluded or in an atypical state are defined as unqualified samples. To ensure the typicality of the stored feature trajectory, unqualified samples need to be filtered out before generating the feature trajectory.
[0100] The quality assessment model uses a simple binary classification network built based on a convolutional neural network, and trains the binary classification network model by making a training set. The training set divides the target screenshots that are not occluded and in a typical state into one category, assigns the label 0 to it, divides the target screenshots that are occluded or in an atypical state into one category, assigns the label 1 to it, and feeds the above training set into the built convolutional neural network. The convolutional neural network will output a predicted value, that is, the probability of the model predicting the category of the target screenshot. The cross-entropy loss is used as the loss function. For the binary classification task, the cross-entropy loss function predicted for each sample is as follows
[0101] L(y,p) = -y·log(p) - (1 - y)·log(1 - p)
[0102] where y is the true label (taking values of 0 or 1), and p is the probability predicted by the model (taking values between 0 and 1). L(y,p) represents the cross-entropy loss between the true label y and the predicted probability p.
[0103] For all samples in one training, its loss function is the sum of the cross-entropy losses of all samples, and the overall loss function is as follows
[0104] Loss = ∑L(y,p) = ∑-y·log(p) - (1 - y)·log(1 - p)
[0105] The trained model can be used to classify target samples (qualified or unqualified), directly filter out unqualified samples, and use qualified samples for feature extraction and feature trajectory generation in the next stage.
[0106] The target samples that have passed the quality assessment are sent into the feature extraction network model for ReID feature extraction. The feature extraction network model uses an existing open-source model, which can be replaced if there is a better-performing model. Calculate the cosine distance, Euclidean distance, and intersection over union (IoU) between the ReID feature vectors extracted from all detected targets and the most recently stored feature vectors of all existing feature trajectories in this video stream. When all of the following conditions are simultaneously met: the cosine distance is greater than a specified threshold (configurable), the Euclidean distance is less than a specified threshold (configurable), and the intersection over union (IoU) is greater than a specified threshold (configurable), it is considered that the target matches an existing feature trajectory. Incorporate the target feature into the corresponding trajectory and assign the target the same ID value as the trajectory (referred to as an update of the feature trajectory). When the target ReID feature fails to meet the conditions, match the IoU value and feature vector of the target in the previous and subsequent frames. When successful matching occurs for three consecutive frames (both the IoU, cosine distance, and Euclidean distance are within the threshold range), it is considered that a new target has been entered, and initialize the trajectory feature for this target (referred to as a newly generated trajectory). Among them, the threshold needs to be adapted according to the specified scenario. Denote the feature vector of the target to be matched as u, the area enclosed by the detection box as A, the feature vector to be matched as v, and the area enclosed by the detection box as B. The calculation formulas for the cosine distance dist(u, v), Euclidean distance ||u - v||, and intersection over union J(A, B) are as follows
[0107]
[0108] The feature trajectory of each target will contain its own feature information for multiple frames (set to 50 frames in the present invention, configurable). As the target moves, its feature trajectory can contain feature information in multiple directions of the target, greatly enhancing the robustness of subsequent feature matching between multiple video streams. Each video stream can generate multiple feature trajectories according to the number of targets appearing in its video frame. When a certain feature trajectory has no new feature vectors stored for a long time (set to 30 minutes in the present invention, configurable), delete this feature trajectory to prevent the accumulation and redundancy of feature trajectories.
[0109] Step S1: Generate the to-be-matched coordinate trajectory β in corresponding to the to-be-matched feature trajectory α in , and based on the to-be-matched feature trajectory α in and the to-be-matched coordinate trajectory β in , generate the to-be-matched comprehensive feature coordinate trajectory γ in , including:
[0110] Step S13: For each visual feature vector in the to-be-matched feature trajectory α in , determine the three-dimensional spatial coordinates (X W , Y W , Z W ) of the target corresponding to this visual feature vector in the world coordinate system, where:
[0111] (X W ,Y W ,Z W ) = R(X C ,Y C ,Z C ) + T
[0112]
[0113] Wherein, X W ,Y W ,Z W are respectively the coordinates of the target corresponding to the visual feature vector on the X-axis, Y-axis, and Z-axis in the world coordinate system, R is the rotation matrix of the camera coordinate system relative to the world coordinate system, X C ,Y C ,Z C are respectively the coordinates of the target in the camera coordinate system, T is the translation matrix of the camera coordinate system relative to the world coordinate system, u 0 , v 0 , f x , f y are all variables related to the internal parameters of the video device, f x , f y are respectively the normalized focal lengths on the u-axis and v-axis, u 0 and v 0 are both the intersection points of the video device and the image plane, f x = f / dX, f y = f / dY, f is the focal length of the video device, dX and dY respectively represent the sizes of unit pixels on the u-axis and v-axis of the pixel coordinate system; i1 is the coordinate value of the target on the u-axis in the pixel coordinate system, j1 is the coordinate value of the target on the v-axis in the pixel coordinate system, and D(i1, j1) is the depth value of the point (i1, j1);
[0114] Step S14: Denote the coordinate trajectory to be matched corresponding to the nth feature trajectory α in in the ith video stream as β in , Wherein, is the time - coordinate pair saved in β in , t inl is the time point when the target is detected, β inl = (X Winl ,Y Winl ,Z Winl ), β inl is the three - dimensional space coordinate of the target obtained by solving in the world coordinate system, X Winl ,Y Winl ,Z Winlare the coordinate values of the target on each coordinate axis in the world coordinate system; the obtained feature trajectory α in and coordinate trajectory β in Combine to obtain the comprehensive trajectory of feature coordinates to be matched γ in =(α in ,β in ).
[0115] In the present invention, the rotation matrix and the translation matrix can be calibrated by the Zhang Zhengyou calibration method. u0 and v0 represent the optical center, that is, the intersection of the video device and the image plane, in units of mm / pixel, usually located at the center of the image, so its value is often half of the resolution of the video device, and u0 and v0 are equal.
[0116] In the present invention, whether it is a target sample used to generate a new feature trajectory or a target sample incorporated into an existing trajectory, while storing its feature vector, it is also necessary to store its position coordinates in time series and generate a coordinate trajectory. The target three-dimensional coordinate trajectory is a set of target three-dimensional coordinates containing a time series. Its primary task is to solve the three-dimensional position coordinates of each target in each frame image. The present invention uses the depth information of the depth sensor to solve the target three-dimensional coordinates. To obtain the three-dimensional coordinates of a pixel point based on the depth map, it is necessary to know the coordinates of the pixel point as well as the internal and external parameters of the camera. Figure 3 Explain the transformation relationship between the four major coordinate systems in vision.
[0117] In step S1, the task of single-channel video stream ReID is to detect targets in the single-channel video stream and assign the same ID to the same targets, while generating a temporal feature trajectory that can characterize its typical characteristics and a temporal coordinate trajectory that can characterize its motion characteristics for the targets with the same ID.
[0118] Step S2: The coordinates of the feature to be matched are integrated into the trajectory γ in Compare the feature coordinates of the target obtained by other video devices and generated globally, and determine the confidence of each re-identification ω in,jm ,include:
[0119] For each target feature coordinate integrated trajectory obtained by other video devices and having generated a global identifier, the following operations are performed:
[0120] Step S21: The feature track γ in the feature coordinate integrated track of the target that has been obtained from other video devices and has generated a global mark jm Randomly extract 3 feature vectors from x1,x2,x3 are γ jm The number of the randomly extracted feature vector in;
[0121] From the coordinates of the feature to be matched, the trajectory γin The corresponding α in Randomly extract 3 eigenvectors from Let y1, y2, and y3 be the numbers of the eigenvectors randomly extracted from α in respectively;
[0122] Calculate the eigen-trajectory similarity
[0123] Step S22: Determine γ jm Check whether the time interval corresponding to γ in overlaps with the time interval corresponding to γ. If so, go to Step S23; otherwise, go to Step S25;
[0124] Step S23: Take the overlapping part of the time interval corresponding to γ jm and the time interval corresponding to γ in as the eigen-coordinate composite sub-trajectory, and denote the eigen-coordinate composite sub-trajectory corresponding to γ jm and the eigen-coordinate composite sub-trajectory corresponding to γ in as {β jm(z1) , β jm(z2) , …, β jm(zo)} and {β in(z1) , β in(z2) , …, β in(zo)} respectively;
[0125] Determine the Euclidean distance between the two eigen-coordinate composite sub-trajectories:
[0126]
[0127] Use the hyperbolic tangent function tanh to transform the value range of the Euclidean distance between the two eigen-coordinate composite sub-trajectories into the interval [0, 1], denoted as
[0128] Calculate {β jm(z2) - β jm(z1) , β jm(z3) - β jm(z2) , …, β jm(zo) - β jm(z(o-1))};
[0129] Calculate {β in(z2) - β in(z1) , β in(z3) - β in(z2) , …, β in(zo) - β in(z(o-1))};
[0130] Obtain the coordinate trajectory similarity, denoted as
[0131]
[0132] Among them, β jm(z1) , β jm(z2) , …, β jm(zo) are respectively the position coordinates at each moment in the comprehensive sub-trajectory of the characteristic coordinates corresponding to γ jm , β in(z1) , β in(z2) , …, β in(zo) are respectively the position coordinates at each moment in the comprehensive sub-trajectory of the characteristic coordinates corresponding to γ in . is the Euclidean distance between two comprehensive sub-trajectories of characteristic coordinates, r is an accumulation variable, are respectively the x, y, and z axis coordinate values of β jm(zr) , are respectively the x, y, and z axis coordinate values of β in(zr) ;
[0133] Step S24: Determine the re-identification confidence ω in,jm , where:
[0134]
[0135] η, μ, and λ are respectively weight parameters;
[0136] The comprehensive trajectory processing of the characteristic coordinates of the target obtained by the other road video devices and with a globally unique identifier generated is completed;
[0137] Step S25: Determine the re-identification confidence ω in,jm , where:
[0138]
[0139] The comprehensive trajectory processing of the characteristic coordinates of the target obtained by the other road video devices and with a globally unique identifier generated is completed.
[0140] For each newly generated characteristic / coordinate trajectory or a characteristic / coordinate trajectory that has been updated once in the present invention, it is necessary to match it with the existing characteristic / coordinate trajectories of other roads. If the weighted average of the similarity between the target and the existing target characteristic trajectories of other roads and the similarity of the coordinate trajectories exceeds a set threshold (configurable and adapted according to different environments), it is considered that the two target screenshots are of the same target taken from different angles, and the same global ID is assigned to both; if the weighted average does not exceed the set threshold, a new global ID is generated for this target.
[0141] For example Figure 4As shown, for each newly generated trajectory or updated trajectory, the re-identification confidence needs to be calculated with all trajectories from other video streams according to the above method. If the re-identification confidence ω of two target trajectories exceeds the threshold (configurable), it is initially considered that the two target trajectories come from the same target, and the global ID of the target to be matched is assigned to the target to be matched, so that the two have the same global ID. When this global ID originally corresponds to multiple target trajectories, it is necessary to calculate the re-identification confidence ω between the target trajectory to be matched and all the targets corresponding to this global ID respectively, and calculate the average value of the obtained re-identification confidence If exceeds the set threshold can this global ID be assigned to the target to be matched. If there are multiple globally matching IDs, take all ω and ω in max of the global ID. If the trajectory to be matched cannot find a trajectory that matches it and meets the confidence threshold, when the trajectory to be matched is a newly generated trajectory, a new global ID is assigned to it; when the curve to be matched is not a newly generated trajectory, first judge whether its original global ID corresponds to other trajectories. If not, the original global ID is used for it. If so, it proves that there may be a mis-match in the front, correct the mis-match, and assign a new global ID to it.
[0142] The present invention adopts a binary classification quality evaluation model, defines the target that is not occluded and in a typical state (standing state for a personnel target) as a qualified sample, and defines the screenshot of the target that is occluded or in an atypical state as an unqualified sample. When saving the feature trajectory of the target sample, filter the features of the unqualified sample to prevent some features that cannot represent the typical state of the sample from being mixed in the feature trajectory, ensure the typicality of the feature trajectory, and improve the robustness of the comparison of the target feature trajectory; use a depth sensor to collect data, provide depth information for the video image, so as to solve the position coordinates of the target in the given world coordinate system, continuously save the position coordinates in time series, form the coordinate trajectory of the target, and use the balanced Euclidean distance between the coordinate trajectories of the targets and the cosine similarity of the motion vector trajectories as weights, and use spatial physical information to prevent different targets with similar appearances from being recognized as the same target.
[0143] According to the actual applicable scenario of the present technology, the following assumptions are made for the scenario: In a venue about the size of half a basketball court, multiple video acquisition devices are installed around the venue, and depth acquisition devices are equipped for the video acquisition devices. The depth acquisition devices and the video acquisition devices are pre-calibrated for coordinates, and depth information can be provided for the pixel points of the image through coordinate transformation (the depth information is relatively sparse compared to the pixels, and the depth of the nearest depth point to the pixel point is taken as the depth of the pixel point); there are multiple targets (taking personnel targets as an example) moving back and forth in the venue, and in some perspectives, there may be mutual occlusion between the targets.
[0144] The key of this technology is to fuse the targets after multi-video source recognition, complete deduplication, merging, consistency verification, and achieve the effect of the same target with the same ID. The cross-border target re-identification (ReID) algorithm of the present invention mainly consists of two stages: single-channel ReID and multi-channel matching.
[0145] The present invention provides a cross-border re-identification device based on trajectory matching, and the device includes:
[0146] An acquisition module: configured to acquire the to-be-matched feature trajectory α corresponding to the video stream obtained by the i-th video device in and generate the to-be-matched coordinate trajectory β corresponding to the to-be-matched feature trajectory α in and generate the to-be-matched feature coordinate comprehensive trajectory γ based on the to-be-matched feature trajectory α in and the to-be-matched coordinate trajectory β in ; where n is the n-th newly detected target in the video stream obtained by the i-th video device; in in in in,jm
[0147] A matching module: configured to compare the to-be-matched feature coordinate comprehensive trajectory γ one by one with the feature coordinate comprehensive trajectories of the targets obtained by other video devices and having generated global identifiers, and determine each re-identification confidence ω in where each video stream corresponds to one or more targets; j represents the j-th video device, m is the m-th target obtained by analyzing the j-th video stream, and each video device corresponds to one video stream; in,jm in,jm
[0148] An identification module: configured to, when there is one or more re-identification confidences ω in,jm greater than or equal to a preset threshold, obtain the maximum value of the re-identification confidence ω in,jm and denote this maximum value as ω in,j'm' and obtain the global identifier corresponding to the m'-th target obtained by analyzing the video stream corresponding to the road number j' of the video device corresponding to ω in,j'm' and use this global identifier as the global identifier corresponding to the to-be-matched feature trajectory γ in ;
[0149] When all the re-identification confidences ω in,jm are less than the preset threshold:
[0150] When the to-be-matched feature coordinate comprehensive trajectory γ in is a continuous trajectory of the first feature coordinate comprehensive trajectory, and the global identifier of the first feature coordinate comprehensive trajectory corresponds to other feature coordinate comprehensive trajectories except the first feature coordinate comprehensive trajectory, for the to-be-matched feature coordinate comprehensive trajectory γ inAssign a new global identifier to the combined trajectory of the first feature coordinates; the continuous trajectory refers to the combined trajectory of feature coordinates generated by the same target that generated the combined trajectory of the first feature coordinates at a later time after the time when the combined trajectory of the first feature coordinates was generated;
[0151] When the combined trajectory of the feature coordinates to be matched γ in is a continuous trajectory of the combined trajectory of the first feature coordinates, and the global identifier of the combined trajectory of the first feature coordinates only corresponds to the combined trajectory of the first feature coordinates, assign the global identifier of the combined trajectory of the first feature coordinates to the combined trajectory of the feature coordinates to be matched γ in ;
[0152] When the combined trajectory of the feature coordinates to be matched γ in is not a continuous trajectory of any combined trajectory of feature coordinates, assign a new global identifier to the combined trajectory of the feature coordinates to be matched.
[0153] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-border re-identification method based on trajectory matching, characterized in that: The method comprises the following steps: Step S1: Obtain the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in , generate the feature trajectory α to be matched in The corresponding coordinate trajectory to be matched β in , based on the feature trajectory to be matched α in And the coordinate trajectory to be matched β in , generate the comprehensive trajectory of the feature coordinates to be matched γ in ; Where n is the nth target newly detected in the video stream obtained by the i-th video device, denoted as target n; Step S2: The coordinates of the feature to be matched are integrated into the trajectory γ in Compare the feature coordinates of the target obtained by other video devices and generated globally, and determine the confidence of each re-identification ω in,jm , where each video stream corresponds to one or more targets; j represents the j-th video device, m is the m-th target obtained by analyzing the j-th video stream, and each video device corresponds to one video stream; Step S3: When there are one or more re-identification confidences ω in,jm When it is greater than or equal to the preset threshold, the re-identification confidence ω is obtained. in,jm The maximum value is denoted as ω in,j'm' , get ω in,j'm' The number j' of the corresponding video device and the global identifier corresponding to the m'th target obtained by analyzing the corresponding video stream are used as the feature trajectory to be matched γ in The corresponding global identifier; When all re-identification confidence ω in,jm When both are less than the preset threshold: When the coordinates of the feature to be matched are integrated into the trajectory γ in When it is a continuation track of the first feature coordinate integrated track, and the global identifier of the first feature coordinate integrated track corresponds to other feature coordinate integrated tracks except the first feature coordinate integrated track, it is the feature coordinate integrated track to be matched γ in and assigning a new global identifier to the first feature coordinate integrated trajectory; the continued trajectory refers to: the feature coordinate integrated trajectory generated by the same target that generates the first feature coordinate integrated trajectory at a later time than the time when the first feature coordinate integrated trajectory is generated; When the coordinates of the feature to be matched are integrated into the trajectory γ in When it is the continuation track of the first feature coordinate integrated track, and the global identifier of the first feature coordinate integrated track only corresponds to the first feature coordinate integrated track, it is the feature coordinate integrated track to be matched γ in assigning a global identifier of the first characteristic coordinate integrated trajectory; When the coordinates of the feature to be matched are integrated into the trajectory γ in When it is not a continuation track of any feature coordinate integrated track, a new global identifier is allocated to the feature coordinate integrated track to be matched.
2. The method according to claim 1, characterized in that In step S1, the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device is obtained. in ,include: Step S11: performing target detection on the video stream of the i-th video device to determine a newly detected target; Step S12: For each newly detected target n, the following operations are performed: The feature extraction network model is used to extract RelD features from the video stream of the i-th video device to obtain the visual feature vector of target n. The visual feature vector Matching with the characteristic tracks of the 1st to n-1th targets determined by the i-th video device; If there is a target matching target n among the 1st to n-1th targets determined by the i-th video device, the local identification of target n in the i-th video device is recorded as the local identification of the target matching target n in the i-th video device; obtain the feature track trace' corresponding to the target matching target n in , the visual feature vector Merge to feature track trace' in The tail of the updated feature trajectory trace' in ; The updated feature trajectory is used as the feature trajectory to be matched corresponding to the video stream obtained by the i-th video device α in ; If there is no target matching target n among the targets 1 to n-1 determined by the i-th video device, then the visual feature vector To verify; The verification method is: obtain several frames after the i-th video device detects the target n, if the target in at least three frames matches the target n, then the visual feature vector The verification is passed; otherwise, the verification fails; After verification, obtain the visual feature vectors corresponding to the video frames matching the target n; and arrange the visual feature vectors in chronological order. The visual feature vectors corresponding to the video frames matching the target n are arranged in chronological order to form a to-be-matched feature trajectory α corresponding to the video stream obtained by the i-th video device. in .
3. The method according to claim 2, characterized in that The visual feature vector Matching with the characteristic tracks of the 1st to n-1th targets determined by the i-th video device includes: Calculate visual feature vector The cosine distance, Euclidean distance and intersection-over-union ratio of the feature trajectories of the 1st to n-1th targets determined by the i-th video device are as follows: When the cosine distance with a target from the 1st to the n-1th target determined by the i-th video device is greater than the first threshold, the Euclidean distance is less than the second threshold, and the intersection and union ratio is greater than the third threshold, the visual feature vector is determined. Matches this target.
4. The method according to claim 3, characterized in that Step S1: generating a feature trajectory α to be matched in The corresponding coordinate trajectory to be matched β in , based on the feature trajectory to be matched α in And the coordinate trajectory to be matched β in , generate the comprehensive trajectory of the feature coordinates to be matched γ in ,include: Step S13: matching the feature trajectory α in For each visual feature vector in the image, determine the three-dimensional space coordinates (X W ,Y W ,Z W ),in: (X W ,Y W ,Z W )=R(X C ,Y C ,Z C )+T Among them, X W ,Y W ,Z W are the coordinates of the target corresponding to the visual feature vector on the X-axis, Y-axis, and Z-axis in the world coordinate system, R is the rotation matrix of the camera coordinate system relative to the world coordinate system, X C ,Y C, Z C are the coordinates of the target in the camera coordinate system, T is the translation matrix of the camera coordinate system relative to the world coordinate system, u0, v0, f x 、f y are variables related to the internal parameters of the video device, f x 、f y are the normalized focal lengths on the u-axis and v-axis respectively, u0 and v0 are the intersection points of the video device and the image plane, f x =f / dX,f y =f / dY, where f is the focal length of the video device, dX and dY represent the size of the unit pixel on the u-axis and v-axis of the pixel coordinate system respectively; i1 is the coordinate value of the target on the u-axis in the pixel coordinate system, j1 is the coordinate value of the target on the v-axis in the pixel coordinate system, and D(i1,j1) is the depth value of the point (i1,j1); Step S14: The nth feature trajectory α in the i-th video stream in The corresponding coordinate trajectory to be matched is recorded as β in , in, β in The time-coordinate pairs stored in t inl is the time point of target detection, β inl =(X Winl ,Y Winl ,Z Winl ), β inl To solve the three-dimensional space coordinates of the target in the world coordinate system, X Winl ,Y Winl ,Z Winl are the coordinate values of the target on each coordinate axis in the world coordinate system; the obtained feature trajectory α in and coordinate trajectory β in Combine to obtain the comprehensive trajectory of feature coordinates to be matched γ in =(α in ,β in ).
5. The method according to claim 4, characterized in that Step S2: The coordinates of the feature to be matched are integrated into the trajectory γ in Compare the feature coordinates of the target obtained by other video devices and generated globally, and determine the confidence of each re-identification ω in,jm ,include: For each target feature coordinate integrated trajectory obtained by other video devices and having generated a global identifier, the following operations are performed: Step S21: The feature track γ in the feature coordinate integrated track of the target that has been obtained from other video devices and has generated a global mark jm Randomly extract 3 feature vectors from x 1, x 2, x 3 are γ jm The number of the randomly extracted feature vector in; From the coordinates of the feature to be matched, the trajectory γ in The corresponding α in Randomly extract 3 feature vectors from y1, y2, y3 are α in The number of the randomly extracted feature vector in; Step S22: Determine γ jm The corresponding time interval and γ in Whether the corresponding time intervals overlap, if so, proceed to step S23; otherwise, proceed to step S25; Step S23: Set γ jm The corresponding time interval and γ in The overlapping parts of the corresponding time intervals are taken as the feature coordinate comprehensive sub-trajectory, γ jm The corresponding characteristic coordinates integrated sub-trajectory and γ in The corresponding feature coordinate comprehensive sub-trajectories are recorded as {β jm(z1) ,β jm(z2) ,…,β jm(zo) } and {β in(z1) ,β in(z2) ,…,β in(zo) }; Determine the Euclidean distance between two feature coordinate synthesis subtrajectories: The hyperbolic tangent function tanh is used to convert the value interval of the Euclidean distance between the two feature coordinate comprehensive sub-trajectories into the [0,1] interval, which is recorded as calculation {b jm(z2) -b jm(z1) ,b jm(z3) -b jm(z2) ,…,b jm(zo) -b jm(z(o-1)) }; calculation {b in(z2) -b in(z1) ,b in(z3) -b in(z2) ,…,b in(zo) -b in(z(o-1)) }; Calculate the similarity of coordinate trajectories, denoted as Among them, β jm(z1) ,β jm(z2) ,…,β jm(zo) They are γ jm The corresponding characteristic coordinates are the position coordinates at each moment in the comprehensive sub-trajectory, β in(z1) ,β in(z2) ,…,β in(zo) They are γ in The corresponding characteristic coordinates are the position coordinates at each moment in the comprehensive sub-trajectory, is the Euclidean distance between the two feature coordinate integrated sub-trajectories, r is the cumulative variable, They are β jm(zr) The x, y, and z axis coordinate values, They are β in(zr) The x, y, and z axis coordinate values; Step S24: Determine the re-identification confidence ω in,jm ,in: η, μ, and λ are weight parameters respectively; The processing of the comprehensive trajectory of characteristic coordinates of the target obtained by the other video devices and having generated a global identifier is completed; Step S25: Determine the re-identification confidence ω in,jm ,in: The processing of the comprehensive trajectory of characteristic coordinates of the target obtained by the other video devices and having generated a global identification is completed.
6. A cross-border re-identification device based on trajectory matching, characterized in that: The device comprises: Acquisition module: configured to obtain the feature trajectory α to be matched corresponding to the video stream obtained by the i-th video device in , generate the feature trajectory α to be matched in The corresponding coordinate trajectory to be matched β in , based on the feature trajectory to be matched α in And the coordinate trajectory to be matched β in , generate the comprehensive trajectory of the feature coordinates to be matched γ in ; Where n is the nth target newly detected in the video stream obtained by the i-th video device, denoted as target n; Matching module: configured to integrate the coordinate trajectory of the feature to be matched γ in Compare the feature coordinates of the target obtained by other video devices and generated globally, and determine the confidence of each re-identification ω in,jm , where each video stream corresponds to one or more targets; j represents the j-th video device, m is the m-th target obtained by analyzing the j-th video stream, and each video device corresponds to one video stream; Identification module: configured to be used when there are one or more re-identification confidences ω in,jm When it is greater than or equal to the preset threshold, the re-identification confidence ω is obtained. in,jm The maximum value is denoted as ω in,j'm' , get ω in,j'm' The number j' of the corresponding video device and the global identifier corresponding to the m'th target obtained by analyzing the corresponding video stream are used as the feature trajectory to be matched γ in The corresponding global identifier; When all re-identification confidence ω in,jm When both are less than the preset threshold: When the coordinates of the feature to be matched are integrated into the trajectory γ in When it is a continuation track of the first feature coordinate integrated track, and the global identifier of the first feature coordinate integrated track corresponds to other feature coordinate integrated tracks except the first feature coordinate integrated track, it is the feature coordinate integrated track to be matched γ in and assigning a new global identifier to the first feature coordinate integrated trajectory; the continued trajectory refers to: the feature coordinate integrated trajectory generated by the same target that generates the first feature coordinate integrated trajectory at a later time than the time when the first feature coordinate integrated trajectory is generated; When the coordinates of the feature to be matched are integrated into the trajectory γ in When it is the continuation track of the first feature coordinate integrated track, and the global identifier of the first feature coordinate integrated track only corresponds to the first feature coordinate integrated track, it is the feature coordinate integrated track to be matched γ in assigning a global identifier of the first characteristic coordinate integrated trajectory; When the coordinates of the feature to be matched are integrated into the trajectory γ in When it is not a continuation track of any feature coordinate integrated track, a new global identifier is allocated to the feature coordinate integrated track to be matched.
7. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 5.
Citation Information
Patent Citations
Cross-border head trajectory tracking method and device and storage medium
CN113689475A
Cross-domain vehicle re-identification and continuous track construction method
CN115205559A
Intersection multi-target cross-domain tracking method based on overlapped view
CN116894855A
Multi-camera multi-target tracking method and device, storage medium and electronic equipment
CN117670939A
Cited By
Object re-identification method, device, electronic equipment and system
CN121366279A
An object re-identification method, device, electronic equipment and system
CN121366279B