A video-based object association method, apparatus, and readable storage medium
By detecting the current frame image in the monitoring video and matching the target in the historical frame image, the problem that the tracking system in the prior art is unable to perform target category analysis and behavior analysis, and intelligent target correlation and tracking are achieved.
Patent Information
- Application Number
- CN202111296373.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing tracking system can only achieve single tracking positioning, and cannot perform category analysis or behavioral analysis of targets, and lacks intelligent detection.
By acquiring the current frame image of the monitoring video, the target to be tracked is detected, the target to be matched in the historical frame image is determined, and the target correlation and tracking is performed based on the similarity of the target detection information.
Intelligent detection and analysis of targets and target correlation are achieved, improving the accuracy and intelligence of target tracking.
Smart Images

Figure CN114219828B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a video-based target association method, apparatus, and readable storage medium. Background Art
[0002] Currently, the application scenarios of tracking systems are becoming increasingly widespread. In the military field, it can be applied to battlefield detection or target tracking; in the security field, it can be used for crime warning or suspect location; in the urban traffic field, it can be applied to traffic violation detection or autonomous driving, etc. Therefore, tracking systems have become a current research hotspot. However, the current tracking systems can only achieve single tracking and positioning, and cannot perform intelligent detections such as target category analysis or behavior analysis. Summary of the Invention
[0003] This application provides a video-based target association method, apparatus, and readable storage medium, which can perform target association in real time and intelligently.
[0004] To solve the above technical problems, the technical solution adopted by this application is: to provide a video-based target association method, which includes: obtaining the current frame image of a surveillance video; performing target detection on the target to be tracked in the current frame image to obtain the first target detection information of the target to be tracked; determining each target to be matched in the historical frame image based on the position information of the target to be tracked in the current frame image; the targets to be matched include the targets within the candidate matching region in the historical frame image, and the candidate matching region is determined based on the position information of the target to be tracked; the historical frame image is the image before the current frame image in the surveillance video; determining the target to be matched associated with the target to be tracked based on the similarity between the second target detection information of each target to be matched and the first target detection information.
[0005] To solve the above technical problems, another technical solution adopted by this application is: to provide a real-time target tracking device, which includes a memory and a processor connected to each other. Among them, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the video-based target association method in the above technical solution.
[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, which is used to store a computer program, and when the computer program is executed by a processor, it is used to implement the video-based target association method in the above technical solution.
[0007] Through the above solution, the beneficial effects of the present application are as follows: First, obtain the current frame image of the surveillance video, then perform object detection on the target to be tracked in the current frame image to obtain the first object detection information of the target to be tracked. Then, based on the position information of the target to be tracked in the current frame image, determine each target to be matched in the historical frame image. Finally, based on the similarity between the second object detection information of each target to be matched and the first object detection information, determine the target to be matched associated with the target to be tracked; through this solution, the object detection information of the target to be tracked can be obtained, realizing intelligent detection and analysis of the target. At the same time, the object detection information (including the first object detection information and the second object detection information) can be used to perform object matching on the target to be tracked in the historical frame image to achieve object association, and then complete object tracking; in addition, by combining object detection analysis and object tracking, the accuracy of object tracking can be further improved and it is also more intelligent. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0009] Figure 1 is a schematic flowchart of an embodiment of the object association method based on video provided by the present application;
[0010] Figure 2 is a schematic flowchart of another embodiment of the object association method based on video provided by the present application;
[0011] Figure 3 is a schematic flowchart of the process of restricting the object detection frame of the non-key frame image provided by the present application;
[0012] Figure 4 is a schematic flowchart of step 27 provided by the present application;
[0013] Figure 5 is a schematic structural diagram of an embodiment of the real-time object tracking device provided by the present application;
[0014] Figure 6 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present application, but do not limit the scope of the present application. Similarly, the following embodiments are only partial embodiments of the present application rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0016] In the present application, the mention of "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0017] It should be noted that the terms "first", "second", and "third" in the present application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0018] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the video-based target association method provided by the present application. The method includes:
[0019] Step 11: Obtain the current frame image in the surveillance video.
[0020] The surveillance area can be monitored by a gun camera, a dome camera, a panoramic camera or other camera devices to obtain the surveillance video. Specifically, the above camera devices can be installed in a top-mounted or oblique-mounted manner. The oblique-mounted manner of the camera device can be selected when the installation height is greater than 2.5 meters, and the top-mounted manner of the camera device can be selected when the installation height is less than 2.5 meters, so as to ensure that the surveillance images obtained by the camera device are complete for subsequent tracking of the targets in the current frame image.
[0021] Step 12: Perform object detection on the target to be tracked in the current frame image to obtain the first object detection information of the target to be tracked.
[0022] The target to be tracked can be a human object, that is, the current frame image obtained from the surveillance video can contain a human object, and then object detection is performed on the human object in the current frame image; specifically, the target to be tracked in the current frame image can be one, two, three or more than three. Object detection is performed on all targets to be tracked in the current frame image to obtain the first object detection information corresponding to all targets to be tracked; it can be understood that the current frame image can be input into the object detection model first, so as to use the object detection model to perform object detection, obtain the object detection boxes corresponding to all targets to be tracked, and then perform detection processing on the object detection boxes to obtain the first object detection information of the target to be tracked, where the object detection model can be a YOLOX model or a CentNet model, etc.
[0023] In a specific embodiment, the user can set a detection area in the current frame image to specify the range of object detection, that is, when performing object detection, only the objects in the pre-set detection area are detected; for example: when the camera device monitors the hotel lobby, a surveillance video containing the hotel front desk, the hotel entrance, and the hotel waiting area can be obtained. If the user only wants to track the pedestrians passing by the hotel entrance, the area of the hotel entrance can be specified as the detection area at this time. Subsequent object detection and object tracking are performed in this detection area, and no object detection and tracking are performed on other areas except the detection area, so as to perform object detection and tracking in a targeted manner, save unnecessary workload, save computer resources, and further improve work efficiency.
[0024] Step 13: Based on the position information of the target to be tracked in the current frame image, determine each target to be matched in the historical frame image.
[0025] The historical frame image is the image before the current frame image in the surveillance video. Generally, the historical frame image can be the previous frame image of the current frame image, but when a situation similar to frame skipping occurs, the historical frame image can be several previous frame images of the current frame image; specifically, the target to be matched includes the target in the candidate matching area in the historical frame image, and the candidate matching area is determined based on the position information of the target to be tracked.
[0026] In a specific embodiment, the candidate matching region can be determined based on the current position information of the target to be tracked in the current frame image. It can, but is not limited to, determining the range of a circle / polygon centered on the current position of the target to be tracked in the current frame image as the above-mentioned candidate matching region. The targets within the candidate matching region in the historical frame image are regarded as the targets to be matched, so as to perform target matching between the targets to be matched and the target to be tracked in the current frame image. Since the position change of the same target in two consecutive frame images is not very large, screening the targets to be matched in the historical frame image can greatly reduce the time consumption of target matching while not affecting the target matching accuracy, and improve the efficiency of target matching. It can be understood that the diameter size of the candidate matching region can be set according to the actual situation.
[0027] Step 14: Determine the target to be matched associated with the target to be tracked based on the similarity between the second target detection information and the first target detection information of each target to be matched.
[0028] The similarity between the second target detection information and the first target detection information of each target to be matched can be compared to obtain the similarity between the second target detection information and the first target detection information of each target to be matched, so as to match the target to be tracked with the target to be matched in the historical frame image. If the target to be tracked is successfully matched with the target to be matched, that is, the same target in the current frame image and the historical frame image is matched, the tracking result of the target to be tracked can be obtained. The tracking result may include the second target detection information of the historical frame image and the first target detection information of the current frame image. Specifically, when there are multiple targets to be tracked in the current frame image, multiple targets to be tracked can be matched simultaneously to obtain their respective tracking results.
[0029] In a specific embodiment, the first target detection information may include the category information and key point information of the target to be tracked. The category information is the information of the category of the target to be tracked, and the category information may include information such as the age or gender of the target to be tracked. The age information can be a specific value or a range. The key point information is the information of the key points of the target to be tracked. The key points of the target to be tracked (i.e., the human object) correspond to the joints or parts of the human body. For example: head, elbow joint, shoulder joint or knee, etc. Generally speaking, when performing key point detection processing on a certain human object, 15 - 17 human key points can be output. It can be understood that the second target detection information may include the category information and key point information of the target to be matched.
[0030] In this embodiment, first obtain the current frame image of the surveillance video, then perform object detection on the target to be tracked in the current frame image to obtain the first object detection information of the target to be tracked, and then determine each target to be matched in the historical frame image based on the position information of the target to be tracked in the current frame image. Finally, based on the similarity between the second object detection information of each target to be matched and the first object detection information, determine the target to be matched associated with the target to be tracked; through this solution, the object detection information of the target to be tracked can be obtained, realizing intelligent detection and analysis of the target, and at the same time, the object detection information can be used to perform object matching on the target to be tracked in the historical frame image to achieve object association, thereby completing object tracking; in addition, by combining object detection analysis and object tracking, the accuracy of object tracking can be further improved and it is also more intelligent.
[0031] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the video-based object association method provided by this application. The method includes:
[0032] Step 21: Obtain the current frame image in the surveillance video.
[0033] This step 21 is the same as step 11 in the above embodiment and will not be elaborated here.
[0034] Step 22: Determine whether the current frame image is a key frame image.
[0035] Key frame images can be set in the surveillance video, and some frame images of the surveillance video can be selected as key frame images, and the remaining images are non-key frame images; during object detection, only the key frame images are processed for object detection, and the object detection of non-key frame images is omitted, thereby saving the time of object detection and improving work efficiency. It can be understood that there is only a difference in the processing of key frame images and non-key frame images in the object detection step, and both key frame images and non-key frame images are executed in subsequent classification and key point detection operations.
[0036] In a specific embodiment, it can be determined whether the current frame image is a key frame image by judging whether the frame number of the current frame image is within a preset numerical set; if the frame number of the current frame image is within the preset numerical set, it is determined that the current frame image is a key frame image; further, the preset numerical set can be selected and set by the user. For example, if the user selects and sets that starting from the first frame image, every four frames is a key frame image, then the preset numerical set can include 1, 6, 11, 16, etc. and so on; it can be understood that in order to ensure that the non-key frame images can set the object detection frame according to the object detection frame of the historical frame image, no matter how many frames are set as the interval between key frame images by the user, the first frame image in the surveillance video is a key frame image, that is, the preset numerical set must include 1.
[0037] Further, in addition to determining whether the current frame image is a key frame image according to the preset numerical set as described above, it can also be determined according to the real-time target tracking result. Denote the current frame as the Nth frame. If a new target appears in the (N - 1)th frame image compared with the (N - 2)th frame image, or if an old target is lost in the (N - 1)th frame image compared with the (N - 2)th frame image, then the Nth image is determined as a key frame image at this time, so as to ensure that new targets can be accurately detected in a timely manner. For example: only targets A, B, and C are included in the (N - 2)th frame image, and a new target D appears in the (N - 1)th image. Then, in order to accurately detect target D, the Nth frame image is marked as a key frame image, and target detection is performed on this key frame image.
[0038] Step 23: If the current frame image is a key frame image, use the target detection model to process the current frame image to obtain a target detection box.
[0039] When the current frame image is a key frame image, use the target detection model to process the current frame image to obtain target detection boxes corresponding to all the targets to be tracked in the current frame image. The target detection model can be a YOLOX model or a CentNet model, etc.
[0040] Step 24: If the current frame image is a non-key frame image, determine the target detection box based on the positions of the targets to be matched in the historical frame images.
[0041] When the current frame image is a non-key frame image, the target detection boxes of each target to be tracked can be determined according to the positions of the targets to be matched in the historical frame images. Specifically, find the target to be matched with the closest position to the current target to be tracked in the historical frame image, and then expand the target detection box of this target to be matched by a preset multiple (such as 1.2 - 2 times), that is, expand the length and width of the target detection box by the same multiple, and use the expanded detection box as the target detection box of the current target to be tracked.
[0042] It can be understood that when there are multiple targets to be matched in the historical frame image that are all very close to the position of the current target to be tracked and it is impossible to determine which one is the closest, the target detection boxes of multiple targets to be matched can be expanded by the same multiple, and then the union area of all the expanded target detection boxes is used as the target detection box, and then this target detection box is processed to obtain the first target detection information including the category information and key point information of the target to be tracked.
[0043] Step 25: Use the classification model to process the target detection box to obtain the category information of the target to be matched and the category information of the target to be tracked.
[0044] Processing the target detection box using a classification model can obtain the category information of the target to be matched and the category information of the target to be tracked. Specifically, when the number of targets to be tracked is multiple, the target detection boxes of all targets to be tracked can be input into the classification model simultaneously to obtain the category information of all targets to be tracked, and this category information can include information such as age or gender.
[0045] Step 26: Process the target detection box using a key point detection model to obtain the key point information of the target to be matched and the key point information of the target to be tracked.
[0046] Processing the target detection box using a key point detection model can obtain the key point information of the target to be matched and the key point information of the target to be tracked. Specifically, when the number of targets to be tracked is multiple, the target detection boxes of all targets to be tracked can be input into the key point detection model simultaneously to obtain the key point information of all targets to be tracked. Further, the key point information includes all key points of the target to be tracked, and the key points correspond to the joints or parts of the human body, such as: head, elbow joint, shoulder joint or knee, etc.
[0047] In a specific embodiment, the target detection information further includes key point position information. When the current frame image is a non-key frame image, after obtaining the category information and key point information of the target to be tracked, the target detection box can be restricted based on the key point position information to obtain a more accurate target detection box, so that when the next frame image is still a non-key frame image, the target detection box can be generated according to the restricted target detection box of the current frame image, and at the same time, it also provides a basis for subsequent similarity calculation.
[0048] Specifically, the position detection model can be used to perform position detection processing on the key point information of the target to be tracked and the key point information of the target to be matched respectively to obtain the key point position information of the target to be tracked and the key point position information of the target to be matched. This key point position information can include the position coordinates of each key point, and the position coordinates include the abscissa value and the ordinate value. For example, the position coordinates of a certain key point are (2, 5), indicating that its abscissa value is 2 and its ordinate value is 5. Further, the position detection model can be a Graph Convolutional Networks (GCNs). That is, the key point information can be subjected to position detection through the graph convolutional neural network to obtain the key point position information, and then the target detection box can be restricted based on the key point position information. The specific steps are as Figure 3 shown:
[0049] Step 31: Obtain the key point with the minimum ordinate value and the key point with the maximum ordinate value among all key points based on the key point position information.
[0050] Find the key point with the smallest y - coordinate value and the key point with the largest y - coordinate value based on the position coordinates of each key point. For example, if the target to be tracked is in a normal standing state, then the key point with the smallest y - coordinate value can be the key point located at the foot, and the key point with the largest y - coordinate value can be the key point located at the head.
[0051] Step 32: Based on the key point position information, obtain the key point with the smallest x - coordinate value and the key point with the largest x - coordinate value among all key points.
[0052] When the target to be tracked is in a standing state with hands in pockets, the key point with the smallest x - coordinate value at this time is the key point located at one elbow joint, and the key point with the largest x - coordinate value is the key point located at the other elbow joint.
[0053] Step 33: Obtain the bounding rectangle of the key point with the smallest y - coordinate value, the key point with the largest y - coordinate value, the key point with the smallest x - coordinate value, and the key point with the largest x - coordinate value, and get the target detection box of the current frame image.
[0054] By obtaining the bounding rectangle of the key point with the smallest y - coordinate value, the key point with the largest y - coordinate value, the key point with the smallest x - coordinate value, and the key point with the largest x - coordinate value, a smallest rectangle containing all key points in the target to be tracked can be formed. Taking this rectangle as the target detection box of the target to be tracked, the limitation of the target detection box can be completed, and a more accurate target detection box can be obtained.
[0055] Step 27: Based on the category information of the target to be matched, the key point information of the target to be matched, the category information of the target to be tracked, and the key point information of the target to be tracked, compare the similarity between the target to be tracked and each target to be matched to obtain the corresponding similarity.
[0056] After obtaining the category information and key point information of the target to be matched and the target to be tracked by using the classification model and the key point detection model respectively, compare the similarity between the target to be tracked and each target to be matched through the category information and key point information, and obtain the corresponding similarity between the target to be tracked and each target to be matched; it can be understood that the similarity is used to represent the matching degree between the target to be tracked and the target to be matched. The higher the similarity, the higher the matching degree between the target to be tracked and the target to be matched.
[0057] The following details the steps of comparing the similarity between the target to be tracked and each target to be matched based on the category information of the target to be matched, the key point information of the target to be matched, the category information of the target to be tracked, and the key point information of the target to be tracked to obtain the corresponding similarity, specifically as Figure 4 shown:
[0058] Step 41: Based on the category information of the target to be matched and the category information of the target to be tracked, compare the degree of matching between the target to be tracked and each target to be matched to obtain the category similarity.
[0059] The category information includes the category of the target to be tracked and the feature vector corresponding to the category. Calculate the deviation between the feature vector of the target to be tracked and the feature vector of the target to be matched to obtain a deviation value, and then match the deviation value with a preset deviation scoring table to obtain the category similarity. Specifically, the preset deviation scoring table can be set according to the actual situation, which can include the deviation value and the corresponding category similarity. Matching the calculated deviation value with the preset deviation scoring table can obtain the corresponding category similarity.
[0060] It can be understood that the category information includes multiple information such as age or gender. Then, calculate the deviation of the corresponding feature vectors of the target to be tracked and the target to be matched under the same category respectively, and then obtain the sub-deviation values and corresponding sub-category similarities corresponding to each category. At this time, in order to obtain the category similarity that can represent the overall category matching degree of the target to be tracked and the target to be matched, sum all the sub-category similarities and then calculate the average value, and use the calculated average value as the final category similarity. For example: calculate the deviation between the age feature vectors of the target to be tracked and the target to be matched to obtain a deviation value, calculate the deviation between the gender feature vectors of the target to be tracked and the target to be matched to obtain another deviation value, then match each deviation value with the preset deviation scoring table to obtain the corresponding sub-category similarity, sum all the sub-category similarities, and then calculate the average value to obtain the final category similarity.
[0061] Step 42: Based on the target detection box of the target to be matched and the target detection box of the target to be tracked, compare the positions of the target to be tracked and each target to be matched to obtain the spatial similarity.
[0062] The change of the spatial position between the target to be tracked and the target to be matched can be judged by calculating the intersection over union (IOU) between the target detection box of the target to be tracked and the target detection box of the target to be matched. Calculate the intersection over union (IOU) between the target detection box of the target to be tracked and the target detection box of the target to be matched to obtain the current IOU, and then match the current IOU with a preset spatial scoring table to obtain the spatial similarity. Specifically, dividing the intersection area of the target detection box in the historical frame image and the target detection box in the current frame image by the union area can obtain the current IOU. It can be understood that when the current frame image is a non-key frame image, the target detection box generated after the shrinking process is used to calculate the spatial similarity.
[0063] Further, the preset space scoring table can be set according to the actual situation. It can include the intersection over union (IoU) and the corresponding space similarity. By matching the calculated current IoU with the preset space scoring table, the corresponding space similarity can be obtained.
[0064] Step 43: Based on the key point information of the target to be matched and the key point information of the target to be tracked, perform pose comparison between the target to be tracked and each target to be matched to obtain the pose similarity.
[0065] The target detection information also includes pose information, which is used to represent the pose of the human object and can include poses such as standing pose, sitting pose, or lying pose, etc. The pose comparison in this embodiment can include the following two methods:
[0066] 1) Perform pose comparison processing on the key point position information of the target to be tracked and the key point position information of the target to be matched based on the pose comparison model to obtain the pose information and the pose similarity.
[0067] The pose comparison model can be a Siamese neural network. The specific structure and working principle of the Siamese neural network are the same as those in the related art and will not be elaborated here. The key point position information of the target to be tracked and the key point position information of the target to be matched can be input into the Siamese neural network, so as to obtain the pose information of the target to be tracked and the pose similarity between the target to be tracked and the target to be matched.
[0068] 2) Select the first reference key point from all the key points of the target to be tracked, select the second reference key point from the target to be matched, then align the first reference key point with the second reference key point, calculate the distances between the remaining key points in the target to be tracked and the corresponding remaining key points in the target to be matched to obtain the pose offset value, and then match the pose offset value with the preset pose scoring table to obtain the pose similarity.
[0069] The second reference key point is located at the same part as the first reference key point. Generally, the first reference key point and the second reference key point located at the head can be selected, that is, align the key points of the target to be tracked and the target to be matched located at the head, and then calculate the distances between the remaining key points of the corresponding parts in the target to be tracked and the target to be matched respectively. For example: taking the key points located at the head as the first reference key point and the second reference key point, at this time, the distance between the hand key points of the target to be tracked and the hand key points of the target to be matched can be calculated respectively, the distance between the foot key points of the target to be tracked and the hand key points of the target to be matched can be calculated, etc., without listing them one by one here, so as to obtain the sub-attitude offset values corresponding to the key points of each part. The sum of all sub-attitude offset values can be processed and averaged to obtain the final attitude offset value, and then the attitude offset value is matched with the preset attitude scoring table to obtain the attitude similarity.
[0070] Further, calculation methods such as calculating the Euclidean distance or the cosine distance can be used to calculate the distance between key points, and the calculation method of the distance is not limited here; the preset attitude scoring table can be set according to the actual situation, and it can include the attitude offset value and the corresponding attitude similarity.
[0071] In other embodiments, the above two schemes can also be combined to determine the attitude similarity. For example: perform weighted summation on the attitude similarities obtained by the two schemes as the final attitude similarity.
[0072] Step 44: Generate a similarity based on the category similarity, the spatial similarity, and the attitude similarity.
[0073] Weighted summation processing can be performed on the category similarity, the spatial similarity, and the attitude similarity to obtain the similarity, so as to match the target based on the similarity between the target to be tracked and each target to be matched, thereby determining the tracking result.
[0074] Step 28: Determine whether there is a target to be matched with a similarity greater than the preset scoring threshold among all targets to be matched.
[0075] The preset scoring threshold can be set according to the actual situation. For example: set the preset scoring threshold to 0.8. When the similarity of the target to be matched is greater than 0.8, it means that the target to be matched is relatively well matched with the target to be tracked.
[0076] Step 29: If there is a target to be matched with a similarity greater than the preset scoring threshold among all targets to be matched, then associate the target to be matched with the largest similarity with the target to be tracked.
[0077] When there are multiple target candidates to be matched with a similarity greater than a preset scoring threshold, the target that best matches the target to be tracked can be selected from the multiple target candidates to be matched, and the target candidate with the highest similarity is associated with the target to be tracked, that is, it is determined that the target candidate with the highest similarity and the target to be tracked are successfully matched, so as to obtain the tracking result of the target to be tracked; specifically, the tracking result may include the second target detection information of the historical frame image and the first target detection information of the current frame image. For the tracking result of a certain target to be tracked, it is the first target detection information of the target to be tracked and the second target detection information of the target candidate to be matched in the matched historical frame image. The target detection information may include category information, key point information, key point position information, and pose information, etc.
[0078] In a specific embodiment, the target detection information further includes identity information (Identity, ID). When the target to be tracked and the target candidate to be matched are successfully matched, the identity information of the target candidate to be matched is determined as the identity information of the target to be tracked, that is, the target ID of the target candidate to be matched is assigned to the target to be tracked; when the target to be tracked fails to be successfully matched with the target candidate to be matched, a new target ID is assigned to the target to be tracked; specifically, when the target to be tracked fails to be successfully matched with all target candidates in the historical frame image, it means that the target to be tracked is a newly emerged target, and at this time, a new target ID can be assigned to the target from a pre-set target ID database.
[0079] In this embodiment, according to whether the current frame image is a key frame image, a corresponding method for obtaining the target detection box is selected. The target detection model can be used to perform target detection on the key frame image, or the target detection box of the historical frame image can be used to delimit the target detection box for the non-key frame image, which can save the time of target detection while ensuring the target detection accuracy and improve work efficiency; by processing the target detection box, the corresponding category information and key point information are obtained, and then the target matching is performed using the category information and key point information of the historical frame image and the current frame image. The similarity evaluation is carried out from three aspects: category similarity, spatial similarity, and pose similarity, and the target matching is performed based on the size of the similarity, making the matching result more accurate and the target tracking more precise; moreover, while accurately tracking the target, multi-faceted tracking results can be generated, obtaining target detection information including category information, key point information, target ID, and pose information, etc., realizing the combination of target detection analysis and target tracking, and making the target tracking more intelligent and comprehensive.
[0080] Please refer to Figure 5 , Figure 5FIG. 0 is a schematic structural diagram of an embodiment of a real-time target tracking device provided by the present application. The real-time target tracking device 50 includes a memory 51 and a processor 52 connected to each other. The memory 51 is used to store a computer program, and when the computer program is executed by the processor 52, it is used to implement the video-based target association method in the above embodiment.
[0081] Please refer to Figure 6 , Figure 6 FIG. 7 is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 60 is used to store a computer program 61, and when the computer program 61 is executed by a processor, it is used to implement the video-based target association method in the above embodiment.
[0082] The computer-readable storage medium 60 may be various media that can store program codes, such as a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0083] In several implementation manners provided by the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the device implementation manner described above is only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0084] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this implementation manner.
[0085] In addition, each functional unit in various implementation manners of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0086] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application in the same way.
Claims
1. A video-based object association method, characterized in that Including: Obtain the current frame image of the surveillance video; Perform object detection on the target to be tracked in the current frame image to obtain the first object detection information of the target to be tracked; And Based on the position information of the target to be tracked in the current frame image, determine each target to be matched in the historical frame image; the target to be matched includes the target within the candidate matching region in the historical frame image, the candidate matching region is determined based on the position information of the target to be tracked, and the candidate matching region is a region with a preset geometric shape centered on the current position of the target to be tracked; the historical frame image is the image before the current frame image in the surveillance video; Based on the similarity between the second object detection information of each target to be matched and the first object detection information, determine the target to be matched associated with the target to be tracked; when there are multiple targets to be tracked in the current frame image, match the multiple targets to be tracked simultaneously to obtain their respective targets to be matched; Wherein, the first object detection information includes the category information and key point information of the target to be tracked, the category information includes the category of the target to be tracked and the feature vector corresponding to the category, and the category includes age or gender; the second object detection information includes the category information and key point information of the target to be matched, the key point information contains all the key points of the target to be tracked, and the key points correspond to the joints or parts of the human body; the step of determining the target to be matched associated with the target to be tracked based on the similarity between the second object detection information of each target to be matched and the first object detection information includes: Based on the category information of the target to be matched and the category information of the target to be tracked, compare the matching degree between the target to be tracked and each target to be matched to obtain the category similarity; Based on the object detection box of the target to be matched and the object detection box of the target to be tracked, compare the positions of the target to be tracked and each target to be matched to obtain the spatial similarity; Select a first reference key point from all the key points of the target to be tracked, and select a second reference key point from the target to be matched, and the second reference key point and the first reference key point are located in the same part; align the first reference key point and the second reference key point, and calculate the distances between the remaining key points in the target to be tracked and the corresponding remaining key points in the target to be matched to obtain the pose offset value; match the pose offset value with a preset pose scoring table to obtain the pose similarity; the pose includes standing pose, sitting pose or lying pose, and the pose scoring table includes the pose offset value and the corresponding pose similarity; Perform weighted summation processing on the category similarity, the spatial similarity and the pose similarity to generate the similarity; Based on the similarity, determine the target to be matched associated with the target to be tracked.
2. The video-based target association method according to claim 1, wherein The step of determining the target to be matched associated with the target to be tracked based on the similarity between the second target detection information of each target to be matched and the first target detection information includes: Judging whether there is a target to be matched with a similarity greater than a preset scoring threshold among all the targets to be matched; If so, associating the target to be matched with the largest similarity with the target to be tracked.
3. The video-based target association method according to claim 1, wherein After the step of determining the target to be matched associated with the target to be tracked, it includes: Determining the identity identification information of the target to be matched as the identity identification information of the target to be tracked.
4. The video-based target association method according to claim 1, wherein The step of comparing the matching degree between the target to be tracked and each target to be matched based on the category information of the target to be matched and the category information of the target to be tracked to obtain the category similarity includes: Calculating the deviation between the feature vector of the target to be tracked and the feature vector of the target to be matched to obtain a deviation value; Matching the deviation value with a preset deviation scoring table to obtain the category similarity.
5. The video-based target association method according to claim 1, characterized in that, The step of comparing the positions of the target to be tracked and each target to be matched based on the target detection frame of the target to be matched and the target detection frame of the target to be tracked to obtain the spatial similarity includes: Calculating the intersection over union between the target detection frame of the target to be tracked and the target detection frame of the target to be matched to obtain the current intersection over union; Matching the current intersection over union with a preset spatial scoring table to obtain the spatial similarity.
6. The video-based target association method according to claim 1, wherein The step of comparing the postures of the target to be tracked and each target to be matched based on the key point information of the target to be matched and the key point information of the target to be tracked to obtain the posture similarity includes: Based on a position detection model, respectively performing position detection processing on the key point information of the target to be tracked and the key point information of the target to be matched to obtain the key point position information of the target to be tracked and the key point position information of the target to be matched; Based on a posture comparison model, performing posture comparison processing on the key point position information of the target to be tracked and the key point position information of the target to be matched to obtain the posture information and the posture similarity.
7. The video-based target association method according to claim 1, wherein Before the step of performing target detection on the target to be tracked in the current frame image to obtain the first target detection information of the target to be tracked, it includes: Judging whether the current frame image is a key frame image; If so, processing the current frame image with a target detection model to obtain a target detection frame; If not, determining the target detection frame based on the positions of the targets to be matched in the historical frame images; Processing the target detection frame with a classification model to obtain the category information of the target to be matched and the category information of the target to be tracked; Processing the target detection frame with a key point detection model to obtain the key point information of the target to be matched and the key point information of the target to be tracked.
8. The video-based target association method according to claim 7, wherein The step of judging whether the current frame image is a key frame image includes: Judging whether the frame number of the current frame image is within a preset numerical set; If so, determining that the current frame image is a key frame image.
9. The video-based target association method according to claim 7, wherein The method further includes: When the current frame image is not a key frame image, restricting the target detection box based on the key point position information of the target to be tracked.
10. The video-based target association method according to claim 9, wherein, The key point position information includes the position coordinates of each key point, and the position coordinates include the abscissa value and the ordinate value. The step of restricting the target detection box based on the key point position information includes: Based on the key point position information, obtaining the key point with the smallest ordinate value and the key point with the largest ordinate value among all the key points; Based on the key point position information, obtaining the key point with the smallest abscissa value and the key point with the largest abscissa value among all the key points; Obtaining the circumscribed rectangle of the key point with the smallest ordinate value, the key point with the largest ordinate value, the key point with the smallest abscissa value, and the key point with the largest abscissa value to obtain the target detection box of the current frame image.
11. A real-time target tracking device, characterized in that, It includes a memory and a processor connected to each other. Among them, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the video-based target association method according to any one of claims 1-10.
12. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the video-based target association method according to any one of claims 1-10.
Citation Information
Patent Citations
Target detecting and tracking method and device
CN106846362A
Multi-target tracking method and device, equipment and storage medium
CN109522843A
Main body tracking method and device, electronic equipment and computer readable storage medium
CN110334635A
Face tracking method, device and equipment and storage medium
CN110705478A