Target tracking method, terminal and computer-readable storage medium
Through the binocular camera system switching detection within a small field of view and large field of view, the target position is determined by using feature similarity comparison, which solves the tracking problem of a single surveillance camera when the target is temporarily lost or blocked, and achieves a more stable target tracking effect.
Patent Information
- Application Number
- CN202111320883.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-11-09
AI Technical Summary
In the prior art, single surveillance cameras cannot effectively track when the target is temporarily lost or blocked, resulting in a decrease in tracking real-time and accuracy.
The binocular camera system is used to detect and extract targets for a small field of view by using the first image acquisition device. If no target is detected, switch to the second image acquisition device to detect large field of view, and determine the target position through feature similarity comparison to achieve stable tracking of the target.
Improve the stability of target tracking, avoid target loss, improve the real-time and accuracy of tracking, and eliminate the need for coordinated calibration of monitoring equipment.
Smart Images

Figure CN114219830B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a target tracking method, a terminal, and a computer-readable storage medium. Background Art
[0002] At present, the commonly used gun-ball linkage, thunder-ball linkage, or panoramic close-up tracking in the surveillance field all use radar or fixed cameras to monitor a certain area and use pan-tilt cameras for close-up zoom. The monitoring method based on the linkage of multiple cameras will, to a certain extent, lead to lower real-time tracking performance. At the same time, if the number of monitoring devices in an area increases or decreases, the coordinates of each device need to be recalibrated. The target tracking system based on a single surveillance camera uses a separate pan-tilt monitoring device to identify and track the target. This can avoid the coordinate calibration process and linkage control process caused by the linkage between multiple monitoring devices, and to a certain extent improve the real-time and accuracy of tracking. However, it is less effective in solving common problems such as occlusion in target tracking. During the process of tracking the target, if the target is temporarily lost or obscured, the single surveillance camera cannot achieve tracking, and it is still necessary to use the pan-tilt parameters of the current device to call other monitoring devices to track the target in a linkage manner. Summary of the Invention
[0003] The main technical problem solved by the present invention is to provide a target tracking method, terminal and computer-readable storage medium to solve the problem in the prior art that a single surveillance camera cannot track a target due to temporary loss or obstruction.
[0004] To solve the above technical problems, the first technical solution adopted by the present invention is: to provide a target tracking method, the target tracking method comprising: performing target detection on a first acquired video frame, the first video frame being acquired by a first image acquisition device; in response to the first video frame not detecting a preset first target object, performing target detection on a second acquired video frame, the second video frame being acquired by a second image acquisition device; wherein, the first monitoring area of the first image acquisition device is a subset of the second monitoring area of the second image acquisition device; in response to the second target object being detected in the second video frame, and the second target object being the same as the first target object, determining the position information of the second target object in the second monitoring area as the position information of the first target object.
[0005] The target tracking method includes: in response to detecting a preset first target object in a first video frame, outputting position information of the first target object in the first video frame.
[0006] In which, in response to the first video frame detecting a preset first target object, the position information of the first target object in the first video frame is output, including: determining whether the previous video frame of the first video frame saves the preset feature map of the first target object; if the previous video frame saves the preset feature map, determining that the first video frame is not the first frame image containing the first target object.
[0007] Among them, in response to the first video frame detecting a preset first target object, the position information of the first target object in the first video frame is output, including: detecting the first video frame to obtain a candidate target object and the position information of the candidate target object; extracting features of the candidate target object to obtain a first feature map; judging whether the similarity between the first feature map and the preset feature map of the first target object saved in the previous video frame is greater than a first threshold; if the similarity between the first feature map and the preset feature map is greater than the first threshold, determining that the candidate target object is the first target object; and outputting the position information of the first target object in the first video frame.
[0008] Among them, in response to the first video frame detecting a preset first target object, the position information of the first target object in the first video frame is output, and it also includes: if the previous video frame does not save the preset feature map, it is determined that the first video frame is the first frame image containing the first target object, then the position information of the first target object in the first video frame is output.
[0009] In which, in response to the first video frame not detecting the preset first target object, target detection is performed on the acquired second video frame, including: if the preset first target object is not detected in the first video frame; acquiring at least one second video frame within a preset time period; performing target detection on at least one second video frame; wherein the preset time period includes a time period including the first video frame acquisition moment and the first time period thereafter; or the preset time period includes a time period including the second time period thereafter after the first video frame acquisition moment.
[0010] Among them, in response to detecting a second target object in the second video frame, and the second target object is the same as the first target object, the position information of the second target object in the second monitoring area is determined as the position information of the first target object, which includes: extracting features of the second target object to obtain a second feature map; judging whether the similarity between the second feature map and the preset feature map exceeds a second threshold; in response to detecting a second target object in the second video frame, and the second target object is the same as the first target object, the position information of the second target object in the second monitoring area is determined as the position information of the first target object, including: if the similarity between the second feature map and the preset feature map exceeds the second threshold, the position information of the second target object in the second monitoring area is determined as the position information of the first target object.
[0011] The target tracking method further includes: if the similarity between the second feature map and the preset feature map does not exceed a second threshold, deleting the preset feature map and outputting preset position information.
[0012] In order to solve the above technical problems, the second technical solution adopted by the present invention is: to provide a terminal, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor is used to execute program data to implement the steps in the above target tracking method.
[0013] In order to solve the above technical problems, the third technical solution adopted by the present invention is: providing a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above target tracking method are implemented.
[0014] The beneficial effects of the present invention are as follows: different from the prior art, a target tracking method, terminal and computer-readable storage medium are provided, wherein the target tracking method performs target detection on a first video frame acquired by a first image acquisition device; in response to the first video frame not detecting a preset first target object, target detection is performed on a second video frame acquired by a second image acquisition device; the second video frame is acquired by a second image acquisition device; wherein the first monitoring area of the first image acquisition device is a subset of the second monitoring area of the second image acquisition device; in response to the second target object being detected in the second video frame and the second target object being the same as the first target object, the position information of the second target object in the second monitoring area is determined as the position information of the first target object. The present application performs target detection on the first video frame acquired by the first image acquisition device; when the preset first target object is not detected in the first video frame, target detection is performed on the second video frame acquired by the second image acquisition device; when the comparison shows that the second target object is the same as the preset first target object, the position information of the second target object in the second monitoring area is determined as the position information of the first target object, thereby avoiding the disappearance of the target object in the first image acquisition device and causing the target object to be lost, thereby improving the stability of target object tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] Figure 1 It is a flowchart of the target tracking method provided by the present invention;
[0017] Figure 2 It is a flowchart of a specific embodiment of the target tracking method provided by the present invention;
[0018] Figure 3 is a schematic block diagram of an embodiment of a terminal provided by the present invention;
[0019] Figure 4 It is a schematic block diagram of an embodiment of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION
[0020] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.
[0021] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0022] The term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship. Furthermore, "many" in this document means two or more than two.
[0023] In order to enable those skilled in the art to better understand the technical solution of the present invention, a target tracking method provided by the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0024] In this application, the image acquisition device is a binocular camera, comprising a first image acquisition device and a second image acquisition device. The first and second image acquisition devices simultaneously track a target object at the same viewing angle and different zoom settings. For example, the first and second image acquisition devices are installed in the same device. The first and second image acquisition devices each correspond to a visual sensor and are connected to a corresponding main chip. The target object can be a pedestrian, a vehicle, or the like. The first image acquisition device serves as a primary lens, which is a zoom lens. It calibrates and tracks targets within a small field of view, capturing more target details. In other words, the first image acquisition device captures images of a first monitoring area. The second image acquisition device serves as an auxiliary lens, which can be a fixed-focus lens or a zoom lens. It tracks targets within a larger field of view. In other words, the second image acquisition device captures images of a second monitoring area. The first monitoring area of the first image acquisition device is a subset of the second monitoring area of the second image acquisition device. In other words, the first monitoring area is a subset of the second monitoring area, facilitating the capture of more target details.
[0025] Among them, the centers of the lenses of the first image acquisition device and the second image acquisition device correspond to the same position in the PTZ (Pan / Tilt / Zoom) coordinate system. That is, the second image acquisition device and the first image acquisition device simultaneously track the same target object, and the coordinates of the target object in the image captured by the first image acquisition device are the same as the coordinates of the target object in the image captured by the second image acquisition device. In other words, the coordinates of the target object in the image captured by the first image acquisition device and the coordinates of the target object in the image captured by the second image acquisition device are both the coordinates of the target object in the second monitoring area. When the first image acquisition device does not capture an image containing the preset target object, the second image acquisition device is used to capture an image, and the position information of the preset target object in the second monitoring area can be determined by the position information of the preset target object in the image captured by the second image acquisition device.
[0026] See also Figure 1 , Figure 1 This is a flow chart of the target tracking method provided by the present invention. This embodiment provides a target tracking method suitable for tracking a target object in a dual-lens scenario. Without requiring coordinated monitoring equipment, the method uses a dual-lens approach to capture images on a local pan-tilt head (PTZ). The motion of the local PTZ is controlled based on its motion parameters to track the target object. The target tracking method includes the following steps.
[0027] S11: Performing target detection on a first video frame acquired, where the first video frame is acquired by a first image acquisition device.
[0028] Specifically, the first video frame is detected to obtain a candidate target object and location information of the candidate target object; features of the candidate target object are extracted to obtain a first feature map; and a determination is made as to whether the similarity between the first feature map and a preset feature map corresponding to the first target object stored in a video frame preceding the first video frame is greater than a first threshold. If the similarity between the first feature map and the preset feature map is greater than the first threshold, the candidate target object is determined to be the first target object; and the location information of the first target object in the first video frame is output. If the similarity between the first feature map and the preset feature map is not greater than the first threshold, it is determined that the preset first target object is not detected in the first video frame.
[0029] S12: In response to the first video frame not detecting the preset first target object, performing target detection on the acquired second video frame, where the second video frame is acquired by a second image acquisition device.
[0030] Specifically, if a preset first target object is not detected in the first video frame, multiple second video frames within a preset time period are acquired; target detection is performed on the multiple second video frames to obtain at least one second target object and position information of the second target object. If a preset first target object is not detected in the first video frame, at least one second video frame within a preset time period is acquired; target detection is performed on the second video frame; wherein the preset time period includes a time period including the time instant when the first video frame was captured and a first duration thereafter; or the preset time period includes a time period including a second duration thereafter when the first video frame was captured.
[0031] In this embodiment, features are extracted from the second target object to obtain a second feature map; and a determination is made as to whether the similarity between the second feature map and a preset feature map exceeds a second threshold. If the similarity between the second feature map and the preset feature map exceeds the second threshold, the second target object is determined to be the first target object. If the similarity between the second feature map and the preset feature map does not exceed the second threshold, the preset feature map is deleted and preset location information is output.
[0032] S13: In response to detecting a second target object in the second video frame, and the second target object is the same as the first target object, determining the position information of the second target object in the second monitoring area as the position information of the first target object.
[0033] Specifically, if the similarity between the second feature map and the preset feature map exceeds a second threshold, the second target object is determined to be the first target object, and the position information of the second target object is determined as the position information of the first target object.
[0034] This embodiment provides a target tracking method, which performs target detection on a first video frame acquired by a first image acquisition device; in response to a preset first target object not being detected in the first video frame, performs target detection on a second video frame acquired by a second image acquisition device; the second video frame is acquired by a second image acquisition device; wherein the first monitoring area of the first image acquisition device is a subset of the second monitoring area of the second image acquisition device; in response to a second target object being detected in the second video frame, and the second target object being identical to the first target object, determines the position information of the second target object in the second monitoring area as the position information of the first target object. This application performs target detection on the first video frame acquired by the first image acquisition device; when the preset first target object is not detected in the first video frame, performs target detection on the second video frame acquired by the second image acquisition device; when the comparison shows that the second target object is identical to the preset first target object, determines the position information of the second target object in the second monitoring area as the position information of the first target object, thereby avoiding the disappearance of the target object in the first image acquisition device, resulting in target object loss, thereby improving the stability of target object tracking.
[0035] See also Figure 2 , Figure 2 This is a flow chart illustrating a specific embodiment of a target tracking method provided by the present invention. This target tracking method is suitable for tracking a target object in a dual-lens scenario. Without requiring coordinated monitoring equipment, a local pan-tilt head (PTZ) uses a dual-lens approach to capture images and controls its motion based on its motion parameters to track the target object. This target tracking method includes the following steps.
[0036] S201: Acquire a first video frame.
[0037] Specifically, a current video frame is acquired by a first image acquisition device, and the current video frame serves as a first video frame. The first video frame captures an image within a small field of view. The first video frame may contain a first target object; or it may not contain the first target object, i.e., the first target object is temporarily lost or obscured by other objects. The first target object is a preset target object. In other words, the first target object is a target object that is set to be tracked. The first target object can be a person or a vehicle.
[0038] S202: Perform target detection on the first video frame to obtain candidate target objects and corresponding position information.
[0039] Specifically, an object detection network is used to perform object detection on the first video frame to obtain a detection frame of the candidate target object. In other words, the candidate target object and its location information in the first video frame are obtained. In a specific embodiment, an object detection network selected from Fast R-CNN, Faster R-CNN, YOLO, and SSD can be used to perform object detection on the first video frame to obtain a detection frame of the candidate target object.
[0040] S203: Extract features of the candidate target object to obtain a first feature map.
[0041] Specifically, a feature extraction network is used to extract features from the detected candidate target object to obtain a first feature map corresponding to the candidate target object. In a specific embodiment, the first feature map can be obtained by extracting features from the candidate target object contained in the detection frame using a feature extraction network selected from the group consisting of Histogram of Oriented Gradient (HOG), Scale Invariant Feature Transform (SIFT), Speeded Up Robust Features (SURF), and Superpoint.
[0042] S204: Determine whether the similarity between the first feature map and a preset feature map corresponding to a preset first target object in a previous video frame is greater than a first threshold.
[0043] Specifically, in order to confirm whether the detected candidate target object is the preset first target object. It is necessary to determine whether the first feature map corresponding to the candidate target object is the same as the preset feature map corresponding to the preset first target object in the previous video frame. That is, determine whether the similarity between the first feature map and the preset feature map is greater than the first threshold. In a specific embodiment, when the first video frame is not the first frame image containing the first target object, the similarity between the first feature map and the preset feature map of the first target object saved in the previous video frame is calculated. The similarity between the first feature map and the preset feature map of the first target object saved in any historical video frame before the first video frame can also be calculated. In another specific embodiment, when the first video frame is the first frame image containing the first target object, the similarity between the first feature map and the pre-stored preset feature map of the first target object is calculated.
[0044] In one embodiment, whether the first feature map is the same as the preset feature map may be determined by using a Fast Library for Approximate Nearest Neighbors (FLANN), a Brute Force (BF) algorithm, and a Euclidean distance.
[0045] If the similarity between the first feature map and the preset feature map is greater than the first threshold, the process directly jumps to step S205; if the similarity between the first feature map and the preset feature map is not greater than the first threshold, the process directly jumps to step S206.
[0046] S205: Determine the candidate target object as the first target object.
[0047] Specifically, if the similarity between the first feature map and the preset feature map is greater than a first threshold, it indicates that the first feature map is the same as the preset feature map, and it is determined that the candidate target object corresponding to the first feature map is the preset first target object.
[0048] S206: Output the position information of the first target object in the first video frame.
[0049] Specifically, when it is determined that the candidate target object is the preset first target object, it indicates that the first target object is detected in the first video frame, and the position information of the first target object in the first video frame is output as the position information of the first target object in the second monitoring area, so as to facilitate tracking of the first target object.
[0050] S207: Determine that no preset first target object is detected in the first video frame.
[0051] Specifically, if the similarity between the first feature map and the preset feature map is not greater than a first threshold, it indicates that the first feature map is different from the preset feature map, and it is determined that the candidate target object corresponding to the first feature map is not the preset first target object. If the comparison shows that all candidate target objects detected in the first video frame are not the preset first target object, it indicates that the first target object is lost in the first video frame or is obscured by other objects.
[0052] Since the first image acquisition device acquires an image containing the first target object in a small field of view, when the first target object is lost in the first video frame acquired by the first image acquisition device, the first image acquisition device cannot continue to track the first target object, and it is necessary to use the second image acquisition device to acquire images in a large field of view to determine the position information of the first target object.
[0053] S208: Acquire a second video frame.
[0054] Specifically, the second image acquisition device captures all video frames within a time period including the time of capture of the first video frame and a first duration thereafter as the second video frame. Alternatively, the second image acquisition device captures all video frames within a time period including a second duration thereafter as the second video frame. The second video frame can be a single frame or multiple frames. The second video frame captures an image with a wide field of view, making it easier to capture an image containing the target object.
[0055] S209: Perform target detection on the second video frame to obtain at least one second target object and position information of the second target object.
[0056] Specifically, the second video frame is subjected to target detection using an object detection network to obtain a detection frame of the second target object. In other words, the second target object and its location information in the second video frame are obtained. In a specific embodiment, the second video frame can be subjected to target detection using an object detection network selected from Fast R-CNN, Faster R-CNN, YOLO, and SSD to obtain a detection frame of the second target object.
[0057] S210: Extract features of the second target object to obtain a second feature map.
[0058] Specifically, a feature extraction network is used to extract features of the detected second target object to obtain a second feature map corresponding to the second target object. In a specific embodiment, a feature extraction network selected from HOG, SIFT, SURF, and Superpoint can be used to extract features of the second target object contained in the detection frame to obtain the second feature map.
[0059] S211: Determine whether the similarity between the second feature map and the preset feature map exceeds a second threshold.
[0060] Specifically, to confirm whether the detected second target object is the preset first target object, it is necessary to determine whether the second feature map corresponding to the second target object is identical to the preset feature map corresponding to the preset first target object. In other words, it is determined whether the similarity between the second feature map and the preset feature map is greater than a second threshold. The second threshold may be the same as or different from the first threshold.
[0061] If the similarity between the second feature map and the preset feature map is greater than the second threshold, jump directly to step S212; if the similarity between the second feature map and the preset feature map is not greater than the second threshold, jump directly to step S213.
[0062] S212: Determine the location information of the second target object as the location information of the first target object.
[0063] Specifically, if the similarity between the second feature map and the preset feature map exceeds a second threshold, the position information of the second target object in the second monitoring area is determined as the position information of the first target object. The position information of the second target object in the second monitoring area and the second feature map are sent to the first image acquisition device, and the first image acquisition device updates the second feature map to the preset feature map in the first video frame and saves it. The first image acquisition device outputs the position information of the second target object in the second monitoring area and continues to track the first target object based on the received position information of the second target object in the second monitoring area.
[0064] S213: Delete the preset feature map and output the preset position information.
[0065] Specifically, if the similarity between the second feature map and the preset feature map does not exceed a second threshold, indicating that the first target object has completely disappeared, there is no need to continue tracking the first target object, the preset feature map is deleted, and preset position information indicating the end of tracking is output. For example, the preset position information is the origin of the coordinate system.
[0066] In an optional embodiment, in order to obtain the preset feature map stored in the video frame closest to the first target object and the current video frame, it is determined whether the previous video frame stores the preset feature map corresponding to the first target object; if the previous video frame stores the preset feature map corresponding to the first target object, it is determined that the first video frame is not the first image frame containing the first target object. If the previous video frame does not store the preset feature map, it is determined that the first video frame is the first image frame containing the first target object, and the position information of the first target object in the first video frame is determined to be output.
[0067] This embodiment provides a target tracking method, which performs target detection on a first video frame acquired by a first image acquisition device; in response to a preset first target object not being detected in the first video frame, performs target detection on a second video frame acquired by a second image acquisition device; the second video frame is acquired by a second image acquisition device; wherein the first monitoring area of the first image acquisition device is a subset of the second monitoring area of the second image acquisition device; in response to a second target object being detected in the second video frame, and the second target object being identical to the first target object, determines the position information of the second target object in the second monitoring area as the position information of the first target object. This application performs target detection on the first video frame acquired by the first image acquisition device; when the preset first target object is not detected in the first video frame, performs target detection on the second video frame acquired by the second image acquisition device; when the comparison shows that the second target object is identical to the preset first target object, determines the position information of the second target object in the second monitoring area as the position information of the first target object, thereby avoiding the disappearance of the target object in the first image acquisition device, resulting in target object loss, thereby improving the stability of target object tracking.
[0068] See Figure 3 , Figure 3 This is a schematic block diagram of an embodiment of a terminal provided by the present invention. Terminal 70 in this embodiment includes: a processor 71, a memory 72, and a computer program stored in memory 72 and executable by processor 71. When executed by processor 71, this computer program implements the target tracking method described above. To avoid repetition, detailed descriptions are omitted here.
[0069] See Figure 4 , Figure 4 It is a schematic block diagram of an embodiment of a computer-readable storage medium provided by the present invention.
[0070] In an embodiment of the present application, a computer-readable storage medium 90 is further provided. The computer-readable storage medium 90 stores a computer program 901. The computer program 901 includes program instructions. The processor executes the program instructions to implement the target tracking method provided in an embodiment of the present application.
[0071] The computer-readable storage medium 90 may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium 90 may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc.
[0072] The above are merely embodiments of the present invention and are not intended to limit the scope of patent protection of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of patent protection of the present invention.
Claims
1. A target tracking method, characterized in that: The target tracking method is suitable for capturing images with a binocular camera, the binocular camera being located on a local pan-tilt platform, and controlling the movement of the local pan-tilt platform according to the motion parameters of the local pan-tilt platform; the binocular camera comprising a first image acquisition device and a second image acquisition device, the centers of the lenses of the first image acquisition device and the second image acquisition device corresponding to the same position in a PTZ coordinate system, and the target tracking method comprising: Performing target detection on a first video frame acquired, where the first video frame is acquired by a first image acquisition device; In response to the first video frame not detecting the preset first target object, performing target detection on a second video frame acquired, the second video frame being acquired by a second image acquisition device; wherein the first monitoring area of the first image acquisition device is a subset of the second monitoring area of the second image acquisition device; In response to detecting a second target object in the second video frame, and the second target object is the same as the first target object, determining the position information of the second target object in the second monitoring area as the position information of the first target object; The performing target detection on the acquired first video frame includes: Detecting the first video frame to obtain candidate target objects and location information of the candidate target objects; Extracting features of the candidate target object to obtain a first feature map; Determining whether a similarity between the first feature map and a preset feature map corresponding to the first target object stored in a video frame previous to the first video frame is greater than a first threshold; If the similarity between the first feature map and the preset feature map is not greater than the first threshold, it is determined that the preset first target object is not detected in the first video frame.
2. The target tracking method according to claim 1, characterized in that The target tracking method comprises: In response to detecting the preset first target object in the first video frame, position information of the first target object in the first video frame is output.
3. The target tracking method according to claim 2, characterized in that In response to detecting the preset first target object in the first video frame, outputting position information of the first target object in the first video frame includes: Determining whether a previous video frame of the first video frame stores a preset feature map of the first target object; If the previous video frame stores the preset feature map, it is determined that the first video frame is not the first frame image containing the first target object.
4. The target tracking method according to claim 3, characterized in that: In response to detecting the preset first target object in the first video frame, outputting position information of the first target object in the first video frame includes: If the similarity between the first feature map and the preset feature map is greater than the first threshold, determining that the candidate target object is the first target object; Output the position information of the first target object in the first video frame.
5. The target tracking method according to claim 3, characterized in that: The method further includes outputting position information of the first target object in the first video frame in response to detecting the first target object in the first video frame, and further includes: If the previous video frame does not save the preset feature map, it is determined that the first video frame is the first frame image containing the first target object, and the position information of the first target object in the first video frame is output.
6. The target tracking method according to claim 3, characterized in that: In response to the first video frame not detecting the preset first target object, performing target detection on the acquired second video frame includes: If the preset first target object is not detected in the first video frame, obtaining at least one second video frame within a preset time period; performing target detection on the at least one second video frame; The preset time period includes a time period including the first video frame acquisition moment and a first duration thereafter; or the preset time period includes a time period including a second duration thereafter and a first video frame acquisition moment.
7. The target tracking method according to claim 6, characterized in that: In response to detecting a second target object in the second video frame, and the second target object is the same as the first target object, determining the position information of the second target object in the second monitoring area as the position information of the first target object, the method comprising: performing feature extraction on the second target object to obtain a second feature map; Determining whether the similarity between the second feature map and the preset feature map exceeds a second threshold; In response to detecting a second target object in the second video frame, and the second target object is the same as the first target object, determining the position information of the second target object in the second monitoring area as the position information of the first target object includes: If the similarity between the second feature map and the preset feature map exceeds the second threshold, the location information of the second target object in the second monitoring area is determined as the location information of the first target object.
8. The target tracking method according to claim 7, characterized in that: The target tracking method further includes: If the similarity between the second feature map and the preset feature map does not exceed the second threshold, the preset feature map is deleted and the preset position information is output.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor is configured to execute program data to implement the steps of the target tracking method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the target tracking method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Multi-target active recognition tracking and monitoring method
CN105407283A
Object tracking method and device, storage medium and electronic device
CN110866480A