Target Automatic Detection and Tracking Method, Device, Computing Device and Storage Medium
By performing object detection and tracking on a single gimbal, the problems of high cost and difficult debugging in the existing methods are solved, and automatic detection and tracking of the target is achieved.
Patent Information
- Application Number
- CN202111490364.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing automatic target detection and tracking methods require search equipment and tracking equipment for target detection, which is costly and difficult to deploy and debug.
Use a single gimbal to realize object detection and tracking, and obtain video images in real time, perform object detection and generate tracking template images, and control the gimbal to track the target.
Reduces cost and deployment and debugging difficulty, and achieves automatic detection and tracking of goals.
Smart Images

Figure CN114140500B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer technology, and particularly to a method, device, computing device, and storage medium for automatic target detection and tracking. Background Art
[0002] At present, video surveillance is one of the important control means for border security and public security prevention and control. Vision algorithms such as object detection and object tracking based on deep learning are applied to video surveillance, enabling video surveillance to have a certain intelligent monitoring ability. Existing fully automatic detection and tracking methods mostly use optoelectronic search and tracking systems, which not only require search devices for target detection but also tracking devices for tracking targets. This not only has a high cost but also high difficulties in device deployment and debugging. Therefore, there is an urgent need for a new method for automatic target detection and tracking. Summary of the Invention
[0003] Based on the problems of existing fully automatic detection and tracking methods, which not only require search devices for target detection but also tracking devices for tracking targets, and have high costs as well as high difficulties in device deployment and debugging, the embodiments of the present invention provide a method, device, computing device, and storage medium for automatic target detection and tracking, which can utilize a single pan-tilt head to achieve automatic detection and tracking of targets, and at the same time can reduce costs and deployment and debugging difficulties.
[0004] In a first aspect, the embodiments of the present invention provide a method for automatic target detection and tracking, including:
[0005] Real-time obtaining video images collected by a pan-tilt head;
[0006] Performing target detection on each frame of the video image to obtain a video image containing the detected target;
[0007] Determining the detected target as a tracking target, and generating a tracking template image by using the video image containing the detected target; the tracking template image includes feature information of the tracking target;
[0008] Controlling the pan-tilt head to track the tracking target according to the tracking template image and the video image obtained in real time, so that the tracking target is located within the video image.
[0009] Preferably, the performing target detection on each frame of the video image includes:
[0010] Inputting each frame of the video image into the backbone network of the YOLOV4 network to obtain a first feature map and a second feature map;
[0011] In the neck network, performing upsampling processing on the first feature map to obtain a third feature map with the same size as the second feature map;
[0012] Perform tensor concatenation on the third feature map and the second feature map, perform convolution processing on the concatenated feature map obtained after tensor concatenation, input the feature map after convolution processing into the head network of the YOLOV4 network, and determine the detection target in the video image according to the output content of the head network.
[0013] Preferably, after determining the detection target as the tracking target and generating the tracking template image, it further includes:
[0014] When it is determined that the set calibration period is reached, determine whether the tracking target has changed according to the video image including the detection target in the current frame and the tracking template image; if it has changed, update the tracking template image.
[0015] Preferably, determining whether the tracking target has changed according to the video image including the detection target in the current frame and the tracking template image includes:
[0016] For the video image including the detection target in the current frame, determine the detection targets included in the current frame video image, and calculate the first similarity between each detection target and the tracking target included in the tracking template image;
[0017] Judge whether the maximum similarity among the obtained first similarities is less than the first threshold;
[0018] If not, obtain the current video image, and calculate the second similarity between the detection target corresponding to the maximum similarity and the tracking target in the current video image;
[0019] Judge whether the second similarity is less than the second threshold. If the second similarity is less than the second threshold, determine that the tracking target has changed;
[0020] Updating the tracking template image includes: updating the tracking template image according to the detection target corresponding to the maximum similarity.
[0021] Preferably, it further includes: determining the moving speed of the detection target; setting the calibration period according to the moving speed; the calibration period is negatively correlated with the moving speed.
[0022] Preferably, controlling the pan-tilt to track the tracking target according to the tracking template image and the video image obtained in real time, so that the tracking target is located within the video image, includes:
[0023] Calculate the distance between the target position of the tracking target in the video image and the specified position;
[0024] Calculate the current off-target amount of the pan-tilt according to the field of view angle of the pan-tilt and the distance;
[0025] Adjust the field of view angle of the pan-tilt according to the off-target amount so that the tracking target is located at the specified position in the video image.
[0026] Preferably, after obtaining the video image containing the detection target, it further includes: when it is determined that the number of consecutive frames of the video image not containing the detection target is greater than the set number of frames, it is determined that the detection target is lost.
[0027] In a second aspect, an embodiment of the present invention further provides an automatic target detection and tracking device, including:
[0028] A video image acquisition unit for real-time acquiring the video image acquired by the pan-tilt.
[0029] A target detection unit for performing target detection on each frame of the video image to obtain a video image containing the detection target.
[0030] A tracking target determination unit for determining the detection target as the tracking target and generating a tracking template image using the video image containing the detection target; the tracking template image includes the feature information of the tracking target.
[0031] A pan-tilt tracking control unit for controlling the pan-tilt to track the tracking target according to the tracking template image and the real-time acquired video image so that the tracking target is located within the video image.
[0032] In a third aspect, an embodiment of the present invention further provides a computing device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.
[0033] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in any embodiment of this specification.
[0034] An embodiment of the present invention provides a method, apparatus, computing device, and storage medium for automatic target detection and tracking. By obtaining in real time the video images collected by a camera in a pan-tilt head, detecting the detection targets in each frame of the video image, determining a tracking template image including the tracking target and its feature information, and then controlling the pan-tilt head to track the target based on the determined tracking template image and the video images obtained in real time. In this solution, the video images collected by the pan-tilt head are used for target detection on the one hand and target tracking on the other hand. Based on target tracking, the shooting direction of the pan-tilt head is controlled, so that a single pan-tilt head can be used to achieve target detection and target tracking at the same time. Therefore, this solution can use a single pan-tilt head to achieve automatic detection and tracking of targets, and at the same time can reduce costs and the difficulty of deployment and debugging. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 is a flowchart of a method for automatic target detection and tracking provided by an embodiment of the present invention;
[0037] Figure 2 is a flowchart of a judgment method for determining whether the tracking target has changed provided by an embodiment of the present invention;
[0038] Figure 3 is a flowchart of a method for controlling a pan-tilt head to track a target provided by an embodiment of the present invention;
[0039] Figure 4 is a hardware architecture diagram of a computing device provided by an embodiment of the present invention;
[0040] Figure 5 is a structural diagram of an apparatus for automatic target detection and tracking provided by an embodiment of the present invention;
[0041] Figure 6 is another structural diagram of an apparatus for automatic target detection and tracking provided by an embodiment of the present invention;
[0042] Figure 7 is still another structural diagram of an apparatus for automatic target detection and tracking provided by an embodiment of the present invention;
[0043] Figure 8 is yet another structural diagram of an apparatus for automatic target detection and tracking provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0045] As described above, existing automatic target detection and tracking methods still require manual auxiliary judgment and cannot stably track for a long time. Currently, most fully automatic detection and tracking methods use optoelectronic search and tracking systems, which not only require search devices for target detection but also tracking devices for target tracking. Such optoelectronic search and tracking systems not only have high costs but also high equipment deployment and debugging difficulties. If the search device and the tracking device can be combined into one, that is, a single pan-tilt unit can both detect and track the target, the cost of target detection and tracking can be greatly reduced. Based on this, the video images collected by the pan-tilt unit can be used for target detection on the one hand and target tracking on the other hand. The shooting direction of the pan-tilt unit can be controlled based on target tracking, so that a single pan-tilt unit can simultaneously achieve target detection and target tracking.
[0046] The following describes the specific implementation of the above concept.
[0047] Please refer to Figure 1 , an embodiment of the present invention provides an automatic target detection and tracking method, which includes:
[0048] Step 100, continuously obtain the video images collected by the pan-tilt unit in real time.
[0049] Step 102, perform target detection on each frame of the video image to obtain a video image containing the detected target.
[0050] Step 104, determine the detected target as the tracking target, and generate a tracking template image using the video image containing the detected target; the tracking template image includes the feature information of the tracking target.
[0051] Step 106, control the pan-tilt unit to track the tracking target according to the tracking template image and the video image continuously obtained in real time, so that the tracking target is located within the video image.
[0052] In the embodiments of the present invention, by obtaining in real time the video images collected by the camera in the pan-tilt, detecting the detection targets in each frame of the video images, determining the tracking template images containing the tracking targets and their feature information, and then controlling the pan-tilt to track the targets according to the determined tracking template images and the video images obtained in real time. This solution uses the video images collected by the pan-tilt for target detection on the one hand and for target tracking on the other hand, and controls the shooting direction of the pan-tilt based on target tracking, so that a single pan-tilt can be used to achieve target detection and target tracking simultaneously. Therefore, this solution can use a single pan-tilt to achieve automatic detection and tracking of targets, and at the same time can reduce costs and the difficulty of deployment and debugging.
[0053] The execution manners of the following Figure 1 shown steps are described.
[0054] First, for step 100, obtain in real time the video images collected by the pan-tilt.
[0055] In the embodiments of the present invention, by using a single pan-tilt, not only can a camera required for obtaining video images in real time be installed on the pan-tilt, but also the pan-tilt can be controlled by the target automatic detection and tracking network to automatically locate and track the target. Compared with ordinary monitoring devices, it can scan and monitor a larger range. It can track and monitor the target under the operation of the personnel in the monitoring center, or the camera on the pan-tilt can automatically scan the monitoring area under the control of the target automatic detection and tracking algorithm.
[0056] Then, for step 102, perform target detection on each frame of the video images to obtain the video images containing the detection targets.
[0057] In practical applications, it is necessary to pre-determine the target categories to be detected and tracked, and train the target detection and target tracking models through the target images in the monitoring scene to obtain the target detection and target tracking models suitable for actual monitoring. For example, if a vehicle is determined as the detection and tracking target, the images of the vehicle in the monitoring scene are required to train the target detection and target tracking models. Then, when performing target detection, the target detection can be performed on the video images obtained in real time through the feature information of the targets in the target detection and target tracking models to determine whether the detection targets are included.
[0058] Similarly, before detecting the target, the detection and tracking areas and rules can also be pre-determined according to the monitoring requirements. For example, if it is necessary to monitor the surrounding of a shopping mall and vehicles and animals are not allowed to enter area B, then the images of vehicles and animals need to be used to train the target automatic detection and tracking network in advance. And area B can be pre-set as the detection area and the tracking area. As long as a vehicle or an animal enters area B, detection, tracking and video evidence collection will start, and even an alarm will be triggered.
[0059] Therefore, this solution can pre-determine the target category or monitoring rules according to the actual situation. After determination, target detection and tracking can be performed on real-time video images based on these preset conditions.
[0060] In an embodiment of the invention, the YOLOV4 network can be used to perform target detection on each frame of video image. Specifically, the following steps S1-S3 can be included:
[0061] S1, Input each frame of video image into the backbone network of the YOLOV4 network to obtain a first feature map and a second feature map.
[0062] In an embodiment of the present invention, in order to make the target detection accuracy of the target automatic detection and tracking network higher and more adaptable to the detection of small targets, a scale is added to the detection network, and the target is subjected to feature extraction and detection through multiple scales. For example, the original YOLOV4 network inputs each frame of video image into the backbone network of the YOLOV4 network, and the backbone network performs feature extraction on each frame of video image and outputs feature maps with sizes of 16*16, 32*32, and 64*64. In an embodiment of the present invention, a feature map with a size of 128*128 is added as the second feature map, that is, the backbone network of the YOLOV4 network outputs feature maps with sizes of 16*16, 32*32, 64*64, and 128*128.
[0063] S2, In the neck network, perform upsampling processing on the first feature map to obtain a third feature map with the same size as the second feature map.
[0064] Specifically, in the neck network, the feature map with a size of 64*64 obtained from the backbone network is subjected to upsampling processing to obtain a third feature map with a size of 128*128. Correspondingly, the original YOLOV4 network also performs upsampling processing on the 16*16 and 32*32 feature maps respectively to obtain corresponding feature maps with sizes of 32*32 and 64*64.
[0065] S3, Perform tensor splicing on the third feature map and the second feature map, perform convolution processing on the spliced feature map obtained after tensor splicing, input the feature map after convolution processing into the head network of the YOLOV4 network, and determine the detection target in the video image according to the output content of the head network.
[0066] In an embodiment of the present invention, through the upsampling process in step S2, a third feature map with a size of 128*128 is newly added. In order to enable full feature fusion, the feature maps with sizes of 32*32, 64*64, and 128*128 after upsampling in the neck network are respectively tensor-concatenated with the feature maps of the same size in the backbone network, and five convolutions are respectively performed on the three sizes of concatenated feature maps obtained after tensor concatenation and the 16*16 feature map in the backbone network. The four sizes of feature maps after convolution processing are input into the head network of the YOLOV4 network, and the detection targets in the video image are determined according to the output content of the head network, and a video image containing the detection targets is obtained.
[0067] In an embodiment of the present invention, when it is determined that the detection target appears in the detection area, the detection and tracking of the target start by using the video image obtained in real time. Since there may be multiple targets entering the detection area simultaneously or continuously, when the determined detection target leaves the detection area or the detection target tracking is lost, in order to detect other targets as soon as possible, when the number of consecutive frames of the video image determined not to contain the detection target is greater than the set number of frames, it is determined that the detection target is lost, the detection and tracking are stopped, and the detection of other targets is restarted.
[0068] In an embodiment of the present invention, after obtaining the video image containing the detection target, whether the target is lost is determined not only by whether the detection target is included in the video image of consecutive frames, but also the target is tracked by using the real-time video image, and tracking correction is continuously performed.
[0069] Next, for step 104, the detection target is determined as the tracking target, and a tracking template image is generated by using the video image containing the detection target; the tracking template image includes the feature information of the tracking target.
[0070] It should be noted that after the YOLOV4 network performs target detection on each frame of the video image, there may be multiple detection targets at the same time. There are many ways to determine the tracking target from these multiple detection targets. For example, the target confidence of each detection target can be calculated, and the detection target with the highest target confidence is selected as the tracking target; or the distance between each detection target and a specified position can be calculated, and the detection target with the closest distance is selected as the tracking target; or the tracking target can be directly selected by the monitoring personnel from the detection targets.
[0071] In an embodiment of the present invention, a tracking target is determined by using any one of the above methods, and a tracking template image is generated from a video image including a detection target in step 102. The SiamRPN++ network based on the Siamese network is used to track the target through a real-time video image collected by a camera in a pan-tilt head and the tracking template image. Among them, the ResNet50 network is used as a feature extraction network to extract feature information of the tracking target in the tracking template image and the real-time video image respectively, and the position information of the tracking target in each frame of the real-time video image is judged according to the feature information of the tracking target in the tracking template image.
[0072] To improve the multi-scale tracking ability of the target, a feature fusion method is used in the ResNet50 network to enhance the fusion of the shallow surface information and the deep semantic information of the target. To reduce the model calculation pressure, in the ResNet50 network, not only the depth cross-correlation calculation method for reducing invalid network parameters is used, but also the feature map size is compressed and the number of feature map channels is reduced.
[0073] When a situation of loss of the tracking target caused by occlusion, mutation, etc. of the tracking target occurs, the "local-to-global" search method is used. With the position of the tracking target in the video image of the previous frame before the loss of the tracking target as the center, the search area is gradually increased with a set step length to alleviate the disordered search phenomenon after the target is occluded or lost.
[0074] To achieve long-term stable tracking of the target, in an embodiment of the present invention, when it is determined that a set calibration period is reached, it is determined whether the tracking target has changed according to the video image including the detection target in the current frame and the tracking template image; if it has changed, the tracking template image is updated. By correcting and updating the tracking target in the tracking template image, long-term stable tracking of the tracking target in the real-time video image is achieved.
[0075] Among them, the calibration period can be set according to the moving speed of the detection target in the video image in actual applications. Although the shorter the calibration period, the higher the calibration accuracy, it will cause a waste of a large amount of computing resources, thus affecting the tracking effect. Therefore, in order to select a suitable calibration period, it is necessary to determine the moving speed of the detection target; set the calibration period according to the moving speed; the calibration period is negatively correlated with the moving speed.
[0076] For example, since a vehicle moves faster than a person in a video image, when the detection target is a vehicle, the correction period needs to be shorter than when the detection target is a person in order to better adapt to the tracking and correction of the vehicle. However, if the detection targets are a vehicle and a person, in order not to affect the tracking effect of each detection target, the correction period needs to be set according to the moving speed of the vehicle in the video image. Therefore, if the detection targets are of multiple types and there is a significant difference in their moving speeds, the correction period needs to be determined according to the detection target with a faster moving speed.
[0077] After determining the correction period, tracking correction is performed intermittently according to the correction period. For each correction period, it is necessary to determine whether the tracking target has changed based on the video image containing the detection target in the current frame and the tracking template image.
[0078] The following describes the method for determining whether the tracking target has changed.
[0079] In the embodiments of the present invention, please refer to Figure 2 , at least the following steps 200-206 can be used to determine whether the tracking target has changed:
[0080] Step 200, for the video image containing the detection target in the current frame, determine the detection targets included in the current frame video image, and calculate the first similarity between each detection target and the tracking target included in the tracking template image.
[0081] In the embodiments of the present invention, in order to calculate the first similarity between each detection target and the tracking target included in the tracking template image, it is necessary to perform feature extraction on the video image containing the detection target in the current frame and the tracking template image containing the tracking target through an improved ResNet50 network to obtain the feature vectors of each detection target and the tracking target in the tracking template image.
[0082] Specifically, for the convenience of subsequent feature extraction, first, the video image containing the detection target in the current frame and the tracking template image containing the tracking target are both adjusted to a size of 255*255*3, and then are respectively input into the improved ResNet50 network for feature extraction.
[0083] In the improved ResNet50 network, there are a total of five convolution sets for feature extraction of the input image. In order to retain more feature information of the detection target and the tracking target, make the calculation of the first similarity more accurate, and reduce the effective stride in the fourth convolution set and the fifth convolution set to reduce the dimensionality reduction scale and retain more feature information; correspondingly, in order not to affect the feature extraction speed, a convolution layer with a convolution kernel of 1*1 is added to the output ends of the third convolution set, the fourth convolution set, and the fifth convolution set to reduce the number of feature map channels and improve the feature extraction speed.
[0084] In an embodiment of the present invention, the size of the feature map of the output layer of the improved ResNet50 network is 32*32*256. Subsequently, the feature map is flattened into a one-dimensional feature vector, that is, the feature vectors of each detection target and the tracking target in the tracking template image are obtained.
[0085] After obtaining the feature vectors of each detection target and the tracking target in the tracking template image, the first similarity between each detection target and the tracking target included in the tracking template image is calculated respectively according to the preset cosine similarity calculation rule.
[0086] Step 202: Determine whether the maximum similarity among the obtained first similarities is less than the first threshold.
[0087] In an embodiment of the present invention, by determining whether the maximum similarity among the obtained first similarities is less than the first threshold, it is determined whether the similarities between the detection targets in the current frame of video image and the tracking targets in the tracking template image are all lower than the first threshold. If they are all lower than the first threshold, it can be determined that the previously determined tracking target has been lost or has left, and the tracking can be stopped.
[0088] Step 204: If not, obtain the current video image, and calculate the second similarity between the detection target corresponding to the maximum similarity and the tracking target in the current video image.
[0089] Next, an explanation is given for the judgment result in step 202. If the maximum similarity among the first similarities is greater than the first threshold, it means that the previously determined tracking target is still in the monitoring area, and it is necessary to continue to detect and track this target.
[0090] Since when judging the position information of the tracking target in each frame of real-time video image according to the feature information of the tracking target in the tracking template image, the feature extraction network used is different from the feature extraction network for target detection of the detection target, it is not excluded that the feature information of the tracked target is different from that of the detection target. Therefore, it is necessary to obtain the current video image, use the feature vector calculation method in step 200, and use the improved ResNet50 network to extract the features of the tracking target in the current video image to obtain the feature vector of the tracking target in the current video image. According to the preset cosine similarity calculation rule, use the feature vector of the detection target corresponding to the maximum similarity in step 202 and the feature vector of the tracking target in the current video image to calculate the second similarity between the detection target corresponding to the maximum similarity and the tracking target in the current video image.
[0091] Step 206: Determine whether the second similarity is less than the second threshold. If the second similarity is less than the second threshold, it is determined that the tracking target has changed.
[0092] Since the tracking template image containing the target feature information is generated when the target first enters the detection area, the state of the target and the environment it is in are constantly changing. Also, the feature extraction network for tracking the target in the real-time video image is different from the feature extraction network for detecting the target during target detection. Therefore, it is very likely that the tracking target in the current video image is inconsistent with the tracking target determined when it first entered the detection area. Thus, by determining whether the second similarity between the detection target corresponding to the maximum similarity obtained in step 204 and the tracking target in the current video image is less than the second threshold, it is determined whether the tracking target has changed. If the second similarity is less than the second threshold, it can be determined that the tracking target has changed; if it is greater than the second threshold, it can be determined that the tracking target has not changed, and the position information of the tracking target in the real-time video image is still judged based on the feature information of the tracking target in the previous tracking template image, thereby controlling the pan-tilt to track the target.
[0093] Steps 200 - 206 complete the determination of whether the tracking target has changed. If it has changed, the tracking template image needs to be updated, specifically including: updating the tracking template image according to the detection target corresponding to the maximum similarity.
[0094] In the embodiment of the present invention, if it is determined that the tracking target has changed, it can be determined that the tracking target in the current video image may be inconsistent with the tracking target determined when it first entered the detection area, and there is a risk of losing the tracking target. Therefore, the detection target corresponding to the maximum similarity in the first similarity in step 202 can be re-determined as the tracking target, that is, the video image containing this detection target is used to replace the original tracking template image, so as to achieve large-range and highly reliable re-identification and tracking of the tracking target.
[0095] Finally, for step 106, according to the tracking template image and the real-time acquired video image, control the pan-tilt to track the tracking target so that the tracking target is located within the video image.
[0096] In the embodiment of the present invention, please refer to Figure 3 , at least the following steps 300 - 304 can be used to control the pan-tilt to track the tracking target:
[0097] Step 300, calculate the distance between the target position of the tracking target in the video image and the specified position.
[0098] In the embodiment of the present invention, in order to ensure the tracking and detection of the tracking target, the specified position can be located in the central area of the image. For example, the center of the tracking target is located at the central pixel point of the image, and according to the position information of the tracking target in the real-time video image, calculate the distance between the central pixel point of the tracking target in the video image and the central pixel point of the video image.
[0099] Step 302: Calculate the current miss distance of the pan-tilt head according to the field of view angle of the pan-tilt head and the distance.
[0100] Specifically, the following formula can be used to calculate the current miss distance of the pan-tilt head:
[0101]
[0102]
[0103] where Δα is the angular miss distance of the horizontal azimuth of the pan-tilt head; Δx is the horizontal distance between the center pixel of the tracking target and the specified position (such as the center pixel) of the video image; θ x is the field of view angle of the current horizontal azimuth of the pan-tilt head; w is the width of the video image; Δβ is the angular miss distance of the pitch angle of the pan-tilt head; Δy is the vertical distance between the center pixel of the tracking target and the specified position (such as the center pixel) of the video image; θ y is the field of view angle of the current pitch azimuth of the pan-tilt head; h is the height of the video image.
[0104] Step 304: Adjust the field of view angle of the pan-tilt head according to the miss distance so that the tracking target is located at the specified position of the video image.
[0105] In the embodiment of the present invention, the pan-tilt head is controlled to move along with the tracking target according to the miss distance obtained in step 302, so as to keep the tracking target always located in the central area of the real-time video image collected by the camera in the pan-tilt head, complete the long-term stable tracking of the detection target, and record evidence by video when necessary.
[0106] As Figure 4 、 Figure 5 shown, the embodiment of the present invention provides an automatic target detection and tracking device. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, as Figure 4 shown, it is a hardware architecture diagram of a computing device where the automatic target detection and tracking device provided by the embodiment of the present invention is located. In addition to Figure 4 the processor, memory, network interface, and non-volatile memory shown, the computing device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking the software implementation as an example, as Figure 5 shown, as a logically meaningful device, it is formed by the CPU of the computing device where it is located reading the corresponding computer program in the non-volatile memory into the memory and running. An automatic target detection and tracking device provided in this embodiment includes:
[0107] A video image acquisition unit 501, configured to acquire a video image collected by a pan-tilt head in real time;
[0108] The target detection unit 502 is configured to perform target detection on each frame of video image to obtain a video image including the detected target;
[0109] The tracking target determination unit 503 is configured to determine the detected target as a tracking target and generate a tracking template image by using the video image including the detected target; the tracking template image includes the feature information of the tracking target;
[0110] The pan-tilt tracking control unit 504 is configured to control the pan-tilt to track the tracking target according to the tracking template image and the video image obtained in real time, so that the tracking target is located within the video image.
[0111] In an embodiment of the present invention, when performing the target detection on each frame of video image, the target detection unit 502 is specifically configured to input each frame of video image into the backbone network of the YOLOV4 network to obtain a first feature map and a second feature map; in the neck network, perform upsampling processing on the first feature map to obtain a third feature map with the same size as the second feature map; perform tensor splicing on the third feature map and the second feature map, and perform convolution processing on the spliced feature map obtained after tensor splicing, and input the feature map after convolution processing into the head network of the YOLOV4 network, and determine the detected target in the video image according to the output content of the head network.
[0112] Please refer to Figure 6 , the target automatic detection and tracking device may further include:
[0113] The tracking target change determination unit 505 is configured to determine whether the tracking target has changed according to the video image including the detected target in the current frame and the tracking template image when it is determined that a set calibration period is reached; if it has changed, update the tracking template image.
[0114] In an embodiment of the present invention, when the tracking target change determination unit 505 determines whether the tracking target has changed according to the video image including the detection target in the current frame and the tracking template image, it is specifically configured to, for the video image including the detection target in the current frame, determine the detection target included in the current frame video image, and calculate the first similarity between each detection target and the tracking target included in the tracking template image; determine whether the maximum similarity among the obtained first similarities is less than a first threshold; if not, obtain the current video image, and calculate the second similarity between the detection target corresponding to the maximum similarity and the tracking target in the current video image; determine whether the second similarity is less than a second threshold, and if the second similarity is less than the second threshold, determine that the tracking target has changed; when the tracking target change determination unit 505 updates the tracking template image, it is further specifically configured to update the tracking template image according to the detection target corresponding to the maximum similarity.
[0115] Please refer to Figure 7 , the target automatic detection and tracking device may further include:
[0116] A correction period setting unit 506, configured to determine the moving speed of the detection target; set a correction period according to the moving speed; the correction period is negatively correlated with the moving speed.
[0117] In an embodiment of the present invention, when the pan-tilt tracking control unit 504 controls the pan-tilt to track the tracking target according to the tracking template image and the video image obtained in real time, so that the tracking target is located within the video image, it is specifically configured to calculate the distance between the target position of the tracking target in the video image and a specified position; calculate the current off-target amount of the pan-tilt according to the field of view angle of the pan-tilt and the distance; adjust the field of view angle of the pan-tilt according to the off-target amount, so that the tracking target is located at the specified position in the video image.
[0118] Please refer to Figure 8 , the target automatic detection and tracking device may further include:
[0119] A continuous frame number determination unit 507, configured to determine that the detection target is lost when the continuous frame number of the video image determined not to include the detection target is greater than a set frame number.
[0120] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on an object automatic detection and tracking device. In other embodiments of the present invention, an object automatic detection and tracking device may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0121] Regarding the information interaction, execution process, etc. between the various modules within the above-mentioned device, since they are based on the same concept as the method embodiments of the present invention, the specific content can be referred to the description in the method embodiments of the present invention, and will not be elaborated here.
[0122] The embodiments of the present invention also provide a computing device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements an object automatic detection and tracking method in any embodiment of the present invention.
[0123] The embodiments of the present invention also provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute an object automatic detection and tracking method in any embodiment of the present invention.
[0124] Specifically, a system or device equipped with a storage medium can be provided. Software program codes for implementing the functions of any one of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium.
[0125] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0126] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0127] Furthermore, it should be clear that not only can the actual operations be completed in part or in whole by executing the program code read by the computer, but also by the operating system, etc. operating on the computer based on the instructions of the program code, so as to implement the functions of any one of the above embodiments.
[0128] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion module connected to the computer. Subsequently, based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion module is made to execute part or all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0129] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0130] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes various media such as ROM, RAM, magnetic disk or optical disc that can store program code.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically detecting and tracking a target, characterized in that, Including: Obtaining in real time the video images collected by the pan-tilt head; Performing object detection on each frame of the video image to obtain a video image containing the detected object; Determining the detected object as a tracking object, and generating a tracking template image by using the video image containing the detected object; the tracking template image includes the feature information of the tracking object; Using the SiamRPN++ network based on the Siamese network to control the pan-tilt head to track the tracking object according to the tracking template image and the video image obtained by real-time collection by the pan-tilt head, so that the tracking object is located within the video image; In the SiamRPN++ network, the ResNet50 network is used as the feature extraction network. The feature extraction network respectively extracts the feature information of the tracking object in the tracking template image and the real-time video image, and judges the position information of the tracking object in each frame of the real-time video image according to the feature information of the tracking object in the tracking template image; in the ResNet50 network, feature fusion is used to enhance the fusion of the shallow surface information and the deep semantic information of the object. The ResNet50 network uses a depth cross-correlation calculation method to reduce invalid network parameters, and compresses the size of the feature map and reduces the channels of the feature map; When the tracking object is lost, a local-to-global search method is used. Taking the position of the tracking object in the previous frame of the video image before the tracking object is lost as the center, the search area is gradually increased with a set step size to avoid disordered search; The controlling the pan-tilt head to track the tracking object according to the tracking template image and the real-time acquired video image, so that the tracking object is located within the video image, includes: Calculating the distance between the target position of the tracking object in the video image and the specified position; Calculating the current off-target amount of the pan-tilt head according to the field of view angle of the pan-tilt head and the distance; Adjusting the field of view angle of the pan-tilt head according to the off-target amount, so that the tracking object is located at the specified position of the video image; Using the following formula to calculate the current off-target amount of the pan-tilt head: Among them, Δα is the angular miss distance of the pan-tilt's horizontal azimuth; Δx is the horizontal distance between the center pixel point of the tracking target and the specified position of the video image; θ x is the field of view angle of the pan-tilt's current horizontal azimuth; w is the width of the video image; Δβ is the angular miss distance of the pan-tilt's pitch angle; Δy is the vertical distance between the center pixel point of the tracking target and the specified position of the video image; θ y is the field of view angle of the pan-tilt's current pitch azimuth; h is the height of the video image.
2. The method according to claim 1, characterized in that, The performing object detection on each frame of the video image includes: Inputting each frame of the video image into the backbone network of the YOLOV4 network to obtain a first feature map and a second feature map; In the neck network, performing upsampling processing on the first feature map to obtain a third feature map with the same size as the second feature map; Performing tensor splicing on the third feature map and the second feature map, performing convolution processing on the spliced feature map obtained after tensor splicing, and inputting the feature map after convolution processing into the head network of the YOLOV4 network, and determining the detected object in the video image according to the output content of the head network.
3. The method according to claim 1, wherein After determining the detected object as a tracking object and generating a tracking template image, it further includes: When it is determined that the set calibration period is reached, determining whether the tracking object has changed according to the video image containing the detected object in the current frame and the tracking template image; if it has changed, updating the tracking template image.
4. The method according to claim 3, wherein Determining whether the tracking target has changed according to the video image including the detection target in the current frame and the tracking template image includes: For the video image including the detection target in the current frame, determining the detection targets included in the video image of the current frame, and calculating the first similarity between each detection target and the tracking target included in the tracking template image; Judging whether the maximum similarity among the obtained first similarities is less than a first threshold; If not, obtaining the current video image, and calculating the second similarity between the detection target corresponding to the maximum similarity and the tracking target in the current video image; Judging whether the second similarity is less than a second threshold, and if the second similarity is less than the second threshold, determining that the tracking target has changed; Updating the tracking template image includes: updating the tracking template image according to the detection target corresponding to the maximum similarity.
5. The method according to claim 3, wherein, Further includes: Determining the moving speed of the detection target; setting a correction period according to the moving speed; the correction period is negatively correlated with the moving speed.
6. The method according to any one of claims 1-5, characterized in that, After obtaining the video image including the detection target, it further includes: when it is determined that the number of consecutive frames of the video image not including the detection target is greater than a set number of frames, determining that the detection target is lost.
7. An automatic target detection and tracking device for implementing the method according to any one of claims 1-6, characterized in that, Includes: A video image acquisition unit for real-time acquisition of video images collected by a pan-tilt; A target detection unit for performing target detection on each frame of video image to obtain a video image including a detection target; A tracking target determination unit for determining the detection target as a tracking target and generating a tracking template image using the video image including the detection target; The tracking template image includes the feature information of the tracking target; A pan-tilt tracking control unit for controlling the pan-tilt to track the tracking target according to the tracking template image and the real-time acquired video image, so that the tracking target is located within the video image.
8. A computing device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1-6 is implemented.
9. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Target real-time tracking method and device, computer device and storage medium
CN108985162A
Target tracking method and device, electronic equipment and storage medium
CN112884809A
Food material identification system based on Internet-of-things refrigerator
CN113269208A