Target tracking method, close-up camera, system and storage medium
By linking panoramic and close-up cameras, the panoramic camera performs coarse detection of the overall region, while the close-up camera performs fine detection of the local region. The bounding boxes are then fused to generate target tracking boxes, which solves the problem of insufficient tracking accuracy caused by the large amount of computation in the existing technology and achieves high-precision and stable target tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, panoramic cameras and close-up cameras detect the human body by scanning the entire frame, which results in a large amount of computation and insufficient frame rate, leading to insufficient tracking accuracy.
By using a panoramic camera to detect the overall bounding box of the target object at a lower frame rate and a close-up camera to detect the bounding box of the local area at a higher frame rate, and then using a Kalman filter to fuse the bounding boxes of the two to generate the target tracking box, high-precision tracking of the target object can be achieved.
It improves the accuracy and stability of target tracking, reduces the consumption of computing resources, avoids screen stuttering, and increases the tracking frame rate.
Smart Images

Figure CN121788560A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of tracking technology, and in particular to a target tracking method, a close-up camera, a system, and a storage medium. Background Technology
[0002] In educational recording scenarios (such as recording open classes and high-quality courses), it is necessary to automatically track and film teachers, outputting high-definition video footage including close-ups of teachers to assist in the production and dissemination of high-quality teaching content.
[0003] Currently, existing solutions typically use panoramic cameras to detect the tracking target in a panoramic image and close-up cameras to detect the human body for tracking. However, since both panoramic and close-up cameras detect the human body by using the entire frame—for example, by running a multi-task human body model to simultaneously detect the human body, pose, and face—the computational load is high. Furthermore, this high computational load leads to insufficient tracking frame rate, resulting in insufficient tracking accuracy. Summary of the Invention
[0004] To address the aforementioned technical problems, embodiments of this application provide a target tracking method, a close-up camera, a system, and a storage medium, which can improve target tracking accuracy.
[0005] To address the aforementioned technical problems, the embodiments of this application provide the following technical solutions: In a first aspect, embodiments of this application provide a target tracking method applied to a close-up camera, the close-up camera being connected to a panoramic camera, the method comprising: A first image of the overall region of a target object sent by a panoramic camera is acquired, wherein the first image includes a first bounding box, which is obtained by the panoramic camera detecting the video frame at a first frame rate, and the first bounding box is used to determine the overall region range of the target object. A second image is acquired, and based on the second frame rate and the first bounding box, the second image is detected to generate a second bounding box of the local region of the target object, wherein the second frame rate is greater than the first frame rate; By fusing the first bounding box and the second bounding box, the target tracking box is obtained; The target object is tracked based on the target tracking bounding box.
[0006] In some embodiments, before acquiring a first image of the target object sent by the panoramic camera, the method further includes: Obtain the coordinate mapping relationship between the coordinate system of the panoramic camera and the coordinate system of the close-up camera; Based on the second frame rate and the first bounding box, the second image is detected to generate a second bounding box for the local region of the target object, including: Based on the coordinate mapping relationship, the first bounding box is transformed into a third bounding box in the coordinate system of the close-up camera; Based on the second frame rate and the third bounding box, the second image is detected to generate a second bounding box of the local region of the target object.
[0007] In some embodiments, the first bounding box and the second bounding box are fused to obtain the target tracking box, including: Based on the Kalman filter, the third bounding box and the second bounding box are fused to obtain the target tracking box.
[0008] In some embodiments, the panoramic camera runs a multi-object detection model at a first frame rate to identify target objects in the video frame and generate a first bounding box corresponding to the overall area of the target object. The close-up camera runs the head and shoulder detection model at a second frame rate to detect local regions within the overall region and generate second bounding boxes corresponding to the local regions, where the local regions include the head and shoulder regions.
[0009] In some embodiments, the method further includes: If the close-up camera does not detect the second bounding box of the local region of the target object, target tracking is performed based on the first recovery strategy, wherein the first recovery strategy includes: The system acquires the first image sent by the panoramic camera in real time to determine the current first bounding box of the overall region of the target object. Based on the head and shoulder detection model and the current first bounding box, it detects the current second image and generates the second bounding box corresponding to the head and shoulder region.
[0010] In some embodiments, the method further includes: If the panoramic camera fails to detect a first bounding box of the overall region of the target object, and the close-up camera fails to detect a second bounding box of a local region of the target object, then target tracking is performed based on a second recovery strategy, wherein the second recovery strategy includes: The panoramic camera is controlled to stop sending the first image to the close-up camera until the panoramic camera re-detects the first bounding box of the overall region of the target object. Then, the panoramic camera is controlled to re-send the first image of the overall region of the target object to the close-up camera. Based on the head and shoulder detection model and the first bounding box in the first image, the second image is detected to generate the second bounding box corresponding to the head and shoulder region.
[0011] In some embodiments, the close-up camera is deployed on a gimbal, which is used to fix and support the close-up camera; Based on the target tracking bounding box, the target object is tracked, including: Based on the target tracking frame, the gimbal is controlled in real time to perform horizontal rotation, and / or vertical rotation, and / or zoom control to track the target object.
[0012] Secondly, embodiments of this application provide a close-up camera, including: At least one processor; At least one memory for storing at least one program; When at least one program is executed by at least one processor, such that at least one processor implements the method of the first aspect.
[0013] Thirdly, embodiments of this application provide a target tracking system, including: A panoramic camera is used to acquire a first image, detect target objects in the first image, and determine a first bounding box of the overall region of the target object; For example, a close-up camera in the second aspect is used to acquire a first image of a target object sent by a panoramic camera, and to acquire a second image based on a second frame rate, and to detect the second image based on a first bounding box in the first image to obtain a second bounding box of a local region of the target object. A gimbal is used to fix and support a close-up camera for tracking a target object.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a processor-executable program that, when executed by a processor, is used to perform the method as described in the first aspect.
[0015] The beneficial effects of the embodiments of this application are as follows: Unlike the prior art, the embodiments of this application provide a target tracking method. This method uses a panoramic camera to detect the first bounding box of the overall region of the target object at a lower frame rate, and then uses a close-up camera to detect the second bounding box of the local region of the target object at a higher frame rate. The first bounding box and the second bounding box are then fused to obtain a target tracking box, which is used to track the target object. By linking the panoramic camera and the close-up camera, the panoramic camera is used to perform coarse detection of the overall region, and the close-up camera is used to perform fine detection of the local region, thereby improving the tracking accuracy. Attached Figure Description
[0016] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0017] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a target tracking method provided in an embodiment of this application; Figure 3 yes Figure 2 A detailed flowchart of step S202 in the process; Figure 4 This is an example schematic diagram of determining a target tracking box provided in an embodiment of this application; Figure 5 This is an example schematic diagram of another method for determining a target tracking box provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the process of recovering target object tracking using a close-up camera, as provided in an embodiment of this application. Figure 7 This is a schematic diagram illustrating the process of restoring target object tracking using a panoramic camera and a close-up camera, as provided in an embodiment of this application. Figure 8 This is a schematic diagram of the structure of a target tracking system provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a target tracking device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a close-up camera provided in an embodiment of this application.
[0018] Explanation of icon numbers: Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. In addition, the terms "first" and "second" used in this application do not limit the data, but only distinguish the same or similar items with basically the same function and effect.
[0021] Before introducing the embodiments of this application, a brief introduction will be given to the target tracking methods known to the inventors of this application, so as to facilitate the understanding of the embodiments of this application later.
[0022] Currently, images are typically acquired using a single close-up camera and then input into a human multi-task model for detection to determine the person's position. The gimbal is then used to track the person. Alternatively, a panoramic camera can be used for person recognition, while a close-up camera can detect the entire area of the person for tracking.
[0023] The drawbacks of the above method include: (1) The computational complexity of the human multi-task model is high, the tracking frame rate of the single close-up camera is low, resulting in stuttering in the final displayed image. In addition, the shooting area of the single close-up camera is small, and it takes a long time to re-track when the target is lost.
[0024] (2) The close-up camera still detects the entire image of the person and still needs to calculate a lot of data, resulting in low tracking accuracy.
[0025] To address the aforementioned issues, this application provides a target tracking method. This method uses a panoramic camera at a lower frame rate to detect a first bounding box of the overall region of the target object, and then uses a close-up camera at a higher frame rate to detect a second bounding box of a local region of the target object. The first and second bounding boxes are then fused to obtain a target tracking box, which is used to track the target object. By linking the panoramic and close-up cameras, the panoramic camera performs coarse detection of the overall region, while the close-up camera performs fine detection of the local region, thereby improving tracking accuracy.
[0026] The technical solution of this application is described in detail below with reference to the accompanying drawings: Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application.
[0027] like Figure 1 As shown, the application environment 100 includes a panoramic camera 10 and a close-up camera 20. The panoramic camera 10 is connected to the close-up camera 20 via a wired interface or a wireless network. The wired interface includes, but is not limited to, a USB interface and an HDMI interface, and the wireless network includes, but is not limited to, Wi-Fi and Bluetooth.
[0028] In this embodiment, the panoramic camera 10 is used to detect the overall area of the target object, and the close-up camera 20 is used to detect the local area of the target object. For example, in an educational recording scenario, the panoramic camera 10 is used to detect the overall area of the teacher, and the close-up camera 20 is used to detect the head and shoulder area of the teacher.
[0029] In this embodiment, the panoramic camera 10 detects the overall area of the target object and sends the overall area detected by the panoramic camera 10 to the close-up camera 20. The close-up camera 20 detects the local area of the target object within the overall area detected by the panoramic camera 10 and generates a tracking box corresponding to the local area of the target object, so as to track the target object based on the tracking box.
[0030] In this application embodiment, the panoramic camera 10 includes, but is not limited to, a 360-degree panoramic camera and a 180-degree wide-view camera, and the close-up camera 20 includes, but is not limited to, an educational recording-grade close-up camera.
[0031] Please see Figure 2 , Figure 2 This is a flowchart illustrating a target tracking method provided in an embodiment of this application.
[0032] The target tracking method is applied to a close-up camera. Specifically, the execution entity of the target tracking method is one or at least two processors of the close-up camera.
[0033] Among them, the close-up camera is connected to the panoramic camera.
[0034] like Figure 2 As shown, the target tracking method includes: Step S201: Acquire the first image of the entire area of the target object sent by the panoramic camera.
[0035] In this embodiment of the application, a panoramic camera is used to capture images and detect the overall area of a target object in the images, and generate a first bounding box corresponding to the overall area of the target object in the images. The first bounding box is used to determine the overall area range of the target object, wherein the target object includes dynamic objects, and dynamic targets include people.
[0036] In this embodiment, the panoramic camera runs a multi-object detection model at a first frame rate to identify target objects in the video frame and generate a first bounding box corresponding to the overall area of the target object. The first frame rate can be 10 frames / second, and the video frame can be a scene captured in an educational setting.
[0037] Specifically, the panoramic camera captures images of the current environment in real time and activates a multi-object detection model to detect the images of the current environment, identify target objects in the images, locate the position of the target objects in the images, determine the bounding box of the overall region of the target objects, obtain a first image of the overall region of the target objects, the first image including the first bounding box, and send the first image of the overall region of the target objects to the close-up camera.
[0038] In this embodiment, the multi-object detection model is a deep learning model used to detect human bodies or faces. When the multi-object detection model is used to detect human bodies, it learns features such as human body contours, limb structures, and dynamics through a deep learning model to detect the location of the human body in the current region. When the multi-object detection model is used to detect faces, it learns facial features through a deep learning model to locate the face region in an image or video and output the bounding box of the face. The multi-object detection model includes Faster Region-Based Convolutional Neural Network (Faster R-CNN), YOLOv5 / v8, Multi-Task Cascaded Convolutional Networks (MTCNN), RetinaFace, etc.
[0039] Step S202: Acquire the second image, and based on the second frame rate and the first bounding box, perform detection on the second image to generate a second bounding box of the local region of the target object.
[0040] Specifically, after the close-up camera receives the first image of the overall area of the target object sent by the panoramic camera, the close-up camera acquires a second image in real time. The second image includes the target object, and based on the second frame rate, it detects the local area of the target object in the second image within the first bounding box to generate a second bounding box of the local area of the target object. The target object includes a human body, and the local area includes the head and shoulder area of the human body.
[0041] In this embodiment of the application, the second frame rate is greater than the first frame rate, and the second frame rate can be 30 frames / second.
[0042] In this embodiment, the close-up camera runs a head and shoulder detection model at a second frame rate to detect local regions in the overall region and generate a second bounding box corresponding to the local region, wherein the local region includes the head and shoulder region.
[0043] In this application embodiment, the head and shoulder detection model is a lightweight target detection model. The head and shoulder detection model is used to locate and identify the region composed of the head and shoulders of a human body from an image or video. The head and shoulder detection model includes YOLO-Lite, MobileNet series, etc.
[0044] In this embodiment of the application, the panoramic camera needs to consume a lot of computing resources to run a multi-object detection model to detect the overall area of the target object. Reducing the detection frame rate of the panoramic camera can reduce the number of images processed per unit time, and can also reduce the storage space occupied by the first bounding box, thereby reducing the data transmission pressure of communicating with the close-up camera.
[0045] In this embodiment, the close-up camera runs the head and shoulder detection model at a high frame rate, which can avoid the problem of screen stuttering caused by low frame rate. Furthermore, the head and shoulder detection model only detects the head and shoulder area of the human body, which can reduce computing resources and quickly obtain the second bounding box of the local area.
[0046] Please see Figure 3 , Figure 3 yes Figure 2 A detailed flowchart of step S202 in the process.
[0047] like Figure 3 As shown, step S202 includes: Step S221: Obtain the coordinate mapping relationship between the coordinate system of the panoramic camera and the coordinate system of the close-up camera.
[0048] In this embodiment, because the physical installation positions and lens parameters (such as focal length and field of view) of the panoramic camera and the close-up camera are completely different, the pixel coordinates of the same target object captured by the panoramic camera and the close-up camera are different in their respective images. Therefore, the first bounding box sent by the panoramic camera cannot be directly reused to locate the target in the close-up camera. Therefore, before obtaining the first bounding box of the target object sent by the panoramic camera, it is necessary to determine the coordinate mapping relationship between the coordinate system of the panoramic camera and the coordinate system of the close-up camera.
[0049] In this embodiment, the intrinsic and extrinsic parameters of the panoramic camera and the close-up camera are obtained through dual-camera positioning technology. Both intrinsic and extrinsic parameters include camera intrinsic parameters and camera extrinsic parameters. The camera intrinsic parameters include focal length, principal point coordinates, and distortion parameters. The camera intrinsic parameters are used to describe the imaging characteristics of the camera itself, such as how the camera projects real-world points onto image pixels. The camera intrinsic parameters can eliminate lens distortion (such as stretching at the edges of panoramic images), unify the imaging standards of the two cameras, and avoid coordinate deviations caused by their own distortion. The camera extrinsic parameters include rotation matrix and translation vector. The camera extrinsic parameters are used to describe the relative positional relationship between the two cameras, such as the panoramic camera on the left and the close-up camera on the right, the distance between them, and the angle between them. The spatial relationship between the coordinate systems of the two cameras can be determined through the camera extrinsic parameters.
[0050] In this embodiment of the application, a homography matrix is calculated based on the intrinsic and extrinsic parameters of the panoramic camera and the close-up camera. The homography matrix is a 3*3 matrix used to transform a point in a plane from one coordinate system to another. For example, the coordinates of the points in the first bounding box are transformed from the pixel coordinate system of the panoramic camera to the pixel coordinate system of the close-up camera, so as to determine the coordinate mapping relationship between the coordinate system of the panoramic camera and the coordinate system of the close-up camera based on the homography matrix.
[0051] Step S222: Based on the coordinate mapping relationship, convert the first bounding box into a third bounding box in the coordinate system of the close-up camera.
[0052] Specifically, based on the coordinate mapping relationship between the coordinate system of the panoramic camera and the coordinate system of the close-up camera, the first bounding box is transformed into a third bounding box in the coordinate system of the close-up camera. That is, the point coordinates corresponding to the first bounding box are transformed into the point coordinates in the coordinate system of the close-up camera. By connecting multiple points in the coordinate system of the close-up camera, the third bounding box is obtained.
[0053] Step S223: Based on the second frame rate and the third bounding box, detect the second image and generate the second bounding box of the local region of the target object.
[0054] Specifically, the close-up camera runs a head and shoulder detection model based on the second frame rate, and detects the target object in the second image within the third bounding box to detect the local region of the target object within the third bounding box. Features of the local region are extracted to obtain multiple boundary points. By connecting the multiple boundary points, the second bounding box of the local region of the target object is obtained, where the local region includes the head and shoulder region.
[0055] In this embodiment, by acquiring a first image sent by a panoramic camera and converting the first bounding box in the first image into a third bounding box in the coordinate system of a close-up camera, detecting a local area of the target object within the third bounding box, and obtaining a second bounding box of the local area of the target object, the area recognized by the panoramic camera can be accurately determined, avoiding the problem of false detection by the close-up camera due to background interference (such as desks, chairs, blackboard decorations, etc. in a classroom), thereby improving detection accuracy.
[0056] Step S203: Fuse the first bounding box and the second bounding box to obtain the target tracking box.
[0057] In this embodiment of the application, the first bounding box is the bounding box in the coordinate system of the panoramic camera, the second bounding box is the bounding box in the coordinate system of the close-up camera, and the first bounding box is converted into the bounding box in the coordinate system of the close-up camera to obtain the third bounding box.
[0058] Specifically, the first and second bounding boxes are merged to obtain the target tracking box, which is also the process of merging the third and second bounding boxes to obtain the target tracking box.
[0059] Specifically, based on the Kalman filter, the third bounding box and the second bounding box are fused to obtain the target tracking box.
[0060] In this embodiment, since the frame rate used by the panoramic camera is lower than that used by the close-up camera, the update frequency of the third bounding box is lower than that of the second bounding box. This results in a mismatch in the update frequencies of the third bounding box and the second bounding box in the same time dimension, causing tracking discontinuity in the third bounding box and jitter in the actual displayed image. Therefore, it is necessary to process the third bounding box and the second bounding box using a Kalman filter.
[0061] In the embodiments of this application, the Kalman filter is used to fuse data with different characteristics to output a smooth and stable result. For example, it fuses trend information of low frame rate and real-time information of high frame rate. For example, the trend information of low frame rate includes a third bounding box, and the real-time information of high frame rate includes a second bounding box.
[0062] Specifically, based on a Kalman filter, the third and second bounding boxes are fused to obtain the target tracking box. This includes: initializing the state vector and error covariance matrix of the Kalman filter based on the third bounding box from the panoramic camera and the second bounding box from the close-up camera; initializing the state vector of the Kalman filter to include the initial position, velocity, and direction of the target object; predicting the position of the target object based on the state vector to obtain the predicted box; calculating the deviation between the predicted box and the second bounding box; calculating the Kalman gain based on the error covariance matrix; and correcting the predicted box based on the Kalman gain and the deviation to obtain the target tracking box. The Kalman gain is used to assign weights to the predicted box and the second bounding box associated with the third bounding box.
[0063] Please see Figure 4 , Figure 4 This is an example schematic diagram of determining a target tracking box provided in an embodiment of this application.
[0064] like Figure 4 As shown, region A is the area captured by the panoramic camera. A multi-object detection model is used to detect target objects within region A to determine the bounding box of the region containing each target object. Figure 4 The bounding box corresponding to region B in the image is sent to the close-up camera. Based on the bounding box corresponding to region B, the close-up camera detects the target object in that region to generate the bounding box corresponding to the close-up camera. The bounding box corresponding to region B and the bounding box corresponding to the close-up camera are fused to obtain the target tracking box. The bounding box corresponding to region C is the target tracking box.
[0065] Please see Figure 5 , Figure 5 This is an example schematic diagram of another method for determining a target tracking box provided in an embodiment of this application.
[0066] like Figure 5As shown, region D is the area captured by the panoramic camera. A multi-object detection model is used to detect target objects within region D to determine the bounding box of the region containing the target object. Figure 5 The bounding box corresponding to region H is sent to the close-up camera. The close-up camera runs the head and shoulder detection model and determines the bounding box corresponding to the head and shoulder region of the target object within the bounding box corresponding to region H. The dashed box corresponding to region E is the bounding box corresponding to the head and shoulder region. The bounding box corresponding to region H and the dashed box corresponding to region E are fused to obtain the target tracking box, which includes the bounding box corresponding to region H and the dashed box corresponding to region E.
[0067] In this embodiment, a head and shoulder detection region is run in the close-up camera, so that when the close-up camera needs to detect a local area of the target object, it can... Figure 4 The system improves the detection of the entire target area by using a close-up camera, so that the close-up camera only needs to focus on a local area, thereby reducing the computational load of the close-up camera.
[0068] Step S204: Track the target object based on the target tracking box.
[0069] In this embodiment, the close-up camera is deployed on a gimbal, which is used to fix and support the close-up camera.
[0070] Specifically, based on the target tracking frame, the gimbal is controlled in real time to rotate horizontally and / or vertically, and / or zoom, in order to track the target object.
[0071] Specifically, the position of the target object is determined based on the target tracking box, and a first image of the overall area of the target object sent by the panoramic camera and a second image sent by the close-up camera are continuously acquired. The target tracking box is continuously updated based on the first bounding box in the first image of the overall area of the target object sent by the panoramic camera and the second bounding box generated by the close-up camera. Control commands are generated in real time based on the target tracking box and sent to the gimbal to control the gimbal to rotate in real time to track the target object.
[0072] In this embodiment, if the target tracking frame moves horizontally in the video frame (horizontal movement includes moving horizontally to the left or horizontally to the right), the gimbal is controlled to rotate horizontally. If the target tracking frame moves vertically in the video frame, the gimbal is controlled to rotate vertically (vertical movement includes moving vertically upward or vertically downward). Alternatively, when the target object moves closer to or further away from the close-up camera, the close-up camera is controlled to zoom in to change the focal length of the close-up camera so as to keep the size of the target object unchanged.
[0073] Please see Figure 6 , Figure 6This is a schematic diagram illustrating the process of recovering target object tracking using a close-up camera, as provided in an embodiment of this application.
[0074] like Figure 6 As shown, the process of recovering target object tracking using this close-up camera includes: Step S601: Acquire images captured in real time by the close-up camera.
[0075] In this embodiment of the application, when the target object is occluded, or when network instability causes a delay in the panoramic camera sending the first bounding box, or when the target object moves too fast, causing the close-up camera to fail to track the target object, it is necessary to readjust the close-up camera so that the close-up camera can re-track the target object.
[0076] Specifically, the system acquires images captured in real-time by a close-up camera and performs detection on these images to determine the presence of a target object.
[0077] Step S602: Determine whether the close-up camera has detected the second bounding box of the local region of the target object.
[0078] Specifically, the head and shoulder detection model detects the real-time captured image and determines whether the close-up camera has detected the second bounding box of the local area of the target object. If it is determined that the close-up camera has not detected the second bounding box of the local area of the target object, the process jumps to step S603. If it is determined that the close-up camera has detected the second bounding box of the local area of the target object, the process jumps to step S604.
[0079] Step S603: Acquire the first image sent by the panoramic camera in real time to determine the current first bounding box of the overall region of the target object. Based on the head and shoulder detection model and the current first bounding box, detect the second image and generate the second bounding box corresponding to the head and shoulder region.
[0080] Specifically, if it is determined that the close-up camera has detected a second bounding box of a local region of the target object, target tracking is performed based on the first recovery strategy.
[0081] In this embodiment of the application, the first recovery strategy includes: acquiring a first image sent by a panoramic camera in real time to determine the current first bounding box of the overall region of the target object, and then detecting the second image based on the head and shoulder detection model and the current first bounding box to generate a second bounding box corresponding to the head and shoulder region.
[0082] Specifically, if the close-up camera detects a second bounding box of a local region of the target object, meaning the target object is not present in the detected region, the close-up camera stops detection and re-acquires the first image sent in real-time by the panoramic camera to determine the current first bounding box of the overall target object region. Based on the head and shoulder detection model and the current first bounding box, the second bounding box is detected to generate the second bounding box corresponding to the head and shoulder region. For the detailed process of generating the second bounding box, please refer to [link to documentation]. Figure 3 This will not be described in detail here.
[0083] Step S604: Continue to merge the first bounding box and the second bounding box to obtain the target tracking box, and track the target object based on the target tracking box.
[0084] Specifically, if the close-up camera detects a second bounding box of a local region of the target object, the first and second bounding boxes are then fused to obtain a target tracking box. Based on this target tracking box, the target object is tracked. For a detailed explanation of the target object tracking process, please refer to [link to relevant documentation]. Figure 2 This will not be described in detail here.
[0085] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating the process of restoring target object tracking using a panoramic camera and a close-up camera, as provided in an embodiment of this application.
[0086] like Figure 7 As shown, the process of restoring target object tracking using the panoramic camera and close-up camera includes: Step S701: Acquire images captured by the close-up camera and the panoramic camera.
[0087] In this embodiment, if the target object is occluded, or the network is unstable, or the target object moves too fast, the panoramic camera and the close-up camera may fail to track the target object. In this case, the panoramic camera and the close-up camera need to be readjusted so that the panoramic camera and the close-up camera can re-detect the target object and continue to track the target object.
[0088] Specifically, it acquires images taken by close-up cameras and panoramic cameras.
[0089] Step S702: Determine whether the close-up camera has detected the second bounding box of the local region of the target object.
[0090] Specifically, the image captured by the close-up camera is detected to determine whether the close-up camera has detected the second bounding box of the local area of the target object. If it is determined that the close-up camera has not detected the second bounding box of the local area of the target object, the process jumps to step S703. If it is determined that the close-up camera has detected the second bounding box of the local area of the target object, the process jumps to step S705.
[0091] Step S703: Determine whether the panoramic camera has detected the first bounding box of the entire area of the target object.
[0092] Specifically, if it is determined that the close-up camera did not detect the second bounding box of the local area of the target object, the image captured by the panoramic camera is detected to determine whether the panoramic camera detected the first bounding box of the overall area of the target object. If it is determined that the panoramic camera did not detect the first bounding box of the overall area of the target object, the process jumps to step S704. If it is determined that the panoramic camera detected the first bounding box of the overall area of the target object, the process jumps to step S705.
[0093] In this embodiment of the application, when the panoramic camera does not detect the first bounding box of the entire area of the target object, it acquires 5 consecutive frames of images captured by the panoramic camera in real time, detects whether the target object exists in the 5 consecutive frames, and if the target object is not found in the 5 consecutive frames, the panoramic camera stops sending the first bounding box to the close-up camera and jumps to step S704.
[0094] Step S704: Control the panoramic camera to stop sending the first image to the close-up camera until the panoramic camera re-detects the first bounding box of the overall region of the target object. Then, control the panoramic camera to re-send the first image of the overall region of the target object to the close-up camera. Based on the head and shoulder detection model and the first bounding box in the first image, detect the second image and generate the second bounding box corresponding to the head and shoulder region.
[0095] Specifically, when it is determined that the close-up camera did not detect the second bounding box of the local area of the target object, and the panoramic camera did not detect the first bounding box of the overall area of the target object, target tracking is performed based on the second recovery strategy.
[0096] In this embodiment, the second recovery strategy includes controlling the panoramic camera to stop sending the first image to the close-up camera until the panoramic camera re-detects the first bounding box of the overall region of the target object, and then controlling the panoramic camera to re-send the first image of the overall region of the target object to the close-up camera, so as to detect the second image based on the head and shoulder detection model and the first bounding box in the first image, and generate the second bounding box corresponding to the head and shoulder region.
[0097] Step S705: Acquire a first image of the overall region of the target object sent by the panoramic camera in real time, and detect the second image based on the head and shoulder detection model and the first bounding box in the first image to generate a second bounding box corresponding to the head and shoulder region.
[0098] Specifically, when it is determined that the close-up camera has not detected the second bounding box of the local area of the target object, the first bounding box of the overall area of the target object sent by the panoramic camera is acquired in real time. Based on the head and shoulder detection model, the current first bounding box is detected to generate the second bounding box corresponding to the head and shoulder area.
[0099] Alternatively, if it is determined that the close-up camera did not detect the second bounding box of the local area of the target object, and the panoramic camera detected the first bounding box of the overall area of the target object, then the first image of the overall area of the target object sent by the panoramic camera is acquired in real time. Based on the head and shoulder detection model and the bounding box in the first image, the second image is detected to generate the second bounding box corresponding to the head and shoulder area.
[0100] In the embodiments of this application, the close-up camera and panoramic camera are restored to track the target object in a timely manner through the first recovery strategy and the second recovery strategy, which can quickly retrieve the target object. This improves the traditional solution of blindly tracking the target object by controlling the gimbal, thereby improving the stability of target tracking.
[0101] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a target tracking system provided in an embodiment of this application.
[0102] like Figure 8 As shown, the target tracking system 800 includes a panoramic camera 10, a close-up camera 20, and a gimbal 30.
[0103] In this embodiment, the panoramic camera 10 communicates with the close-up camera 20 via a wired interface or a wireless network.
[0104] The panoramic camera 10 is used to acquire a first image, detect target objects in the first image, and determine a first bounding box of the overall region of the target object.
[0105] Close-up camera 20 is used to acquire a first image of the target object sent by panoramic camera 10, and to acquire a second image based on a second frame rate, and to detect the second image based on a first bounding box in the first image to obtain a second bounding box of a local region of the target object.
[0106] The gimbal 30 is used to fix and support the close-up camera 20 for tracking the target object.
[0107] In this embodiment, the panoramic camera 10 captures images in real time, performs detection on the captured images based on a multi-object detection model to generate a first bounding box of the overall region of the target object, and sends the first bounding box of the overall region of the target object to the close-up camera 20. The close-up camera 20 detects the first bounding box of the overall region of the target object sent by the panoramic camera 10 based on a head and shoulder detection model to generate a second bounding box of the head and shoulder region of the target object. The close-up camera 20 fuses the first bounding box and the second bounding box based on a Kalman filter to obtain a target tracking box. The target object is tracked based on the target tracking box.
[0108] In this embodiment, a first bounding box of the overall region of the target object is detected by a multi-object detection model, and a second bounding box of the local region of the target object is detected by a head and shoulder detection model, so as to quickly generate a target tracking box. This allows the close-up camera to focus only on the local region of the target object, reducing the area analyzed by the close-up camera, reducing the computational load of the close-up camera, and improving the accuracy of the close-up camera in tracking the target.
[0109] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a target tracking device provided in an embodiment of this application.
[0110] like Figure 9 As shown, the target tracking device 900 includes: The acquisition unit 901 is used to acquire a first image of the overall area of the target object sent by the panoramic camera. The first image includes a first bounding box, which is obtained by the panoramic camera detecting the video frame at a first frame rate. The first bounding box is used to determine the overall area range of the target object.
[0111] The bounding box generation unit 902 is used to acquire a second image and, based on a second frame rate and a first bounding box, detect the second image to generate a second bounding box of a local region of the target object, wherein the second frame rate is greater than the first frame rate.
[0112] The fusion unit 903 is used to fuse the first bounding box and the second bounding box to obtain the target tracking box.
[0113] The tracking unit 904 is used to track the target object based on the target tracking box.
[0114] In the embodiments of this application, the tracking accuracy is improved by the cooperation of the various units of the target tracking device.
[0115] In this embodiment, the target tracking device can be a software module, which includes several instructions stored in a memory. The processor can access the memory and execute the instructions to complete the target tracking methods of the above embodiments.
[0116] In the embodiments of this application, the target tracking device can also be constructed from hardware devices. For example, the target tracking device can be constructed from one or more chips, and the chips can work together to complete the target tracking methods described in the above embodiments. Furthermore, the target tracking device can also be constructed from various logic devices, such as general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, ARM (Acorn RISC Machine) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0117] The target tracking device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0118] The target tracking device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0119] The target tracking device provided in this application embodiment can achieve... Figure 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0120] It should be noted that the above-described device can execute the target tracking method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the device embodiments can be found in the target tracking method provided in the embodiments of this application.
[0121] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of a close-up camera provided in an embodiment of this application.
[0122] like Figure 10 As shown, the close-up camera 20 includes one or more processors 21 and a memory 22. Among them, Figure 10 Take a processor 21 as an example.
[0123] Processor 21 and memory 22 can be connected via a bus or other means. Figure 10 Taking the example of a connection between China and Israel via a bus.
[0124] A processor is configured to execute a target tracking method in any embodiment of this application, the method being applied to a close-up camera connected to a panoramic camera, the method comprising: A first image of the overall region of a target object sent by a panoramic camera is acquired, wherein the first image includes a first bounding box, which is obtained by the panoramic camera detecting the video frame at a first frame rate, and the first bounding box is used to determine the overall region range of the target object. A second image is acquired, and based on a second frame rate and a first bounding box, the second image is detected to generate a second bounding box of the local region of the target object, wherein the second frame rate is greater than the first frame rate; By fusing the first bounding box and the second bounding box, the target tracking box is obtained; The target object is tracked based on the target tracking bounding box.
[0125] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the target tracking method in this embodiment of the invention. The processor 21 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 22, thereby implementing the target tracking method of the above-described method embodiment.
[0126] Memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 22 may optionally include memory remotely located relative to processor 21. Examples of the above-described networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0127] One or more modules are stored in memory 22. When executed by one or more processors 21, they perform the target tracking method in any of the above method embodiments, for example, the method described above. Figure 2 The steps shown.
[0128] This application also provides a computer program product, which includes one or more lines of program code stored in a non-volatile computer-readable storage medium. The processor of an electronic device reads the program code from the non-volatile computer-readable storage medium and executes the program code to complete the steps of the target tracking method provided in the above embodiments.
[0129] Based on the above description of the embodiments, those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program or program code related to hardware. The program can be stored in a non-volatile computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0130] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The non-volatile computer-readable storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations as described above in different aspects of this application, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A target tracking method, characterized in that, Applied to a close-up camera connected to a panoramic camera, the method includes: A first image of the overall area of the target object sent by the panoramic camera is obtained, wherein the first image includes a first bounding box, the first bounding box is obtained by the panoramic camera detecting the video frame at a first frame rate, and the first bounding box is used to determine the overall area range of the target object. A second image is acquired, and based on a second frame rate and a first bounding box, the second image is detected to generate a second bounding box of a local region of the target object, wherein the second frame rate is greater than the first frame rate; The first bounding box and the second bounding box are merged to obtain the target tracking box; The target object is tracked based on the target tracking bounding box.
2. The method according to claim 1, characterized in that, Before acquiring the first image of the target object sent by the panoramic camera, the method further includes: Obtain the coordinate mapping relationship between the coordinate system of the panoramic camera and the coordinate system of the close-up camera; The step of detecting the second image based on the second frame rate and the first bounding box to generate a second bounding box for the local region of the target object includes: Based on the coordinate mapping relationship, the first bounding box is converted into a third bounding box in the coordinate system of the close-up camera; Based on the second frame rate and the third bounding box, the second image is detected to generate a second bounding box of the local region of the target object.
3. The method according to claim 2, characterized in that, The process of fusing the first bounding box and the second bounding box to obtain the target tracking box includes: Based on the Kalman filter, the third bounding box and the second bounding box are fused to obtain the target tracking box.
4. The method according to claim 1, characterized in that, The panoramic camera runs a multi-object detection model at a first frame rate to identify target objects in the video frame and generate a first bounding box corresponding to the overall area of the target object. The close-up camera runs a head and shoulder detection model at a second frame rate to detect local regions within the overall region and generate a second bounding box corresponding to the local region, wherein the local region includes the head and shoulder region.
5. The method according to claim 4, characterized in that, The method further includes: If the close-up camera does not detect the second bounding box of the local region of the target object, target tracking is performed based on a first recovery strategy, wherein the first recovery strategy includes: The system acquires the first image sent by the panoramic camera in real time to determine the current first bounding box of the overall region of the target object. Based on the head and shoulder detection model and the current first bounding box, the system detects the second image and generates the second bounding box corresponding to the head and shoulder region.
6. The method according to claim 4, characterized in that, The method further includes: If the panoramic camera fails to detect a first bounding box of the entire region of the target object, and the close-up camera fails to detect a second bounding box of a local region of the target object, then target tracking is performed based on a second recovery strategy, wherein the second recovery strategy includes: The panoramic camera is controlled to stop sending the first side image to the close-up camera until the panoramic camera re-detects the first bounding box of the overall region of the target object. Then, the panoramic camera is controlled to re-send the first image of the overall region of the target object to the close-up camera. Based on the head and shoulder detection model and the first bounding box in the first image, the second image is detected to generate the second bounding box corresponding to the head and shoulder region.
7. The method according to any one of claims 1-6, characterized in that, The close-up camera is deployed on a gimbal, which is used to fix and support the close-up camera; The step of tracking the target object based on the target tracking bounding box includes: Based on the target tracking frame, the gimbal is controlled in real time to rotate horizontally and / or vertically and / or zoom to track the target object.
8. A close-up camera, characterized in that, include: At least one processor; At least one memory for storing at least one program; When at least one of the programs is executed by at least one of the processors, the at least one of the processors implements the method of any one of claims 1-7.
9. A target tracking system, characterized in that, include: A panoramic camera is used to acquire a first image, detect target objects in the first image, and determine a first bounding box of the overall region of the target object; The close-up camera as described in claim 8 is used to acquire a first image of a target object sent by the panoramic camera, and to acquire a second image based on a second frame rate, and to detect the second image based on a first bounding box in the first image to obtain a second bounding box of a local region of the target object; A gimbal is used to fix and support the close-up camera for tracking the target object.
10. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to perform the method as described in any one of claims 1-7.