Following shot device control method, device and storage medium
By integrating motion and perspective adjustment components into the tracking device, autonomous pose adjustment is achieved using distance and coordinate information, thus solving the stability and accuracy problems of shooting equipment caused by manual operation and realizing the technical effect of autonomous motion tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN EMEET TECH CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-08
AI Technical Summary
Existing mobile tracking and shooting equipment relies on manual operation, which leads to fatigue of the shooting personnel, and makes it difficult to guarantee stability and accuracy. Auxiliary devices cannot achieve real-time judgment and manual adjustment.
The tracking device, which includes a motion component, a perspective adjustment component, and a shooting component, adjusts the initial pose using distance and coordinate information, and adjusts the motion component and the shooting component based on the pose offset obtained in a loop, thereby achieving autonomous motion tracking.
It enables autonomous movement and tracking of the shooting equipment, improving shooting accuracy and stability, and reducing the fatigue and tediousness of manual operation.
Smart Images

Figure CN122002137A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent tracking and shooting technology, and in particular to a control method, device and storage medium for a tracking shooting device. Background Technology
[0002] Currently, in mobile tracking shooting, manual operation of shooting equipment is still the core method. The shooting personnel need to hold or operate the shooting equipment throughout the process and adjust the spatial position and shooting angle of the equipment in real time to complete the tracking shooting of dynamic targets.
[0003] To alleviate the burden of equipment load and the cumbersome operation for filming personnel, auxiliary devices such as exoskeletons and dedicated control controllers have emerged in related technologies. However, these auxiliary methods can only reduce the physical burden on filming personnel at a physical level; they cannot fundamentally replace real-time judgment and manual adjustment of movements. The accuracy and consistency of filming are highly dependent on the operator's proficiency. Furthermore, continuous manual operation is prone to fatigue, directly leading to difficulties in ensuring the stability of tracking filming. Summary of the Invention
[0004] The main purpose of this application is to provide a control method, device and storage medium for a tracking device, which aims to solve the technical problem of the lack of motion tracking capability in the prior art.
[0005] To achieve the above objectives, this application provides a tracking shooting device, which includes a moving component, a viewing angle adjustment component, and a shooting component. The moving component is provided with a moving unit and a driving unit. The first end of the viewing angle adjustment component is connected to the moving component, and the second end is connected to the shooting component. The viewing angle adjustment component includes a telescopic unit so that the relative distance between the shooting component and the moving component can be adjusted by the telescopic unit.
[0006] In one embodiment, the second end of the viewing angle adjustment component is provided with a rotation unit so that the relative angle between the shooting component and the viewing angle adjustment component can be adjusted by rotating the rotation unit.
[0007] Furthermore, to achieve the above objectives, this application also provides a control method for a tracking camera, the control method comprising:
[0008] In response to the target object, determine the distance information between the shooting component of the tracking device and the target object and / or the coordinate information of the target object in the field of view of the shooting component; Adjust the tracking device to its initial pose based on distance and / or coordinate information; The pose offset of the target object is obtained in a loop. Based on the pose offset obtained each time, the movement component of the tracking device and / or the shooting component of the tracking device are adjusted.
[0009] In one embodiment, adjusting the tracking device to an initial pose based on distance information and / or coordinate information includes: Based on the difference between the distance information and the preset target distance, the moving components of the following camera are controlled to move by the distance corresponding to the difference, so as to adjust the following camera to the initial pose; and / or, Based on the horizontal deviation between the coordinate information and the coordinates of the field of view center, control the rotation angle of the tracking device's rotating unit or moving component corresponding to the horizontal deviation, thereby adjusting the tracking device to its initial pose; and / or, Based on the vertical deviation between the coordinate information and the coordinates of the center of the field of view, the telescopic unit of the following camera is controlled to adjust the distance corresponding to the horizontal deviation, so as to adjust the following camera to the initial pose.
[0010] In one embodiment, the pose offset is distance information. The pose offset of the target object is obtained cyclically, and the moving components of the tracking device are adjusted according to the pose offset obtained each time, including: The distance information between the shooting components of the tracking device and the target object is collected based on the first frequency. The coordinate information of the target object in the field of view of the shooting component is collected based on the second frequency; Historical location data of the target object in the device coordinate system is constructed based on distance and coordinate information. Determine the movement speed corresponding to historical location data; The predicted movement path of the target object is obtained based on the fitting model corresponding to the movement speed. Adjust the movement components of the tracking device based on the predicted movement path.
[0011] In one embodiment, obtaining the predicted movement path of the target object based on the fitting model corresponding to the movement speed includes: If the moving speed is less than or equal to the speed threshold, the k-th degree polynomial curve is constructed from historical position data using the least squares method. The predicted movement path of the target object is obtained based on the solution results.
[0012] In one embodiment, obtaining the predicted movement path of the target object based on the fitting model corresponding to the movement speed includes: If the moving speed is greater than the speed threshold, construct the state vector of the target object based on historical location data; The state equation is constructed based on the state vector, the uniform acceleration transfer matrix, and the first covariance matrix. The predicted movement path of the target object is obtained based on the state equation and historical location data.
[0013] In one embodiment, the pose offset is a coordinate deviation. The pose offset of the target object is obtained iteratively, and the shooting components of the tracking device are adjusted according to the pose offset obtained each time, including: The target object's coordinate information in the field of view of the shooting component is collected according to the second frequency. Based on the coordinate deviation between the coordinate information collected each time and the coordinates of the center of the field of view, at least one of the telescopic unit, the rotation unit, and the moving component is adjusted to adjust the angle of view of the shooting component.
[0014] In one embodiment, before determining the distance information between the shooting component of the tracking device and the target object and / or the coordinate information of the target object in the field of view of the shooting component in response to the target object, the method includes: In response to the gesture information recognized by the shooting component, determine the target object to which the gesture information is pointing; and / or, In response to the received positioning signal, determine the target object to which the positioning signal points.
[0015] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the control method of the tracking device described above.
[0016] This application provides a control method for a tracking camera. First, in response to a target object, the method determines the distance information between the camera component of the tracking camera and the target object, and / or the coordinate information of the target object within the field of view of the camera component. Based on the distance information and / or coordinate information, the tracking camera is adjusted to an initial pose. The method then iteratively acquires the pose offset of the target object, and based on each acquired pose offset, adjusts the moving component and / or the camera component of the tracking camera. By determining the target distance and / or coordinate information for initial pose calibration, and adjusting the moving component and / or the camera component based on the iteratively acquired pose offsets, the method achieves the technical effect of autonomous movement tracking of the camera. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the device structure provided in Embodiment 1 of the tracking device of this application; Figure 2 A flowchart illustrating the control method for the tracking camera device in Embodiment 3 of this application; Figure 3 A flowchart illustrating the control method for the tracking camera device in Embodiment 9 of this application; Figure 4 This is a schematic diagram of the following equipment used in this application.
[0020] Label Explanation: 1. Moving unit; 2. Driving unit; 3. Telescopic unit; 4. Shooting component.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] Currently, mobile tracking photography still relies primarily on manual operation of the equipment. Photographers must hold or manipulate the equipment throughout the entire process, constantly adjusting its spatial position and shooting angle to track moving targets. To alleviate the burden on photographers and the tediousness of operation, auxiliary devices such as exoskeletons and dedicated controllers have emerged. However, these aids only physically reduce the burden on photographers and cannot fundamentally replace real-time judgment and manual adjustments. The accuracy and consistency of the footage remain highly dependent on the operator's skill level. Furthermore, continuous manual work easily leads to fatigue, directly compromising the stability of tracking photography.
[0025] The main solution of this application is as follows: First, in response to the target object, determine the distance information between the shooting component of the tracking device and the target object, and / or the coordinate information of the target object in the field of view of the shooting component; adjust the tracking device to an initial pose based on the distance information and / or coordinate information; cyclically acquire the pose offset of the target object, and adjust the moving component and / or the shooting component of the tracking device based on each acquired pose offset. By determining the target distance and / or coordinate information for initial pose calibration, and adjusting the moving component and / or the shooting component based on the cyclically acquired pose offset, the technical effect of achieving autonomous movement tracking of the shooting device is achieved.
[0026] It should be noted that the executing entity in this embodiment can be a tracking camera device, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a control device capable of performing the above functions. This embodiment does not specifically limit this. The following uses a tracking camera device as the executing entity to describe this embodiment and the following embodiments.
[0027] Based on this, Embodiment 1 of this application proposes a tracking camera device, please refer to... Figure 1 , Figure 1 This is a schematic diagram of the device structure provided in Embodiment 1 of the following shooting device of this application. The following shooting device includes a moving component, a viewing angle adjustment component, and a shooting component. The moving component is provided with a moving unit and a driving unit. The first end of the viewing angle adjustment component is connected to the moving component, and the second end is connected to the shooting component. The viewing angle adjustment component includes a telescopic unit so that the relative distance between the shooting component and the moving component can be adjusted by the telescopic unit.
[0028] In this embodiment, the mobile component includes a mobile unit, a drive unit, and an environmental sensing unit. The mobile unit uses four omnidirectional wheels for movement, supporting in-situ rotation, crab-like movement, and omnidirectional translation, allowing for flexible adjustment of the device's spatial position and orientation within a two-dimensional plane. The drive unit has a built-in motor and motion controller, capable of receiving commands from the device's central controller to drive the omnidirectional wheels to perform forward, backward, and turning movements. The environmental sensing unit is equipped with 12 ultrasonic radars and 1 2D lidar. The 12 ultrasonic radars are evenly distributed around the mobile component, each with a horizontal coverage of 120°, enabling real-time omnidirectional 360-degree detection of obstacles within a range of 0.1-5 meters and outputting obstacle distance information. The 2D lidar is installed at the front or top of the chassis for environmental scanning within a range of 0-20 meters, generating point cloud data to create an environmental map. The imaging component includes an image acquisition unit and a distance sensing unit, employing a binocular camera setup: a wide-angle camera and a telephoto camera. The wide-angle camera has a field of view ≥120°, used for wide-range scene coverage of 0-5 meters, capturing the overall environment of the target object; the telephoto camera has a focal length ≥50mm, used for detail capture of 5-20 meters. Together, they achieve full-range imaging coverage. The distance sensing unit incorporates a TOF (Time of Flight) sensor array, such as a VCSEL (Vertical-Cavity Surface-Emitting Laser) light source and a SPAD (Single-Photon Avalanche Diode) detector. It outputs data at 8x8 points, covering 120 degrees horizontally and vertically. It can detect the straight-line distance between the target and the device in real time, with a detection accuracy ≤1cm and a coverage range of 0.3-10 meters, outputting target distance data. Before leaving the factory, the area array TOF sensor establishes a mapping relationship with the imaging position of the binocular camera through a calibration method that corresponds point cloud to camera pixels, achieving the correlation and fusion of distance data and image position data. The viewing angle adjustment component includes a telescopic unit, which adopts an electrically telescopic structure, such as a multi-stage electric push rod or lead screw drive mechanism, with an adjustable height range of 0.5-2 meters. By automatically telescopically adjusting the relative distance between the shooting component and the moving component, and in conjunction with the omnidirectional movement of the moving chassis, the position of the camera device can be adjusted in three-dimensional space to meet the shooting needs of targets at different heights. The tracking device, through the moving component, the viewing angle adjustment component with telescopic unit, and the shooting component, realizes the adjustment of the spatial position of the shooting component and the movement tracking of the device, effectively improving the accuracy of tracking shooting.
[0029] As an optional implementation, the top of the moving component is equipped with a standardized mechanical interface, such as a flange, for fixed connection to the first end of the viewing angle adjustment component, and the shooting component is connected to the second end of the viewing angle adjustment component. In the environmental perception unit, signals from the ultrasonic radar, 2D lidar, and drive unit are all connected to the central controller to receive and provide feedback on control commands. In the viewing angle adjustment component, signals from the telescopic rod displacement sensor are connected to the central controller to receive height adjustment commands and provide feedback on the current status. In the shooting component, signals from the binocular camera and the area array TOF sensor are both connected to the central controller to output images and distance data, and to receive shooting parameter adjustment commands, such as focal length and exposure time. Through the stable connection of each component and its connection to the central controller, the accuracy, real-time performance, and stability of the device's autonomous movement, viewing angle adjustment, and intelligent tracking are ensured.
[0030] Based on Embodiment 1, in Embodiment 2 of this application, the same or similar content as in Embodiment 1 can be referred to the above description, and will not be repeated hereafter. Furthermore, a rotation unit is provided at the second end of the viewing angle adjustment component, so that the relative angle between the shooting component and the viewing angle adjustment component can be adjusted by rotating the unit.
[0031] In this embodiment, a rotating unit is provided at the second end of the viewing angle adjustment component. The rotating unit connects the top of the telescopic unit of the viewing angle adjustment component to the bottom of the shooting component. It can adjust the relative angle between the shooting component and the viewing angle adjustment component through its own rotation, so as to realize the 360° rotation adjustment of the shooting component and further expand the adjustment range and flexibility of the shooting angle.
[0032] As an optional implementation, the second end of the viewing angle adjustment component is provided with a rotation unit.
[0033] Specifically, one end of the rotating unit is fixedly connected to the top of the viewing angle adjustment component, and the other end is fixedly connected to the bottom of the shooting component. That is, the rotating unit is separately mounted at the end of the second end of the viewing angle adjustment component. The fixed part of the rotating unit is fixedly connected to the second end of the viewing angle adjustment component, and the rotating part of the rotating unit is fixedly connected to the shooting component. This independently configured rotating unit allows for angle adjustment of the shooting component relative to the viewing angle adjustment component. No additional adjustment components are needed, which effectively simplifies the overall structure of the viewing angle adjustment component, reduces hardware assembly difficulty and subsequent maintenance costs, and enables independent angle adjustment of the shooting component relative to the viewing angle adjustment component, ensuring that the target object is centered in the shooting field of view.
[0034] As another alternative implementation, the viewing angle adjustment component includes a telescopic unit and a rotation unit.
[0035] Specifically, the viewing angle adjustment component includes a telescopic unit and a rotating unit. The telescopic unit and the rotating unit are connected sequentially. The bottom of the telescopic unit is connected to a moving component, and the top of the telescopic unit is connected to the rotating unit. The end of the rotating unit furthest from the telescopic unit is connected to the shooting component. The telescopic unit adjusts the height of the shooting component, while the rotating unit adjusts its angle. Together, they achieve multi-degree-of-freedom shooting angle adjustment. By integrating the telescopic and rotating units within the viewing angle adjustment component, the vertical height of the shooting component can be flexibly adjusted via the telescopic unit, and the horizontal angle of the shooting component can be adjusted via the rotating unit, achieving precise two-way control of both height and angle.
[0036] Furthermore, to adapt to different scenarios, some modules in the above embodiments can be replaced. For example, the omnidirectional wheel can be replaced with a lower-cost differential wheel, but steering requires dual-motor differential control, which sacrifices the stationary rotation and crabbing functions. Alternatively, it can be replaced with a more precise steering wheel, which can support more flexible steering control. The area array TOF sensor can be replaced with monocular or binocular visual ranging, but it relies on the image parallax of the binocular camera to calculate the distance, and its accuracy is slightly lower than that of the area array TOF sensor. The electric telescopic structure can be replaced with a hydraulic telescopic structure with a stronger load capacity, suitable for scenarios that need to support heavy cameras, such as cinema-grade cameras, but the response speed is slightly slower. It can also be replaced with a robotic arm structure. The above are only preferred embodiments of this application and do not limit the patent scope of this application. All equivalent substitutions, conventional modifications, and similar structural designs using the above modules are similarly included within the patent processing scope of this application.
[0037] Based on any of the above embodiments of this application, Embodiment 3 of this application proposes a control method for a tracking camera. Please refer to... Figure 2 , Figure 2 This is a flowchart illustrating the control method for the tracking camera device according to Embodiment 3 of this application. The control method for the tracking camera device includes steps S10 to S30: Step S10: In response to the target object, determine the distance information between the shooting component of the tracking device and the target object and / or the coordinate information of the target object in the field of view of the shooting component.
[0038] In this embodiment, distance information refers to the straight-line physical distance between the shooting component and the target object, which is the actual distance measurement value in three-dimensional space, used to reflect the distance between the tracking device and the target object. Coordinate information refers to the pixel coordinates of the target object in the image coordinate system of the shooting component, with the unit being pixels, that is, the horizontal and vertical pixel positions of the target object within the image frame, used to reflect the degree of offset of the target object relative to the center of the shooting field of view.
[0039] In this embodiment, the distance information to the target object is determined by a distance sensing unit, such as a TOF area array sensor, in the shooting component of the tracking device, and / or the coordinate information of the target object within the field of view of the shooting component is acquired by an image acquisition unit, such as a binocular camera. By acquiring the target distance and / or the coordinate information within the field of view, data is provided for the initial pose adjustment of the tracking device.
[0040] As one implementation method, the distance information between the shooting components of the tracking device and the target object is determined.
[0041] Specifically, the linear physical distance between the tracking device and the target object is directly measured using a distance sensing unit within the camera's imaging components, such as an area-array TOF sensor, a laser ranging module, or an infrared ranging module. Taking an area-array TOF sensor as an example, it emits modulated light pulses and measures the flight time of the light pulses reflected back from the target, calculating the linear distance between the tracking device and the target. This distance value, characterizing the proximity of the two objects, serves as the basis for the tracking device to adjust its posture, such as moving forward and backward or maintaining distance.
[0042] As another implementation method, the coordinate information of the target object in the field of view of the shooting component is determined.
[0043] Specifically, the image acquisition unit in the shooting component of the tracking device, such as a binocular camera, captures images of the target object within the field of view of the shooting component. The pixel coordinates of the target object in the image coordinate system are then calculated, thus obtaining the target object's coordinate information within the field of view of the shooting component. Taking a binocular camera as an example, image recognition algorithms, such as the YOLO object detection model, identify the target object from the image and determine its contour or feature points. The algorithm outputs the target's position in the field of view coordinate system, typically represented by the pixel coordinates (x, y) of the center point of the bounding box or a specific feature point. Using this positional information in a two-dimensional plane, the target's offset direction and magnitude within the image can be obtained, thereby driving the movement component and the viewpoint adjustment component to center the target.
[0044] As another implementation method, the distance information between the shooting component of the tracking device and the target object and the coordinate information of the target object in the field of view of the shooting component are determined.
[0045] Specifically, the distance sensing unit in the tracking device's shooting component measures the flight time from signal transmission to reception, thereby calculating the straight-line distance to the target object. The image acquisition unit in the shooting component acquires an image of the target object within the shooting component's field of view. Based on image recognition algorithms, the pixel coordinates of the target object in the image coordinate system are determined, thus obtaining the target object's coordinate information within the shooting component's field of view. By simultaneously acquiring target distance and coordinate information, multi-dimensional pose data is provided for the tracking device's pose adjustment, improving the accuracy of initial positioning and tracking adjustments.
[0046] The above are only three feasible implementation methods of step S10 provided in this embodiment. This embodiment does not specifically limit the specific implementation method of step S10.
[0047] Furthermore, the data obtained in step S10 corresponds to the follow-up adjustments performed in steps S20 and S30. That is, if distance information and coordinate information are obtained in step S10, the following camera is adjusted in steps S20 and S30 based on the distance information and coordinate information obtained in S10.
[0048] Step S20: Adjust the following camera to its initial pose based on distance information and / or coordinate information.
[0049] In this embodiment, the initial pose refers to the baseline pose that meets the preset tracking conditions before the tracking device enters the formal tracking state. For example, the distance between the tracking device and the target object is maintained at a preset tracking distance, and the target object is located at the center of the shooting field of view of the shooting component, so that the device has an initial state of stable tracking of the target object.
[0050] In this embodiment, the moving component of the tracking device is controlled to move based on a comparison between the distance information and a preset target distance or a preset target distance range, adjusting the distance between the tracking device and the target object to an initial pose, and / or adjusting the tracking device to an initial pose based on coordinate information and the coordinate information of the field of view center. Adjusting the tracking device to an initial pose based on distance information and / or coordinate information can establish a stable and reliable tracking reference state, providing a good foundation for subsequent tracking.
[0051] As one implementation method, the rotating unit or moving component of the tracking camera is controlled to adjust the horizontal deviation.
[0052] Specifically, the first end of the perspective adjustment component is connected to the movement component, and the second end is connected to the shooting component. The second end of the perspective adjustment component is equipped with a rotation unit. Based on the horizontal deviation between the coordinate information and the coordinates of the field of view center, the rotation unit or movement component of the tracking device is adjusted. Specifically, the rotation direction of the movement component (clockwise or counterclockwise) is determined based on the sign of the horizontal deviation. By rotating the tracking device, the target is horizontally centered in the frame. Alternatively, based on the horizontal deviation, the central controller sends a rotation command to the rotation unit to perform a horizontal rotation, directly adjusting the shooting angle of the shooting component to horizontally center the target in the frame. This allows the tracking device to complete the initial horizontal deviation adjustment. By flexibly selecting and adjusting the rotation unit or movement component based on the horizontal deviation, the target can be quickly horizontally centered in the frame, improving the efficiency and flexibility of initial horizontal alignment.
[0053] As another implementation method, the telescopic unit of the tracking camera is controlled to adjust the vertical deviation.
[0054] Specifically, the perspective adjustment component connects the movement component and the shooting component. The perspective adjustment component includes a telescopic unit, within which a built-in displacement sensor connects to the central controller, receiving height adjustment commands and providing feedback on the current height status. Based on the vertical deviation between the coordinate information and the coordinates of the field of view center, the telescopic unit adjusts its height, thereby adjusting the vertical height of the shooting component and enabling the tracking device to complete the initial pose vertical deviation adjustment. By adjusting the shooting component height according to the vertical deviation using the telescopic unit, the target can be quickly vertically centered in the frame.
[0055] Step S30: Loop through the positional offset of the target object, and adjust the moving components and / or shooting components of the tracking device according to the positional offset obtained each time.
[0056] In this embodiment, by acquiring the distance and coordinate information of the target object in real time, the pose offset of the target object is calculated. The pose offset includes distance information and coordinate deviation. The predicted movement path of the target object is predicted using historical position data, and the moving component of the tracking device is adjusted to the predicted position. The viewing angle of the shooting component is adjusted by the coordinate deviation between the coordinate information acquired each time and the coordinates of the center of the field of view, so that the target is located in the center of the field of view. The adjustment operation is executed cyclically, enabling the tracking device to respond to the pose offset of the target in real time and achieve real-time tracking of the target object by dynamically adjusting its own pose.
[0057] As one implementation method, the movement method is determined based on the target object's movement speed.
[0058] Specifically, the system determines whether the target object's moving speed is less than a speed threshold. If it is, the tracking device is adjusted to the same initial pose based on distance and coordinate information. If it is greater than a speed threshold, the motion prediction algorithm predicts the movement path, and the tracking device is then moved to the next position. In other words, when the movement speed is low, it is determined to be low-speed or constant-speed motion, making the target highly predictable. Tracking is then performed using the initial pose adjustment based on distance and coordinate information. If the target's movement speed is high, its movement may be more complex. Relying solely on pose adjustment can easily lead to tracking lag. In this case, motion prediction effectively compensates for the delay and tracking lag, allowing the tracking device to predict the target's movement and achieve accurate and continuous advance tracking. Choosing between initial pose adjustment and path prediction based on the target object's movement speed ensures tracking accuracy while improving tracking response efficiency, achieving a balance between precision and efficiency.
[0059] As another implementation method, a fitting model is determined based on the target object's moving speed to obtain the predicted moving path of the target object.
[0060] Specifically, when the target object moves at a low speed, a polynomial fitting model is used. The k-th degree polynomial curve, constructed from historical position data, is solved using the least squares method. The resulting polynomial expression is the predicted movement path, which is then used to adjust the movement components of the tracking device. When the target object moves at a high speed, a Kalman filter model is used. Based on historical position data and the target's state information at a certain moment, predictions are made for the next moment to obtain the predicted movement path. Selecting different fitting models according to the target object's different movement speeds balances the fitting accuracy for low-speed scenes with the real-time prediction performance for high-speed scenes, improving the accuracy and applicability of the predicted movement path.
[0061] As another implementation method, safe obstacle avoidance is based on predicted movement paths.
[0062] Specifically, during the movement of the tracking device, the environmental perception unit in the moving component continuously updates the environmental map and generates obstacle avoidance paths based on the predicted movement path using a path planning algorithm, controlling the moving component to bypass obstacles. Simultaneously, when the environmental perception unit detects an obstacle, it prioritizes triggering local obstacle avoidance to maintain a safe distance from the obstacle. This proactive obstacle avoidance based on predicted paths combined with path planning algorithms effectively ensures the safety of the tracking device, avoids collisions, and guarantees tracking continuity.
[0063] This embodiment provides a control method for a tracking camera. First, in response to a target object, the method determines the distance information between the camera component of the tracking camera and the target object, and / or the coordinate information of the target object within the field of view of the camera component. Based on the distance information and / or coordinate information, the tracking camera is adjusted to an initial pose. The pose offset of the target object is then repeatedly acquired, and based on each acquired pose offset, the moving component and / or the camera component of the tracking camera are adjusted. By determining the target distance and / or coordinate information for initial pose calibration, and adjusting the moving component and / or the camera component based on the repeatedly acquired pose offsets, the method achieves the technical effect of autonomous movement tracking of the camera.
[0064] Based on Embodiment 3 of this application, Embodiment 4 of this application proposes a control method for a following camera device, which can be referred to the above description and will not be repeated hereafter. Based on this, the following method adjusts the following camera device to its initial pose according to distance information and / or coordinate information, including at least one of steps S21, S22, and S23: Step S21: Based on the difference between the distance information and the preset target distance, control the moving component of the following camera to move the distance corresponding to the difference, so as to adjust the following camera to the initial pose.
[0065] In this embodiment, the preset target distance is the optimal straight-line distance between the tracking device and the target being filmed, which is pre-set through engineering experiments for different shooting scenarios. It is used to determine whether the initial pose of the tracking device is reasonable and to drive the moving component to perform position calibration.
[0066] In one implementation, the tracking camera uses a distance sensing unit, such as a Time-of-Flight (TOF) sensor, to measure the current target distance in real time and compare it with a preset target distance. If the current target distance is greater than the preset target distance, the central controller sends a forward command to the drive unit in the moving component, driving the moving unit to move towards the target. If the current target distance is less than the preset target distance, the central controller sends a backward command to the drive unit in the moving component, driving the moving unit to move away from the target. By using the preset target distance as a reference and converting the real-time distance into movement commands for the moving component, the tracking camera can be quickly calibrated to the initial pose at the optimal shooting distance, improving the initial accuracy and image quality of the tracking shot.
[0067] Step S22: Based on the horizontal deviation between the coordinate information and the coordinates of the center of the field of view, control the rotation unit or moving component of the following camera to rotate the angle corresponding to the horizontal deviation, so as to adjust the following camera to the initial pose.
[0068] In this embodiment, the field-of-view center coordinates are the pixel coordinates corresponding to the pre-set geometric center of the field of view in the image captured by the following camera's shooting component. These coordinates are used to determine whether the target is horizontally centered in the image. This coordinate is a fixed reference value of the shooting component, determined by camera intrinsic parameters such as resolution and optical center position. For example, when the resolution of the image captured by the shooting component is 1920×1080, the field-of-view center coordinates are pixel (960, 540).
[0069] In one implementation, the tracking device acquires the target's position coordinates (x, y) in the image using a camera component, such as a binocular camera, and compares them with the coordinates of the center of the field of view (x0, y0). The horizontal deviation Δx = x between the target coordinates and the center of the field of view is calculated. If |Δx| is greater than a preset horizontal threshold, such as 50 pixels, the central controller controls the rotating unit or moving component of the tracking device to rotate by the angle corresponding to the horizontal deviation until the target is horizontally centered in the field of view. The angle corresponding to the horizontal deviation is calculated by the horizontal deviation Δx and the focal length f of the shooting component. By calculating the horizontal deviation of the target, the rotating unit or moving component is adjusted to complete the angle compensation, ensuring that the target is in the horizontal center of the frame at the initial stage of shooting.
[0070] Step S23: Based on the vertical deviation between the coordinate information and the coordinates of the center of the field of view, control the telescopic unit of the following camera to adjust the distance corresponding to the horizontal deviation, so as to adjust the following camera to the initial pose.
[0071] As one implementation method, the vertical deviation Δy=y between the target coordinate information and the image center coordinates is calculated. If |Δy| is greater than a preset vertical threshold, such as 30 pixels, the central controller controls the telescopic unit to extend or retract, adjusting the distance corresponding to the vertical deviation until the target is vertically centered in the field of view. By calculating the vertical deviation between the target and the center of the field of view, the telescopic unit is controlled to adjust the height of the shooting components, achieving initial pose calibration of the tracking device in the vertical direction, ensuring that the target is always stably in the optimal vertical position for imaging in the shooting frame.
[0072] In this embodiment, pose calibration is completed based on distance difference, horizontal deviation, and vertical deviation, so as to achieve accurate adjustment of the initial pose of the tracking device, lay the best initial imaging foundation for subsequent intelligent tracking and shooting, and improve the initial imaging effect of tracking and the accuracy of subsequent dynamic tracking.
[0073] Based on any of the above embodiments of this application, Embodiment 5 of this application proposes a control method for a following camera device, which can be referred to the above description and will not be repeated hereafter. Based on this, the pose offset is distance information; the pose offset of the target object is obtained cyclically; and the moving components of the following camera device are adjusted according to the pose offset obtained each time, including: Step S31: Collect distance information between the shooting components of the tracking device and the target object according to the first frequency.
[0074] In this embodiment, the first frequency is a preset fixed sampling frequency, which is used to control the periodic working timing of the distance sensing unit.
[0075] In one implementation, the distance sensing unit in the shooting component, such as a TOF area array sensor, transmits a ranging signal to the target object and receives the echo signal under the timing drive of the central controller at a first frequency, such as 10Hz. The central controller analyzes and processes the echo signal to calculate the straight-line physical distance between the shooting component and the target object at each sampling moment, and adds a corresponding timestamp to the distance value to form distance information with a time stamp, thus completing a single distance acquisition and obtaining the distance information D between the shooting component of the tracking device and the target object. t It continuously collects distance information between the imaging component and the target object at the highest frequency, providing historical location data of the target.
[0076] Step S32: Collect the coordinate information of the target object in the field of view of the shooting component according to the second frequency.
[0077] In one implementation, the image acquisition unit in the imaging component, such as a binocular camera, acquires field-of-view images containing the target object at a second frequency, such as 30fps, driven by the central controller. A corresponding timestamp is added to each frame to ensure temporal alignment with distance information. The central controller performs target detection and feature localization processing on the timestamped field-of-view images, identifying the target region of the target object in the image and extracting the center pixel of that region. The horizontal and vertical pixel coordinates of this center pixel are used as the coordinate information of the target object in the field of view of the imaging component. The image acquisition, target detection, and coordinate extraction steps are repeatedly executed at the second frequency to continuously acquire the temporal coordinate information of the target object, thus obtaining the coordinate information (x, y, y) of the target object in the field of view of the imaging component. t ,y t The target's coordinate information within the shooting field of view is acquired at a second frequency, providing a basis for constructing the target's three-dimensional coordinate data.
[0078] Step S33: Construct historical location data of the target object in the device coordinate system based on distance information and coordinate information.
[0079] As one implementation method, based on the camera intrinsic parameters in the shooting component, such as focal length f, optical center coordinates (u0, v0), and distance information D... t The pinhole imaging model is converted into the target's three-dimensional coordinates in the device coordinate system. :
[0080] The calculated historical location data of the target is represented as a time series {(t1,P1),(t2,P2),...,(t... n ,P n )},in Let be the three-dimensional position of the target at time t. Let be the three-dimensional coordinates of the target in the device coordinate system at time t. Based on distance and coordinate information, historical position data of the target in the device coordinate system is constructed to record the target's movement trajectory, laying a data foundation for subsequent analysis of the target's movement characteristics.
[0081] Step S34: Determine the movement speed corresponding to the historical location data.
[0082] As one implementation method, data preprocessing is performed based on the historical location information of the target object. The data collected by the distance sensing unit and the image acquisition unit are uniformly resampled to a preset time interval, such as 20Hz. Missing points are filled in using linear interpolation to obtain a uniform and continuous time series. Sliding median filtering is used to eliminate outliers. A window size is set, for example, a window size of 5, which covers 0.25s of data. After filtering, the position P... t ' is the median of the five original positions within the window, i.e. . This represents the three-dimensional position of the target in the device coordinate system at time t. Let `median` be the 3D position of the target in the device coordinate system at time `t` after preprocessing, and `median` be the median value function, which sorts the five consecutive data points within the window to obtain the median value. Based on the preprocessed... The spatial displacement is calculated, and the moving speed at time t is obtained by combining it with the time interval. A continuous time series is obtained through preprocessing historical location data, which in turn determines the target's moving speed, accurately quantifies the target's dynamic characteristics, and provides a basis for subsequent prediction of the target's path.
[0083] As another implementation, data preprocessing is performed based on the historical location information of the target object to establish a unified time series for data from distance sensing units such as TOF sensors (10Hz) and image acquisition units such as binocular cameras (30fps). For example, when the high-frequency frequency is an integer multiple of the low-frequency frequency (i.e., the number of data points acquired by the image acquisition unit at 30fps in the same time period is an integer multiple of the number of data points acquired by the distance sensing unit at 10Hz), the higher-frequency data stream is selected as the time reference. For example, since the 30fps frequency of the image acquisition unit is higher, its timestamp sequence is used as the target alignment timeline. The preset time interval is the sampling interval of this high-frequency data stream. After resampling, the image coordinate data directly uses the value of its original sampling time. Each target time point corresponds to a coordinate information. For lower-frequency distance data, it is aligned to the high-frequency timestamp by repeating the previous value. For each target time point, the last distance sampling point with a timestamp no later than that time point is found in the distance data sequence, and the distance value of that distance sampling point is assigned to the target time point as the distance value at that time. The spatial displacement is calculated based on the preprocessed coordinate information, and the movement speed at time t is obtained by combining the coordinates with the time interval. By determining a high-frequency time reference, retaining high-frequency data, and aligning low-frequency data to maintain the previous value, historical position information can be processed quickly to obtain the movement speed, reducing computational overhead.
[0084] Step S35: Obtain the predicted movement path of the target object based on the fitting model corresponding to the movement speed.
[0085] As one implementation method, a corresponding fitting model is selected based on the movement speed. For example, when the movement speed is low, a multinomial fitting model is selected for prediction; when the movement speed is high, a Kalman filter model is first used for prediction to obtain the predicted movement path of the target object. By matching the corresponding fitting model to the target's movement speed and generating the predicted movement path, accurate prediction of the target's subsequent movement trajectory can be achieved.
[0086] Step S36: Adjust the moving components of the following camera based on the predicted movement path.
[0087] As one implementation method, after obtaining the predicted movement path of the target object, it is combined with real-time mapping by LiDAR, and then... An obstacle avoidance path is generated using an algorithm (A-Star Algorithm) or DWA (Dynamic Window Approach) to ensure the tracking device maintains a preset target distance. The central controller converts the obstacle avoidance path of the target object into the desired tracking trajectory of the tracking device. Based on the deviation between the current positioning information of the tracking device and the desired tracking trajectory, it generates movement control commands and sends them to the drive module of the moving component. The drive module controls the direction, speed, and start / stop actions of the moving component, enabling the tracking device to move along the desired tracking trajectory. During movement, the actual pose information of the moving component is fed back in real time, and the height is adjusted simultaneously through the viewpoint adjustment component to keep the target vertically centered in the field of view, forming a closed-loop control. The tracking trajectory is continuously corrected until the tracking device and the target object maintain a preset tracking pose relationship, completing the dynamic tracking adjustment based on the predicted movement path. By adjusting the moving component of the tracking device in advance based on the target's predicted movement path, the tracking device can predictively track the target, significantly improving the real-time performance and accuracy of the tracking device for dynamic targets.
[0088] In this embodiment, by collecting target distance and coordinate information at multiple frequencies, constructing three-dimensional historical position data and calculating movement speed, adapting to the corresponding fitting model, and then predicting the target movement path and adjusting the movement components in advance, dynamic tracking of the target object is achieved, which greatly improves the real-time performance, smoothness and accuracy of the tracking device and effectively avoids tracking lag problems.
[0089] Based on any of the above embodiments of this application, Embodiment Six of this application proposes a control method for a tracking camera, which can be referred to the above description and will not be repeated hereafter. Based on this, the predicted movement path of the target object is obtained according to the fitting model corresponding to the movement speed, including: Step S351: If the moving speed is less than or equal to the speed threshold, solve the k-th degree polynomial curve constructed from the historical position data using the least squares method.
[0090] As one implementation method, if the moving speed is less than or equal to a speed threshold, i.e., for a low-speed moving target, a polynomial fitting model is selected for path prediction. Assuming the path is a k-th degree polynomial curve, the coefficients are solved using the least squares method. The target's movement path in the two-dimensional XY plane is represented as follows:
[0091] a k ,a k-1 ...a1,a0 are the coefficients of the polynomial to be solved, X and Y are the coordinates of the target position, and the goal is to minimize the sum of squared fitting errors J:
[0092] n is the number of historical data points used for fitting, ( () represents the i-th observation point, and the values within the parentheses represent the values of the i-th observation point. The difference between the observed value and the predicted value from the polynomial model is J, which is the sum of squares of the errors at all observation points. The coefficient vector is obtained by matrix inversion. :
[0093] Where X is the design matrix and Y is the observation location vector, the optimal order k is determined by AIC (Akaike Information Criterion) to avoid overfitting:
[0094] Where n is the number of data points, J is the sum of squared fitting errors under the current k-order model, k is the current model order, and 2(k+1) is a penalty term. The more complex the model, i.e., the larger k is, the larger the penalty term value. The smaller the AIC value, the better the overall performance of the model. The optimal order k that minimizes AIC is selected as the final model order. In the scenario of low-speed target movement, the least squares method is used to perform k-th order polynomial curve fitting on historical position data, which can achieve smooth and accurate trajectory modeling with low computational cost, providing a reliable foundation for path prediction.
[0095] Step S352: Obtain the predicted movement path of the target object based on the solution results.
[0096] As one implementation method, the polynomial expression is determined based on the obtained optimal order k and polynomial coefficients, thus obtaining the predicted movement path of the target object. T represents the prediction duration, such as 1 second. Based on the polynomial curve fitting results, a predicted movement path is generated, which can accurately predict the subsequent movement trend of low-speed targets.
[0097] In this embodiment, by determining the moving speed and using the least squares method for polynomial fitting in low-speed scenarios, efficient and accurate prediction of the target moving path is achieved, taking into account both real-time computation and tracking accuracy, and adapting to the needs of stable tracking shooting scenarios.
[0098] Based on any of the above embodiments of this application, Embodiment Seven of this application proposes a control method for a tracking camera, which can be referred to the above description and will not be repeated hereafter. Based on this, the predicted movement path of the target object is obtained according to the fitting model corresponding to the movement speed, including: Step S353: If the moving speed is greater than the speed threshold, construct the state vector of the target object based on historical location data.
[0099] As one implementation method, if the moving speed is greater than a speed threshold, i.e., for a high-speed moving target, a Kalman filter model is used to predict the next position. The state vector includes position, velocity, and acceleration. ,in Let be the two-dimensional position of the target at time t. Let be the two-dimensional velocity of the target at time t. Let be the two-dimensional acceleration of the target at time t. A state vector is constructed based on the historical position data of the high-speed moving target to represent its real-time motion state, providing standardized input for subsequent Kalman filter prediction.
[0100] Step S354: Construct the state equation based on the state vector, the uniform acceleration transfer matrix, and the first covariance matrix.
[0101] As one implementation method, the state equation is:
[0102] It describes the process of the target state evolving over time. Let be the state vector at time t. The process noise, i.e., the first covariance matrix Q, and the uniform acceleration transfer matrix F describe the uniform acceleration motion model:
[0103] in The uniform acceleration transfer matrix F describes the X and Y directions respectively, taking the X direction (first three rows and first three columns) as an example. The first row indicates that the new position = original position + original velocity × time + 0.5 × original acceleration × time²; the second row indicates that the new velocity = original velocity + original acceleration × time; and the third row indicates that the new acceleration = original acceleration. The same logic applies to the Y direction. By constructing the state equation using the state vector, the uniform acceleration transfer matrix, and the first covariance matrix, a mathematical model conforming to the laws of high-speed motion is established, improving the stability of trajectory prediction.
[0104] Step S355: Obtain the predicted movement path of the target object based on the state equation and historical location data.
[0105] As one implementation method, based on the state equation and historical location data, the target's state at time t is determined. Using the uniform acceleration transfer matrix F describing the motion law, the predicted state of the target object at time t+1 is predicted, thus obtaining the predicted movement path of the target object. Based on the state equation and historical position data, high-speed target trajectory prediction is achieved, outputting a high-precision predicted movement path, providing a reliable basis for dynamic tracking.
[0106] Specifically, after obtaining the predicted movement path of the target object based on the state equation and historical location data, the steps include: The observation equation is constructed based on the observation matrix and the second covariance matrix.
[0107] The predicted movement path of the target object is obtained based on the observation equation, the state equation, and historical location data.
[0108] As one implementation method, observation vector That is, the actual target position measured by the sensor, the observation matrix. , It is a 2x2 identity matrix, representing only the observed position. This indicates that velocity and acceleration are not directly observed. , i.e., the second covariance matrix R, represents the sensor error. The observation equation is constructed based on the observation matrix and the second covariance matrix, and is:
[0109] Prediction is based on the state equation, using the optimal estimate from the previous time step to predict the current state and uncertainties: ,
[0110] Obtain the predicted state at the current time without observation correction. The covariance of the current state estimation error The Kalman gain is calculated based on the observation equation and the current observation vector. Kalman gain balances the weights of model predictions and current observations:
[0111] Using the current observation vector By correcting the predicted state, we obtain the optimal posterior state estimate for the current time step and the uncertainty of the estimate: ,
[0112] Finally, the optimal state estimate is obtained. ,in This represents the difference between the observed value and the predicted observed value. It represents the current optimal state estimate. Starting from this point, the state equation is substituted back into the equation for forward recursion to obtain the predicted movement path. That is, the position component is extracted from the predicted state sequence to obtain the predicted movement path of the target object over a period of time in the future. This yields the corrected and updated predicted movement path. By constructing observation equations and combining them with state equations and historical location data, the prediction and correction of the target's motion state are achieved, improving the accuracy of the predicted movement path.
[0113] Specifically, after obtaining the predicted movement path of the target object, the process also includes: The mean square error is determined based on the predicted movement path and the actual movement path.
[0114] If the mean square error is greater than the error threshold, update the first covariance matrix and the second covariance matrix.
[0115] As one implementation method, the mean square error between the predicted position and the actual observed position is calculated:
[0116] in This is the target location prediction sequence output from the fitted model, i.e., the polynomial fitted model or the Kalman filter model. The target position data is derived from actual observations by the imaging component and preprocessed. MSE (Mean Squared Error) is the average of the squared distances between the predicted and observed positions; a smaller MSE indicates a more accurate prediction. If the MSE exceeds an error threshold (e.g., 0.2m), model adjustment is triggered. When the target velocity is less than or equal to the velocity threshold and the MSE is less than or equal to the error threshold, a polynomial fitting path is used. When the target velocity is greater than the velocity threshold or the MSE is greater than the error threshold, Kalman filtering prediction is switched, and the first covariance matrix Q and the second covariance matrix R are dynamically updated. Q represents the degree to which the actual target motion deviates from the preset model. If the MSE is too large, it indicates that the actual target motion is stronger and more uncertain than the model predicts, thus increasing the value of Q. After Q increases, the model prediction is considered less reliable in the next calculation, thus assigning higher weights to new observation data, allowing the state estimation to more quickly match the target's true motion. R represents the noise or unreliability of the device measurements. If the MSE is too large, and the analysis suggests increased data fluctuation, it is inferred that the observation noise has increased, thus increasing the value of R. Increasing R in the next calculation will lead to the assumption that the current observation data is relatively noisy, thus appropriately reducing the confidence in the current observation and relying more on the model's prediction, thereby smoothing and stabilizing the estimation. By calculating the mean square error between the predicted path and the actual path, the model is dynamically and adaptively switched and the covariance matrix is updated to improve the accuracy and environmental adaptability of target tracking. By continuously making the set values of Q and R approximate the true noise statistics, the gain of the Kalman filter is ensured. It is at or near the optimal value, thus continuously outputting a high-quality estimate of the target state.
[0117] In this embodiment, a state equation-based Kalman filter model is used for high-speed moving scenarios, which can accurately characterize the high-speed motion characteristics of the target, effectively suppress noise and sudden interference, and achieve reliable path prediction under complex dynamic conditions.
[0118] Based on any of the above embodiments of this application, Embodiment Eight of this application proposes a control method for a tracking camera, which can be referred to the above description and will not be repeated hereafter. Based on this, the pose offset is the coordinate deviation. The pose offset of the target object is obtained iteratively, and the shooting components of the tracking camera are adjusted according to the pose offset obtained each time, including: Step S37: Collect the coordinate information of the target object in the field of view of the shooting component according to the second frequency, and adjust at least one of the telescopic unit, rotation unit, and moving component according to the coordinate deviation between the coordinate information obtained each time and the coordinate of the center of the field of view, so as to adjust the angle of view of the shooting component.
[0119] As one implementation method, the coordinate information of the target object in the field of view of the shooting component is acquired cyclically at a second frequency. The difference between each acquired coordinate information and the coordinates of the center of the field of view is calculated to obtain the coordinate deviation. Based on the coordinate deviation, at least one of the telescopic unit, rotation unit, and moving component is selected and driven to perform an adjustment action, thereby adjusting the shooting angle of the shooting component to keep the target object always in the center of the field of view. By acquiring the coordinate deviation at high frequency and adjusting the shooting angle in real time, dynamic tracking of the target object is achieved, effectively eliminating pose offset and ensuring the stability of the captured image and the centering effect of the target.
[0120] As another implementation method, the telescopic unit, rotating unit, and moving component are controlled collaboratively. The central controller converts the acquired coordinate deviations into a comprehensive adjustment value including horizontal angle, vertical height, and device pose. This comprehensive adjustment value is then synchronously distributed to the telescopic unit, rotating unit, and moving component, enabling multiple actuators to move collaboratively according to their linkage relationship. Through the coupled adjustment of multiple actuators, coordinate deviations are eliminated, the target object is kept centered in the field of view, and the adjustment load and motion trajectory of each actuator are optimized. This avoids tracking jitter or overshoot caused by frequent actions of a single component, improving the smoothness, stability, and overall control accuracy of the dynamic tracking process.
[0121] In this embodiment, the deviation between the target coordinates and the center of the field of view is calculated in real time, and the telescopic, rotating or moving components are driven to make adjustments to ensure that the target object is always in the center of the screen, thus achieving accurate and stable dynamic tracking.
[0122] Based on any of the above embodiments of this application, Embodiment Nine of this application proposes a control method for a tracking camera, which can be referred to the above description and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3This is a flowchart illustrating Embodiment Nine of the control method for the tracking camera of this application. Before determining the distance information between the shooting component of the tracking camera and the target object and / or the coordinate information of the target object within the field of view of the shooting component in response to a target object, the method includes at least one of steps S101 and S102: Step S101: In response to the gesture information recognized by the shooting component, determine the target object pointed to by the gesture information.
[0123] As one implementation method, the user makes a preset gesture within the field of view of the shooting component, such as spreading and waving their hand. The tracking device uses image recognition algorithms, such as YOLO object detection and gesture classification models, to recognize the gesture and lock onto the target at the corresponding position, such as a person or vehicle—that is, the target object pointed to by the gesture information. By recognizing gesture information and locking onto the target object through the shooting component, rapid target selection is achieved, which is convenient, responsive, and improves target locking efficiency.
[0124] Specifically, the real-time image is processed, and all hand regions in the image are located using image recognition algorithms. These hand regions are then classified to determine if they correspond to a pre-defined trigger gesture. Hand keypoint detection is performed; for the identified hand, a pose estimation model is used to extract the pixel coordinates of multiple keypoints, such as the index fingertip, the base of the index finger, and the palm. Based on these keypoints, the geometric information of the pointing direction is calculated. The coordinates of the index fingertip are used as the most direct pointing point, and the vector from the base of the index finger to the fingertip is calculated as the pointing direction. A pointing geometric confidence score is calculated. For example, the hand keypoint detection itself has a confidence score, which is averaged; or the pointing intention is determined by the degree of finger extension, i.e., by calculating the angle between keypoints. The higher the extension, the higher the pointing geometric confidence score. Candidate targets that the gesture might point to are searched in the image, and the association confidence score for each candidate target is calculated. An object detection model is used to identify all potential targets in the image, obtaining their bounding boxes. For each candidate target, its association score with the gesture pointing direction is calculated; this score is used as the association confidence score for that target. The association confidence score is calculated based on whether the pointing point falls within the target bounding box, the consistency of the direction of the vector from the pointing point to the center or nearest edge of the target box with the pointing direction vector, and the Euclidean distance between the pointing point and the target box. Higher association confidence scores are associated with the pointing point falling within the target bounding box, higher direction consistency, and smaller Euclidean distance. A comprehensive confidence score is calculated by weighting the association confidence score and the pointing geometric confidence score. Based on the calculated comprehensive confidence scores of each candidate target, if the comprehensive confidence score is greater than a confidence threshold, the gesture is identified as pointing to the target with the highest confidence score. If the highest comprehensive confidence score is greater than the confidence threshold but very close to the confidence scores of one or more other targets, the pointing is considered ambiguous. In this case, the gesture may not be locked, and the user may be asked to make a clearer gesture or select the target through other means, or interactive inquiry may be initiated. If all comprehensive confidence scores are below the confidence threshold, a valid pointing target is not identified, and the gesture locking fails. Based on the comprehensive confidence level, the system accurately determines the pointing position of the gesture and makes a decision based on the comparison of thresholds. This system can handle ambiguous or complex pointing scenarios more intelligently, effectively reduce false locking, and improve the accuracy and reliability of human-computer interaction.
[0125] Step S102: In response to the received positioning signal, determine the target object to which the positioning signal points.
[0126] In one implementation, the user sends a target location signal via a dedicated positioning remote control, such as Bluetooth or 2.4G wireless signal. The device receives the signal and combines it with the coordinates from the positioning remote control, which are obtained through a built-in GPS (Global Positioning System) or UWB (Ultra Wide Band) positioning module, to lock onto nearby objects as the target—the object the positioning signal points to. By locking onto the target object upon receiving the positioning signal, target selection can be achieved at long distances and outside the field of view, expanding the applicable scenarios and operational flexibility of target locking for tracking devices.
[0127] Specifically, the user holds a dedicated positioning remote control and presses a specific trigger button on the remote control. The dedicated positioning remote control then transmits a target positioning signal containing its identification identifier through its wireless module.
[0128] The tracking device continuously monitors the wireless channel. When its wireless receiving module captures and identifies a valid positioning signal from the paired remote controller, it triggers the target locking process. Through its built-in GPS module, it calculates its own geographic coordinates in real time and encodes these coordinates in the signal or transmits them via other links; alternatively, through its built-in UWB positioning module, in areas where UWB positioning base stations are deployed, the dedicated positioning remote controller acts as a UWB tag, calculating its own high-precision 3D coordinates in real time. Using the received coordinates from the dedicated positioning remote controller as the center, the tracking device defines a spherical or cylindrical search area with a preset radius in 3D space. Within this search area, it identifies and locks onto the target. The tracking device drives its environmental perception unit to perform a focused scan of the defined search area, analyzing the data within the area and identifying the objects within it—the target object indicated by the positioning signal. By operating the remote controller near the target, the tracking object can be quickly and accurately designated, expanding its applicability and operational efficiency in complex and large-scale scenarios.
[0129] In this embodiment, target locking is achieved through gesture recognition or positioning signals, enabling rapid and accurate selection of target objects across multiple scenarios, thereby improving the device's scene adaptability and ease of operation.
[0130] Based on any of the above embodiments of this application, Embodiment Ten of this application proposes a control method for a tracking camera, which can be referred to the above description and will not be repeated hereafter. During the target's movement, there is a possibility that the target may suddenly turn, resulting in the inability to collect distance and coordinate information.
[0131] As one implementation method, the weights of data fusion are dynamically adjusted when data anomalies or loss occur.
[0132] Specifically, initial confidence levels are assigned to distance information, coordinate information, and predicted state, where distance information... Coordinate information derived from distance sensing unit The two-dimensional position observations are derived from the image acquisition unit after coordinate transformation, and the predicted state is derived from the state equation prediction of the Kalman filter model. A real-time confidence score between 0 and 1 is calculated for each data source; a higher score indicates greater reliability of the data at the current moment. For the observation data containing distance and coordinate information, innovation is calculated based on the observation matrix corresponding to the distance and coordinates. Innovation is the difference between the observed value and the model's predicted value, used to characterize the "new information" brought by the observation data, and is also the core basis for judging observation anomalies, calculating Mahalanobis distance, and adaptively adjusting filter weights—that is, the difference between the observed value and the predicted value. The Mahalanobis distance of the innovation is calculated and compared with a threshold based on the chi-square distribution to determine whether the observation is significantly abnormal. The status flags returned by each unit are checked; if an invalid or erroneous status is returned, the corresponding confidence is set to zero. Confidence is calculated based on the status flags (1 for valid, 0 for invalid), the decay coefficient (which controls the rate at which confidence decreases with the degree of anomaly), and the Mahalanobis distance threshold of the observation innovation. For the predicted state, the uncertainty of the predicted state is determined by the prediction error covariance matrix. The model describes a method for measuring prediction uncertainty by taking the trace or the largest eigenvalue. It calculates statistics on recent prediction errors, such as the mean squared error (MSE) within a sliding window, to evaluate the model's long-term performance. Confidence is calculated based on prediction uncertainty and MSE; higher prediction uncertainty and recent prediction errors result in lower model confidence. For each observation data source, the observation noise covariance matrix R (the second covariance matrix) is adjusted according to its real-time confidence, thereby reducing the weight of the unreliable observation in state updates. When observation confidence is generally low while prediction confidence is relatively high, the process noise covariance matrix Q (the first covariance matrix) can be appropriately reduced; conversely, if observation confidence is high but prediction error is large, Q can be appropriately increased to allow the filter to respond more quickly to observation changes. By adjusting Kalman filter parameters through real-time confidence assessment, adaptive weight allocation is achieved. This enables the tracking device to automatically reduce or eliminate the impact of unreliable data in complex situations such as sudden target changes, environmental interference causing data anomalies or loss, and enhance the role of the remaining reliable data, thus significantly improving the adaptability and fault tolerance of the tracking device.
[0133] Furthermore, to achieve faster and more accurate tracking of target objects, wireless positioning tags or beacons can be installed on them. These devices continuously or on demand emit wireless signals containing their unique IDs. The tracking device receives these signals and can directly identify and locate specific target objects. While tracking may be lost when the target moves rapidly, suddenly changes direction, enters a blind spot, or is briefly obscured by other objects, wireless signals possess diffraction and penetration capabilities, maintaining a connection under non-extreme conditions. This allows for continued tracking of the target even during periods of visual interruption, relying on wireless positioning signals to achieve continuous tracking coverage and significantly reduce the probability of losing the target.
[0134] This application provides a control device for a tracking camera, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the control method of the tracking camera in the first embodiment described above.
[0135] The following is for reference. Figure 4 The diagram illustrates a structural schematic of a control device suitable for implementing the tracking camera device in the embodiments of this application. The control device for the tracking camera device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablets, and in-vehicle terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The control device of the tracking camera shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0136] like Figure 4As shown, the control device of the tracking camera may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the control device of the tracking camera. The processing unit 1001, the read-only memory 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the control device of the tracking equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the control device of the tracking equipment with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0137] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0138] The control device for the tracking camera provided in this application, employing the control method for the tracking camera in the above embodiments, can solve the technical problem of the lack of motion tracking capability in the prior art. Compared with the prior art, the beneficial effects of the control device for the tracking camera provided in this application are the same as those of the control device for the tracking camera provided in the above embodiments, and other technical features in the control device for the tracking camera are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0139] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0140] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0141] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the control method of the tracking device in the above embodiments.
[0142] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), or any suitable combination thereof.
[0143] The aforementioned computer-readable storage medium may be included in the control device of the tracking camera; or it may exist independently and not be assembled into the control device of the tracking camera.
[0144] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the control device of the tracking camera, cause the control device to: determine, in response to a target object, the distance information between the shooting component of the tracking camera and the target object and / or the coordinate information of the target object in the field of view of the shooting component; adjust the tracking camera to an initial pose based on the distance information and / or the coordinate information; cyclically acquire the pose offset of the target object, and adjust the moving component of the tracking camera and / or adjust the shooting component of the tracking camera based on each acquired pose offset.
[0145] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation that may be implemented in systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0147] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0148] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the control method of the above-described tracking device, thereby solving the technical problem of the lack of motion tracking capability in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the control method of the tracking device provided in the above embodiments, and will not be repeated here.
[0149] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the control method for the tracking device described above.
[0150] The computer program product provided in this application can solve the technical problem of the lack of motion tracking capability in the prior art. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the control method of the tracking device provided in the above embodiments, and will not be repeated here.
[0151] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A tracking camera device, characterized in that, The following shooting device includes a moving component, a perspective adjustment component, and a shooting component. The moving component is provided with a moving unit and a driving unit. The first end of the perspective adjustment component is connected to the moving component, and the second end is connected to the shooting component. The perspective adjustment component includes a telescopic unit so as to adjust the relative distance between the shooting component and the moving component.
2. The following camera device as described in claim 1, characterized in that, The second end of the viewing angle adjustment component is provided with a rotation unit so that the relative angle between the shooting component and the viewing angle adjustment component can be adjusted by the rotation unit.
3. A control method for a tracking camera, characterized in that, The control method for the following filming device includes: In response to a target object, the distance information between the shooting component of the tracking device and the target object and / or the coordinate information of the target object in the field of view of the shooting component are determined; The following camera device is adjusted to its initial pose based on the distance information and / or the coordinate information. The pose offset of the target object is obtained cyclically, and the moving component and / or shooting component of the following camera are adjusted according to the pose offset obtained each time.
4. The control method for the tracking camera as described in claim 3, characterized in that, Adjusting the following camera to its initial pose based on the distance information and / or the coordinate information includes: Based on the difference between the distance information and the preset target distance, the moving component of the following camera is controlled to move the distance corresponding to the difference, so as to adjust the following camera to the initial pose; and / or, Based on the horizontal deviation between the coordinate information and the coordinates of the field of view center, the rotation unit or moving component of the tracking device is controlled to rotate by the angle corresponding to the horizontal deviation, so as to adjust the tracking device to the initial pose; and / or, Based on the vertical deviation between the coordinate information and the coordinates of the center of vision, the telescopic unit of the following camera is controlled to adjust the distance corresponding to the horizontal deviation, so as to adjust the following camera to the initial pose.
5. The control method for the tracking camera as described in claim 3, characterized in that, The pose offset is distance information. The pose offset of the target object is obtained cyclically. Based on the pose offset obtained each time, the movement components of the tracking device are adjusted, including: The distance information between the shooting component of the tracking device and the target object is collected according to a first frequency; The coordinate information of the target object in the field of view of the shooting component is acquired according to the second frequency; Based on the distance information and the coordinate information, construct the historical position data of the target object in the device coordinate system; Determine the movement speed corresponding to the historical location data; The predicted movement path of the target object is obtained based on the fitting model corresponding to the movement speed; The movement components of the tracking device are adjusted based on the predicted movement path.
6. The control method for the tracking camera as described in claim 5, characterized in that, The step of obtaining the predicted movement path of the target object based on the fitting model corresponding to the movement speed includes: If the moving speed is less than or equal to the speed threshold, the k-th degree polynomial curve constructed from the historical position data is solved using the least squares method. The predicted movement path of the target object is obtained based on the solution results.
7. The control method for the tracking camera as described in claim 5, characterized in that, The step of obtaining the predicted movement path of the target object based on the fitting model corresponding to the movement speed includes: If the moving speed is greater than the speed threshold, a state vector of the target object is constructed based on the historical location data; The state equation is constructed based on the state vector, the uniform acceleration transfer matrix, and the first covariance matrix. The predicted movement path of the target object is obtained based on the state equation and the historical location data.
8. The control method for the tracking camera as described in claim 3, characterized in that, The pose offset is a coordinate deviation. The pose offset of the target object is obtained iteratively. Based on the pose offset obtained each time, the shooting components of the tracking device are adjusted, including: The target object's coordinate information in the field of view of the shooting component is collected according to the second frequency. Based on the coordinate deviation between the coordinate information collected each time and the coordinates of the center of the field of view, at least one of the telescopic unit, the rotation unit, and the moving component is adjusted to adjust the angle of view of the shooting component.
9. The control method for the tracking camera as described in claim 3, characterized in that, Before determining the distance information between the shooting component of the tracking device and the target object and / or the coordinate information of the target object in the field of view of the shooting component in response to the target object, the following steps are included: In response to the gesture information recognized by the shooting component, the target object to which the gesture information points is determined; and / or, In response to a received positioning signal, the target object to which the positioning signal points is determined.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the control method for the tracking device as described in any one of claims 3 to 9.
Citation Information
Patent Citations
System and method of recognizing, positioning and tracking ping-pong trajectory
CN106780620A
Lens control apparatus and control method for tracking moving object
CN110022432A
Automatic following audio and video acquisition system and method
CN111182221A
Underwater low-speed small target tracking method and system under complex conditions
CN117745760A
Multi-angle tracking type monitoring device
CN221196867U