Intelligent tracking shooting method, shooting holder, equipment and storage medium
Through AI cameras and multi-objective tracking algorithms combined with facial and human body feature extraction, the problems of accurate and high hardware cost of athletes' individual recognition in dynamic scenes are solved, and low-cost and efficient intelligent tracking and shooting are achieved.
Patent Information
- Application Number
- CN202510419862.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
The existing video tracking and shooting methods cannot achieve independent shooting decisions in dynamic scenes, it is difficult to accurately distinguish individual athletes from specific teams, and the hardware storage cost is high, making it difficult to meet the real-time needs of live events.
AI cameras are used to combine object detection models and multi-object tracking algorithms to achieve intelligent tracking and shooting through distortion correction and physical coordinate conversion, and to perform secondary confirmation by combining face and human body feature extraction, reducing hardware costs and improving tracking accuracy.
It realizes efficient and intelligent tracking and shooting in dynamic scenarios, reduces costs, and improves the accuracy of athletes' individual recognition and the real-time broadcast of events.
Smart Images

Figure CN120302158A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of image recognition technology, and in particular, to an intelligent tracking shooting method, a shooting pan-tilt, a device, and a storage medium. Background Art
[0002] Tracking shooting technology is an important means for video acquisition in sports events and is widely used in professional scenarios such as athlete motion capture and event live broadcast. However, there are still significant technical bottlenecks in the intelligent tracking and multi-angle acquisition of dynamic scenes in currently commonly used video tracking shooting methods. First, traditional tracking shooting pan-tilt systems mostly use fixed-angle shooting or manual control shooting, and cannot make autonomous shooting decisions in dynamic scenes, easily resulting in the lack of multi-dimensional image information. Second, tracking algorithms generally adopt the principle of maximum target priority or feature recognition based on the color of human clothing, and can only track the overall crowd in the field, unable to accurately distinguish individual athletes of a specific team, lacking facial feature extraction and multi-modal verification mechanisms, and are extremely prone to target misjudgment or loss in scenarios where multiple people have similar clothing, resulting in insufficient dynamic tracking accuracy. In addition, existing automatic shooting methods for sports events generally adopt panoramic stitching technology, constructing a panoramic image through multi-wide-angle lens distortion correction and image stitching, with a large hardware storage cost and difficult to meet the real-time requirements of event live broadcast.
[0003] Therefore, there is an urgent need for a real-time intelligent tracking shooting method to meet the intelligent, low-cost, and high-efficiency shooting requirements of sports events. Summary of the Invention
[0004] According to the embodiments of the present application, an intelligent tracking shooting solution is provided, which can achieve efficient intelligent tracking shooting and greatly reduce the cost of sports event shooting.
[0005] In the first aspect of the present application, an intelligent tracking shooting method is provided. The method includes:
[0006] Obtain a first frame containing key targets captured by an AI camera in a shooting pan-tilt, input the first frame into a target detection model, and obtain the positions of the key targets;
[0007] Convert the positions of the key targets into first physical coordinates according to the distortion parameters and rotation angles of the AI camera;
[0008] Intelligently track and shoot the key targets according to the first physical coordinates and a multi-target tracking algorithm.
[0009] Optionally, the key targets include but are not limited to people and balls.
[0010] In a possible implementation, in addition to intelligently tracking and photographing key targets, it is also necessary to continuously obtain the moving images of key targets and analyze the regions of interest in the images, and send rotation commands to the shooting pan-tilt according to the movement of the regions of interest in the images.
[0011] Optionally, the method for selecting the region of interest in the image includes but is not limited to user-defined selection and automatic algorithm selection. In user-defined selection, the user can manually smear or outline the range of the region of interest in the interactive interface.
[0012] In a possible implementation, intelligently tracking and photographing key targets according to the first physical coordinates and multi-object tracking algorithm includes:
[0013] The multi-object tracking algorithm predicts the movement of the key target according to the first physical coordinates to obtain the second physical coordinates of the key target;
[0014] Automatically adjust the pitch angle and yaw angle of the shooting pan-tilt according to the second physical coordinates, and continuously obtain the second image containing the key target.
[0015] Optionally, the method further includes:
[0016] Extract the first face feature data by performing face feature extraction on the first image through the face recognition and face bone point detection model and the face feature extraction model;
[0017] Extract the first human body feature data by performing human body feature extraction on the first image through the human body feature extraction model;
[0018] Input the second image into the human body feature extraction model in sequence for feature extraction to obtain the second human body feature data, screen the third human body feature data similar to the first human body feature data from the second human body feature data, and perform intelligent target tracking and photographing on the third image corresponding to the third human body feature data;
[0019] Extract the second face feature data by performing face feature extraction on the third image through the face recognition and face bone point detection model and the face feature extraction model, and perform secondary confirmation on the third image according to the first face feature data and the second face feature data.
[0020] Optionally, extracting the first face feature data by performing face feature extraction on the first image through the face recognition and face bone point detection model and the face feature extraction model includes:
[0021] Perform face positioning and alignment on the first image through the face recognition and face bone point detection model to obtain the first face data and the first face bone point data;
[0022] Input the first face data and the first face bone point data into the face feature extraction model to obtain the first face feature data.
[0023] Optionally, perform a secondary confirmation on the third screen according to the first face feature data and the second face feature data, including:
[0024] Perform a similarity match between the first face feature data and the second face feature data;
[0025] If the first face feature data and the second face feature data are similar, continue with intelligent target tracking shooting;
[0026] If the first face feature data and the second face feature data are not similar, re-screen the third screen.
[0027] Optionally, the method further includes:
[0028] When performing intelligent target tracking shooting on a key target, randomly extract the third screen, input the third screen into the human body feature extraction model for feature extraction, and obtain the fourth human body feature data;
[0029] Perform a similarity match between the fourth human body feature data and the target feature queue;
[0030] If the fourth human body feature data and the target feature queue are similar, continue with intelligent target tracking shooting;
[0031] If the fourth human body feature data and the target feature queue are not similar, the intelligent target tracking shooting fails, pause the intelligent target tracking shooting, and re-select the third screen;
[0032] The target feature queue is a feature sequence set of all the third human body feature data collected historically.
[0033] In a possible implementation manner, the shooting pan-tilt is communicatively connected to the mobile terminal, and the method further includes:
[0034] Obtain the operation instruction of the user on the mobile terminal;
[0035] Adjust the pitch angle and yaw angle of the shooting pan-tilt according to the operation instruction, and zoom in on the screen selected in the operation instruction on the mobile terminal.
[0036] In a possible implementation manner, the shooting pan-tilt further includes a display screen, and the display screen can directly display the execution result captured after the shooting pan-tilt receives the drive instruction.
[0037] In a possible implementation, the AI camera captures the specified actions of the user as the start of the operation instruction, and the AI camera recognizes that the user has completed the action instruction on the display screen within the specified time before starting to recognize the user's operation instruction.
[0038] In a second aspect of the present application, an intelligent tracking and shooting pan-tilt is provided, including:
[0039] An AI camera, configured to obtain a first picture containing a key target, input the first picture into a target detection model to obtain the position of the key target; convert the position of the key target into a first physical coordinate according to the distortion parameter and rotation angle of the AI camera; perform intelligent tracking on the key target according to the first physical coordinate and the multi-target tracking algorithm; and send a rotation instruction to the driving mechanism.
[0040] A clamping structure, configured to fix the intelligent terminal.
[0041] A driving mechanism, configured to receive the rotation instruction and drive the shooting pan-tilt to rotate.
[0042] In a third aspect of the present application, an electronic device is provided. The electronic device includes: a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, the method as described above is implemented.
[0043] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present application is implemented.
[0044] The intelligent tracking and shooting method provided by the embodiments of the present application obtains a first picture containing a key target captured by the AI camera in the shooting pan-tilt, inputs the first picture into a target detection model to obtain the position of the key target. Then, the position of the key target is converted into a first physical coordinate according to the distortion parameter and rotation angle of the AI camera, and intelligent tracking shooting is performed on the key target according to the first physical coordinate and the multi-target tracking algorithm, realizing efficient intelligent tracking shooting, improving the efficiency of event shooting, and reducing the shooting cost at the same time.
[0045] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In combination with the drawings and referring to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present application will become more obvious. In the drawings, the same or similar reference numerals represent the same or similar elements, where:
[0047] Figure 1 Flow chart of the intelligent tracking shooting method according to an embodiment of the present application;
[0048] Figure 2 Flow chart of the confirmation of the key target for intelligent tracking shooting according to an embodiment of the present application;
[0049] Figure 3 Schematic structural diagram of the interaction between an intelligent terminal and a shooting gimbal according to an embodiment of the present application;
[0050] Figure 4 Block diagram of the intelligent tracking shooting gimbal device according to an embodiment of the present application;
[0051] Figure 5 Schematic structural diagram of a terminal device or a server suitable for implementing the embodiments of the present application. Detailed implementation manners
[0052] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0053] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0054] Figure 1 Flow chart of the intelligent tracking shooting method according to an embodiment of the present application. Refer to Figure 1 , the method includes:
[0055] S101, obtain a first picture including a key target captured by an AI camera in a shooting gimbal, input the first picture into a target detection model, and obtain the position of the key target.
[0056] Among them, the key target includes but is not limited to people and balls in sports events. The selection of the key target can be customized by the user on a mobile terminal (including but not limited to a smart phone and a tablet device) through methods such as touch, voice control, and gesture control. Then, the mobile terminal sends the key target information to the shooting gimbal.
[0057] The object detection model is used to identify and locate the category and location information of key objects from images and video streams. The object detection model adopted in this application includes, but is not limited to, the YOLO (You Only Look Once) model. By dynamically dividing the first frame image containing key objects into variable-density grids of 52×52 to 104×104, the grid granularity is automatically adjusted according to the initial size of the object to ensure the detection recall rate of small objects (such as 10×10 pixels). Then, multiple bounding boxes and class probabilities are predicted for each grid, thereby achieving end-to-end real-time detection.
[0058] In this embodiment, the object detection model is used to achieve accurate identification of key objects, has the anti-interference ability in complex scenarios, and effectively overcomes the problem of missed detection in traditional visual detection.
[0059] S102. Convert the key object position into the first physical coordinate according to the distortion parameter and rotation angle of the AI camera.
[0060] In a possible implementation manner, the key object position is (u, v). First, undistort the key object position. The calculation formula for the undistorted coordinates (u', v') is as follows:
[0061]
[0062] r 2 = u d 2 + v d 2 ,
[0063] u c = u d (1 + k1r 2 + k2r 4 + k3r 6 ) + 2p1uv d v d + p2(r 2 + 2u d 2 ),
[0064] v c = v d (1 + k1r 2 + k2r 4 + k3r 6 ) + p1(r 2 + 2v d 2 ) + 2p2uv d v d ,
[0065] u' = f x uc +c x , v' = f y v c +c y ,
[0066] Among them, (c x , c y ) is the coordinate of the center point of the screen, f x , f y are the focal lengths of the AI camera, k1, k2, and k3 are the radial distortion coefficients, and p1 and p2 are the tangential distortion coefficients. Then, the normalized transformation is performed on the undistorted coordinates (u', v') to obtain the normalized coordinates:
[0067]
[0068] Finally, the normalized coordinates (x, y) are rotated to the direction of the first physical coordinate system. The calculation method of the first physical coordinates (x dir , y dir , z dir ) is as follows:
[0069]
[0070] Among them, R is the preset rotation matrix mapped to the direction of the first physical coordinate system.
[0071] In this embodiment, through the dual processing of distortion correction and spatial coordinate mapping, an accurate physical space coordinate system is established, eliminating the positioning deviation caused by the distortion of the AI camera, and providing standardized spatial data input for subsequent intelligent tracking shooting.
[0072] S103, perform intelligent tracking shooting on the key target according to the first physical coordinates and the multi-target tracking algorithm.
[0073] In a possible implementation manner, in addition to performing intelligent tracking shooting on the key target, continuously obtain the moving images of the key target and analyze the region of interest in the images, and send a rotation instruction to the shooting pan-tilt according to the movement of the region of interest in the images. Generating an intelligent control rotation instruction for the shooting pan-tilt based on the analysis of the region of interest does not require achieving the tracking purpose through panoramic stitching or dynamic cropping, reducing the cost of event shooting while ensuring high-resolution shooting.
[0074] Among them, the Region of Interest (ROI) is a local area that needs to be processed by outlining the image to be processed in the form of a square, circle, ellipse, irregular polygon, etc. The selection methods of the Region of Interest include but are not limited to user-defined selection and automatic algorithm selection. In user-defined selection, the user can manually smear or outline the range of the Region of Interest in the interactive interface. In automatic algorithm selection, after the user selects the key target, the target detection model outputs the initial bounding box of the key target, and this bounding box is the initial ROI region. For example, when the user frames the athlete running in the picture, the target detection model will generate a rectangular box covering the whole body of the athlete, and this box is the initial ROI region. Then, by analyzing the spatio-temporal gradient change between adjacent frames containing the key target through the optical flow method, the movement of the key target is inferred, and a new ROI region is obtained to cover the position where it may reach in the next frame, avoiding the key target moving out of the detection range. Among them, the optical flow method aims to improve the efficiency and accuracy of motion estimation through local calculation. First, corner points or feature points with significant texture are extracted within the initial ROI region. Then, based on the gradient equation constructed under the assumption of brightness constancy and the feature points of the initial ROI region extracted, the key feature points of the new ROI region are calculated, thereby obtaining the new ROI region.
[0075] Furthermore, according to the movement of the Region of Interest in the picture, a rotation instruction is sent to the shooting pan-tilt head, generating an AI camera rotation instruction that combines the analysis of the position change of the Region of Interest and the physical space coordinate conversion. First, calculate the pixel offset of the center point (x c , y c ) of the ROI in the current picture frame and the center point (x0, y0) of the initial ROI region. The calculation formula is: Δx = x c - x0, Δy = y c - y0. Then, according to the internal parameters of the camera and the current pitch angle and yaw angle of the shooting pan-tilt head, the pixel offset is converted into the angle that the pan-tilt head needs to adjust. The calculation formula is as follows:
[0076]
[0077] Among them, Δφ is the calculated yaw angle that needs to be adjusted, Δθ is the calculated pitch angle that needs to be adjusted, s is the pixel size of the AI camera, f is the focal length of the AI camera, and φ is the current initial yaw angle of the shooting pan-tilt head.
[0078] In this embodiment, the multi-target tracking algorithm organically combines key target trajectory prediction and motion compensation, ensuring tracking continuity in complex scenarios.
[0079] Optionally, intelligent tracking shooting of the key target is performed according to the first physical coordinate and the multi-target tracking algorithm, including:
[0080] The multi-object tracking algorithm predicts the movement of the key object based on the first physical coordinates to obtain the second physical coordinates of the key object;
[0081] Automatically adjust the pitch angle and yaw angle of the shooting pan-tilt according to the second physical coordinates, and continuously obtain the second picture containing the key object.
[0082] Among them, the multi-object tracking algorithm (Simple Online and Realtime Tracking, SORT) is an efficient multi-object tracking algorithm based on motion modeling and data association. Its core idea is to achieve continuous tracking of cross-frame object identities by fusing the object motion prediction in the time dimension and the detection box position information in the space dimension. First, obtain the position and bounding box of the key object detected by the object detection model in each frame. Then, use the Kalman filter to predict the bounding box position of the key object in the next frame according to its historical trajectory (i.e., the first physical coordinates). Finally, introduce the Hungarian algorithm to establish the association relationship between the current frame containing the key object and the previous frame containing the key object, and match the key object in the current frame and the key object in the previous frame by minimizing the association cost, ensuring that the bounding box of each key object is associated with at most one predicted bounding box, so as to obtain the second physical coordinates of the key object.
[0083] In a possible implementation, if the first physical coordinates are (x1, y1) and the second physical coordinates predicted by the multi-object tracking algorithm are (x2, y2), then the yaw angle Δφ that the shooting pan-tilt needs to adjust is:
[0084]
[0085] where D is the target horizontal parameter. The pitch angle Δθ that the shooting pan-tilt needs to adjust is:
[0086]
[0087] In addition, the rotation speed of the shooting pan-tilt can be controlled by a PID controller, and a smooth rotation speed command is generated according to the pitch angle difference or yaw angle difference to avoid the shaking of the shooting pan-tilt. The calculation formula of the smooth rotation speed u(t) is as follows:
[0088]
[0089] where e(t) is the pitch angle difference or yaw angle difference, K p is a hyperparameter with a set value of 0.8, K i is a hyperparameter with a set value of 0.2, K d is a hyperparameter with a set value of 0.1.
[0090] In this embodiment, through the deep collaboration between the multi-target tracking algorithm and the shooting pan-tilt control in the shooting pan-tilt, high-robustness intelligent tracking in complex event shooting scenarios is achieved.
[0091] Optionally, the method further includes:
[0092] Performing face feature extraction on the first frame through a face recognition, face bone point detection model, and face feature extraction model to obtain first face feature data;
[0093] Performing human body feature extraction on the first frame through a human body feature extraction model to obtain first human body feature data;
[0094] Sequentially inputting the second frame into the human body feature extraction model for feature extraction to obtain second human body feature data, screening third human body feature data similar to the first human body feature data from the second human body feature data, and performing intelligent target tracking shooting on the third frame corresponding to the third human body feature data;
[0095] Performing face feature extraction on the third frame through a face recognition, face bone point detection model, and face feature extraction model to obtain second face feature data, and performing secondary confirmation on the third frame according to the first face feature data and the second face feature data.
[0096] Figure 2 It is a flowchart for confirming key targets in intelligent tracking shooting according to an embodiment of the present application, as Figure 2 shown.
[0097] Among them, in the face recognition and face bone point detection model, the lightweight model RetinaFace-MobileNet is adopted to quickly locate the face bounding box of the key target in the captured image. Then, based on the MediaPipe Face Mesh model, 468 face key points (including but not limited to eyelids, corners of the mouth, and the tip of the nose) are extracted to construct the face topology of the key target. The face feature extraction model includes but not limited to arcface, OpenFace, and Dlib, which are used to extract the face feature data in the image. The human body feature extraction model includes but not limited to CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), and LSTM (Long Short-Term Memory), which are used to extract the human body feature data in the image. In the present invention, when the shooting pan-tilt performs intelligent target tracking shooting on the key target, by comparing the similarity of the human body feature data, it is confirmed that the tracking of the key target is correct. At the same time, the shooting pan-tilt synchronizes the image of the key target captured by tracking to the mobile terminal, and the mobile terminal reconfirms that the tracking shooting of the key target is correct by comparing the similarity of the face feature data, realizing the matching of the key target in the intelligent terminal and the shooting pan-tilt.
[0098] In this embodiment, by combining the dual biometric features of the face and the human body, the accuracy of key target locking and tracking is improved, and interferences such as occlusion and changing clothes are avoided. In addition, through the preliminary screening of human body feature data and the secondary matching of face feature data, the false tracking rate of shooting is reduced.
[0099] Optionally, the face feature extraction model for face recognition and face bone point detection is used to extract face feature data from the first image, and the obtained first face feature data includes:
[0100] The face recognition and face bone point detection model is used to perform face positioning and alignment on the first image to obtain the first face data and the first face bone point data;
[0101] The first face data and the first face bone point data are input into the face feature extraction model to obtain the first face feature data.
[0102] In a possible implementation manner, if the face recognition and face bone point detection model cannot detect the first face data and the first face bone point data in the first image, continuous detection is performed in the subsequent frames of the first image until the first face data and the first face bone point data of the key target are obtained, and the first face data and the first face bone point data are saved in the mobile terminal.
[0103] In this embodiment, geometric alignment is achieved through face recognition and a face bone point detection model, which eliminates the interference of pose differences on face feature extraction and improves the discriminability of the face feature data extracted by the face feature extraction model.
[0104] Optionally, the secondary confirmation of the third picture based on the first face feature data and the second face feature data includes:
[0105] Perform similarity matching between the first face feature data and the second face feature data;
[0106] If the first face feature data and the second face feature data are similar, continue with intelligent target tracking shooting;
[0107] If the first face feature data and the second face feature data are not similar, re-screen the third picture.
[0108] For example, if the first face feature data is [0.12, 0.34, 0.56,... 0.78] and the second face feature data is [0.13, 0.35, 0.57,... 0.79], the similarity matching value between the first face feature data and the second face feature data calculated using the Euclidean distance is 0.05, and the similarity threshold of the face feature data is set to 0.5, then the first face feature data and the second face feature data are not similar, and the tracking shooting of the key target by the AI camera fails. It is necessary to compare the human body feature data in the captured pictures again and re-screen the third picture.
[0109] In this embodiment, the accurate confirmation of the key target is achieved through the similarity matching of face feature data, ensuring the continuity of shooting and tracking in complex competition scenarios.
[0110] Optionally, the method further includes:
[0111] When performing intelligent target tracking shooting on the key target, randomly select the third picture, input the third picture into the human body feature extraction model for feature extraction, and obtain the fourth human body feature data;
[0112] Perform similarity matching between the fourth human body feature data and the target feature queue;
[0113] If the fourth human body feature data and the target feature queue are similar, continue with intelligent target tracking shooting;
[0114] If the fourth human body feature data and the target feature queue are not similar, the intelligent target tracking shooting fails, pause the intelligent target tracking shooting, and reselect the third picture;
[0115] The target feature queue is a feature sequence set of all the third human body feature data collected historically.
[0116] For example, the fourth human body feature data is [0.1, 0.6, 0.8,..., 0.9], and the target feature queue is [[0.1, 0.3, 0.5,..., 0.2],..., [0.2, 0.4, 0.7,..., 0.3]]. The similarity values between the fourth human body feature data and each feature data in the target feature queue are calculated in turn through the Euclidean distance, and then all the similarity values are averaged to obtain the similarity value of 0.3 between the fourth human body feature data and the target feature queue. Then, the fourth human body feature data is not similar to the target feature queue, and the tracking and shooting of the key target of the AI camera fails. It is necessary to compare the human body feature data in the captured picture again and re-screen the third picture.
[0117] In this embodiment, the fourth human body feature data is randomly sampled and matched with the target feature queue, ensuring that the AI camera can continuously track the correct key target.
[0118] Optionally, the shooting pan-tilt head is communicatively connected to the mobile terminal, and the method further includes:
[0119] Obtaining an operation instruction of the user on the mobile terminal;
[0120] Adjusting the pitch angle and yaw angle of the shooting pan-tilt head according to the operation instruction, and zooming in on the picture selected in the operation instruction on the mobile terminal.
[0121] Among them, the operation instructions of the user include but are not limited to text input, voice input, and gesture control. If the user selects a specific point in the shooting picture on the display screen to zoom in, the shooting pan-tilt head automatically adjusts the shooting focal length. If the user rotates the picture by swiping left and right on the display screen, the shooting pan-tilt head automatically and evenly adjusts the pitch angle and yaw angle to provide the picture required by the user.
[0122] In this embodiment, adjusting the pitch angle and yaw angle of the shooting pan-tilt head according to the operation instruction of the user realizes real-time user interaction and improves the user experience.
[0123] Figure 3 FIG. is a schematic structural diagram of the interaction between the intelligent terminal and the shooting pan-tilt head according to an embodiment of the present application, as Figure 3 shown:
[0124] There is real-time communication between the mobile terminal and the shooting gimbal, thereby realizing real-time intelligent tracking shooting of key targets. The mobile terminal includes but is not limited to smartphones and cameras. The shooting gimbal integrates an NPU chip, a display screen, and an AI camera. In the NPU chip, various machine learning models and algorithms such as a target detection model, a face recognition and face bone point detection model, a face feature extraction model, a human body feature extraction model, a multi-target tracking algorithm, and a position coordinate conversion algorithm can be integrated, enabling real-time scene analysis and avoiding the burden and latency of cloud processing. The AI camera uses a wide-angle lens + AISoC (AI System on Chip) chip, which can be used for intelligent tracking shooting of events and can also be used to collect and integrate image data captured by another camera in the mobile terminal or the shooting gimbal. The display screen can directly display the execution results captured by the shooting gimbal after receiving the drive command, and it is in the same direction as the AI camera so that the user can directly control the shooting gimbal through the display screen and then interact.
[0125] In a possible implementation, in order to enable the shooting gimbal to directly interact with the user and correctly recognize the user's operation instructions, the AI camera is used to capture the user's specified actions as the start of the operation instructions, preventing accidental interruption or accidental start of shooting due to misrecognition of the user's actions or gestures in the normal shooting and tracking scenario. For example, action instructions and a specified completion time are displayed on the display screen of the shooting gimbal. The action instructions include but are not limited to opening the palm. The AI camera recognizes that the user has completed the action instructions on the display screen within the specified time before starting to recognize the user's operation instructions.
[0126] According to the embodiments of the present disclosure, the following technical effects are achieved:
[0127] 1) Realize real-time automated intelligent tracking control of the AI camera in the shooting gimbal, significantly improving the tracking shooting efficiency of key targets.
[0128] 2) Integrate multi-dimensional perception data and various machine learning algorithms, ensuring the accuracy of key target positioning and the high resolution of event shooting.
[0129] 3) Establish a dynamic feedback adjustment mechanism, and through multiple authentications, ensure the effectiveness of the intelligent tracking shooting of the AI camera for key targets.
[0130] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0131] Figure 4 The block diagram of the intelligent tracking and shooting pan-tilt device according to an embodiment of the present application is shown. As Figure 4 shown, it includes:
[0132] An AI camera 401, configured to obtain a first picture including a key target, input the first picture into a target detection model to obtain the position of the key target; convert the position of the key target into a first physical coordinate according to the distortion parameter and rotation angle of the AI camera; perform intelligent tracking on the key target according to the first physical coordinate and a multi-target tracking algorithm; send a rotation instruction to the driving mechanism;
[0133] A clamping structure 402, configured to fix an intelligent terminal;
[0134] A driving mechanism 403, configured to receive the rotation instruction and drive the shooting pan-tilt to rotate.
[0135] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0136] Figure 5 The structural schematic diagram of a terminal device or a server suitable for implementing the embodiments of the present application is shown.
[0137] As Figure 5 shown, the terminal device or the server includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the terminal device or the server are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0138] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.
[0139] Specifically, according to an embodiment of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a machine-readable medium, the computer program including program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by a central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.
[0140] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0142] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0143] As another aspect, this application also provides a computer-readable storage medium. The computer-readable storage medium can be included in the electronic device described in the above embodiments; it can also exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors, the methods described in this application are implemented.
[0144] The above description is only a preferred embodiment of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the application involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing application concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions described in this application.
Claims
1. An intelligent tracking shooting method, characterized in that, Including: Obtain a first frame containing a key target captured by an AI camera in a shooting gimbal, input the first frame into a target detection model, and obtain the position of the key target; Convert the position of the key target into a first physical coordinate according to the distortion parameter and rotation angle of the AI camera; Intelligently track and shoot the key target according to the first physical coordinate and a multi-object tracking algorithm.
2. The intelligent tracking and shooting method according to claim 1, wherein The intelligently tracking and shooting the key target according to the first physical coordinate and a multi-object tracking algorithm includes: Predict the movement of the key target by the multi-object tracking algorithm according to the first physical coordinate to obtain a second physical coordinate of the key target; Automatically adjust the pitch angle and yaw angle of the shooting gimbal according to the second physical coordinate, and continuously obtain a second frame containing the key target.
3. The intelligent tracking and shooting method according to claim 2, characterized in that, The method further includes: Extract first face feature data by performing face feature extraction on the first frame through a face recognition, face bone point detection model, and face feature extraction model; Extract first human body feature data by performing human body feature extraction on the first frame through a human body feature extraction model; Input the second frame into the human body feature extraction model in sequence for feature extraction to obtain second human body feature data, screen third human body feature data similar to the first human body feature data from the second human body feature data, and perform intelligent target tracking and shooting on a third frame corresponding to the third human body feature data; Extract second face feature data by performing face feature extraction on the third frame through the face recognition, face bone point detection model, and the face feature extraction model, and perform secondary confirmation on the third frame according to the first face feature data and the second face feature data.
4. The intelligent tracking and shooting method according to claim 3, characterized in that The extracting first face feature data by performing face feature extraction on the first frame through a face recognition, face bone point detection model, and face feature extraction model includes: Perform face positioning and alignment on the first frame through the face recognition, face bone point detection model to obtain first face data and first face bone point data; Input the first face data and the first face bone point data into the face feature extraction model to obtain the first face feature data.
5. The intelligent tracking and shooting method according to claim 3, characterized in that, The performing secondary confirmation on the third frame according to the first face feature data and the second face feature data includes: Perform similarity matching on the first face feature data and the second face feature data; If the first face feature data and the second face feature data are similar, continue the intelligent target tracking and shooting; If the first face feature data and the second face feature data are not similar, re-screen the third frame.
6. The intelligent tracking and shooting method according to claim 3, characterized in that, The method further includes: When performing the intelligent target tracking and shooting on the key target, randomly extract the third frame, input the third frame into the human body feature extraction model for feature extraction to obtain fourth human body feature data; Perform similarity matching on the fourth human body feature data and a target feature queue; If the fourth human body feature data is similar to the target feature queue, continue with the intelligent target tracking shooting; If the fourth human body feature data is not similar to the target feature queue, the intelligent target tracking shooting fails, pause the intelligent target tracking shooting, and reselect the third picture; The target feature queue is a feature sequence set of all the third human body feature data collected historically.
7. The intelligent tracking and shooting method according to claim 3, wherein The shooting pan-tilt is communicatively connected to the mobile terminal, and the method further includes: Obtain an operation instruction of the user on the mobile terminal; Adjust the pitch angle and yaw angle of the shooting pan-tilt according to the operation instruction, and zoom in on the picture selected in the operation instruction on the mobile terminal.
8. An intelligent tracking and shooting pan-tilt, characterized in that, Includes: An AI camera for obtaining a first picture containing a key target, inputting the first picture into a target detection model, and obtaining the position of the key target; Convert the key target position to a first physical coordinate according to the distortion parameter and rotation angle of the AI camera; perform intelligent tracking on the key target according to the first physical coordinate and a multi-target tracking algorithm; send a rotation instruction to the drive mechanism; A clamping structure for fixing the intelligent terminal; A drive mechanism for receiving the rotation instruction and driving the shooting pan-tilt to rotate.
9. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Competition video editing method and system based on tracking holder data, and storage device
CN120916033A