Intelligent tracking shooting holder, shooting method, equipment and storage medium
The intelligent tracking and shooting gimbal, which combines AI cameras and multi-target tracking algorithms, solves the problems of autonomous shooting decisions and high hardware costs in dynamic scenes, and enables precise tracking of athletes from specific teams and real-time live broadcasts of events.
Patent Information
- Application Number
- CN202580001751.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-04-09
- Filing Date
- 2025-04-22
- Publication Date
- 2025-11-07
AI Technical Summary
Existing tracking and shooting technologies cannot make autonomous shooting decisions in dynamic scenes, make it difficult to accurately distinguish individual athletes from specific teams, and have high hardware costs, making it difficult to meet the real-time requirements of live event broadcasts.
The intelligent tracking and shooting gimbal combines an AI camera with a multi-target tracking algorithm. It uses the AI camera to obtain the position of key targets and uses the drive mechanism to achieve 360-degree horizontal and 90-degree vertical rotation. It also combines facial and human feature extraction for accurate tracking, reducing hardware costs.
It enables precise tracking and filming of athletes from specific teams in dynamic scenes, reduces hardware costs, and meets the real-time requirements of live event broadcasts.
Smart Images

Figure CN120917262A_ABST
Abstract
Description
[0001] The present application claims priority to the following Chinese patent applications: Application No. 2025104198626, entitled "Intelligent Tracking and Shooting Method, Shooting Gimbal, Device and Storage Medium", filed on April 3, 2025; Application No. 2025104388584, entitled "Real-time Tracking and Shooting Method and Device Based on Jersey Number", filed on April 9, 2025; and Application No. 202423143844X, entitled "Tracking Gimbal and Tracking and Shooting Device", filed on December 9, 2024, the contents of which are hereby incorporated by reference in their entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of image recognition technology, and in particular to an intelligent tracking and shooting gimbal, a shooting method, a device and a storage medium. BACKGROUND
[0003] Tracking and shooting technology is an important means of sports event image acquisition, and is widely used in professional scenes such as athlete motion capture and live broadcast of events. However, the current video tracking and shooting method still has significant technical bottlenecks in intelligent tracking and multi-angle acquisition in dynamic scenes. First, the traditional tracking and shooting gimbal system mostly uses fixed-angle shooting or manual control shooting, which cannot realize autonomous shooting decision in dynamic scenes, and is prone to missing multi-dimensional image information. Second, the tracking algorithm generally uses the maximum target priority principle or feature recognition based on human clothing color, which can only track the entire group in the field, and cannot accurately distinguish individual athletes of a specific team, lacks face feature extraction and multi-modal verification mechanism, and is prone to target misjudgment or loss in scenes where multiple people wear similar clothes, resulting in insufficient dynamic tracking accuracy.
[0004] Furthermore, the traditional personnel tracking algorithm relies on clothing color and appearance features for target association, which is prone to feature confusion in scenes where team uniforms are similar in color and athletes are densely interlaced, leading to broken tracking trajectories or incorrect identity judgments, which seriously affects the accuracy of data detection. In addition, some studies attempt to integrate features such as athlete hairstyles and body shapes to improve robustness, but complex multi-modal models are large in size and difficult to run in real time on mobile devices with low computing power, limiting their application in live broadcast, instant playback and other scenarios.
[0005] In addition, the driving device of the shooting gimbal in the prior art is usually arranged on the holder connected to the handle, and full-range tracking is achieved by controlling the rotation of the holder, which has the technical problem of complex driving structure arrangement. Moreover, the existing automatic shooting method for sports events generally uses panoramic stitching technology to construct a panoramic image through multi- wide-angle lens distortion correction and image stitching, which has a large hardware storage cost and is difficult to meet the real-time requirements of live broadcast of events.
[0006] Therefore, there is an urgent need for a real-time intelligent tracking shooting method to meet the intelligent, low-cost and efficient shooting requirements of sports events. SUMMARY
[0007] According to the embodiments of the present application, an intelligent tracking shooting cloud platform and a shooting method are provided, which can realize efficient intelligent tracking shooting and greatly reduce the cost of sports event shooting.
[0008] In the first aspect of the present application, an intelligent tracking shooting cloud platform is provided, which comprises a cloud platform base, a cloud platform shell, a tracker and a driving mechanism, wherein the tracker and the driving mechanism are arranged in the cloud platform shell;
[0009] The tracker comprises an AI camera, which is a tracking sensor in the application No. 202423143844X. The AI camera is used to acquire a first image containing a key target, input the first image into a target detection model, and acquire a key target position. The key target position is converted into a first physical coordinate according to the distortion parameter and the rotation angle of the AI camera. The key target is intelligently tracked according to the first physical coordinate and a multi-target tracking algorithm, and a rotation instruction is sent to the driving mechanism.
[0010] The driving mechanism is used to receive the rotation instruction and drive the shooting cloud platform to rotate. The driving mechanism comprises a first driving mechanism and a second driving mechanism. The first driving mechanism is drivingly connected to the cloud platform base to realize 360-degree horizontal rotation of the tracker. The second driving mechanism is drivingly connected to the tracker to realize 90-degree vertical tilting rotation of the tracker.
[0011] Optionally, the key target includes but is not limited to a person and a ball.
[0012] In a possible implementation, in addition to intelligently tracking the key target, moving images of the key target are continuously acquired and the image interesting region is analyzed, and a rotation instruction is sent to the shooting cloud platform according to the movement of the image interesting region.
[0013] Optionally, the selection method of the image interesting region includes but is not limited to user-defined selection and algorithm automatic selection. In the user-defined selection, the user can manually smear or outline the range of the image interesting region in the interactive interface.
[0014] In a possible implementation, intelligently tracking the key target according to the first physical coordinate and the multi-target tracking algorithm comprises:
[0015] The second physical coordinates of the key target are obtained by predicting the movement of the key target according to the first physical coordinates by the multi-target tracking algorithm;
[0016] The pitch angle and the yaw angle of the shooting holder are automatically adjusted according to the second physical coordinates, and the second picture containing the key target is continuously obtained.
[0017] Optionally, the system further comprises a face recognition and face skeleton point detection model, a face feature extraction model, and a human body feature extraction model;
[0018] The face feature extraction model is used to extract face features from the first picture to obtain first face feature data;
[0019] The human body feature extraction model is used to extract human body features from the first picture to obtain first human body feature data;
[0020] The second picture is sequentially input into the human body feature extraction model for feature extraction to obtain second human body feature data, third human body feature data similar to the first human body feature data is screened from the second human body feature data, and a third picture corresponding to the third human body feature data is intelligently tracked and shot;
[0021] The face feature extraction model is used to extract face features from the third picture to obtain second face feature data, and the third picture is twice confirmed according to the first face feature data and the second face feature data.
[0022] Optionally, the face feature extraction model is used to extract face features from the first picture to obtain first face feature data, and the method comprises the following steps:
[0023] The face recognition and face skeleton point detection model is used to locate and align the face of the first picture to obtain first face data and first face skeleton point data;
[0024] The first face data and the first face skeleton point data are input into the face feature extraction model to obtain the first face feature data.
[0025] Optionally, the third picture is twice confirmed according to the first face feature data and the second face feature data, and the method comprises the following steps:
[0026] The first face feature data and the second face feature data are matched in similarity;
[0027] If the first face feature data and the second face feature data are similar, the intelligent target tracking and shooting is continued;
[0028] If the first facial feature data and the second facial feature data are not similar, the third picture is reselected.
[0029] Optionally, when the key target is photographed by the intelligent target tracking, the third picture is randomly selected, the third picture is input into the human feature extraction model for feature extraction, and fourth human feature data is obtained.
[0030] The fourth human feature data is matched with the target feature queue in similarity.
[0031] If the fourth human feature data is similar to the target feature queue, the intelligent target tracking continues to be performed.
[0032] If the fourth human feature data is not similar to the target feature queue, the intelligent target tracking fails, the intelligent target tracking is paused, and the third picture is reselected.
[0033] The target feature queue is a feature sequence set of all third human feature data collected historically.
[0034] In a possible implementation, the shooting gimbal is in communication connection with the mobile terminal, and specifically:
[0035] An operation instruction of a user on the mobile terminal is obtained.
[0036] The pitch angle and the yaw angle of the shooting gimbal are adjusted according to the operation instruction, and the picture selected in the operation instruction is zoomed in the mobile terminal.
[0037] In a possible implementation, the shooting gimbal further includes a display screen, which can directly display the execution result of the shooting gimbal after receiving the driving instruction.
[0038] In a possible implementation, the AI camera captures the specified action of the user as the start of the operation instruction, and the AI camera starts to recognize the operation instruction of the user only when the user completes the action instruction in the display screen within a specified time.
[0039] Further, the tracker is fixed above the tracker and is provided with a clamping structure for clamping the camera.
[0040] Further, the tracker further includes a tracking housing for fixedly installing the AI camera, one side of the tracking housing is provided with an arc-shaped rack, and the second driving mechanism is provided with a driving gear matched with the arc-shaped rack.
[0041] Further, the middle part of the gimbal base is provided with a bearing seat, and the driver of the first driving mechanism is rotationally fixedly connected to the bearing seat.
[0042] Further, a control main plate is arranged between the driver of the first driving mechanism and the bearing seat, and the control main plate is fixedly connected to the driver of the first driving mechanism, and the control main plate and the driver of the first driving mechanism are fixedly connected to the holder shell.
[0043] Further, a support assembly is arranged in the accommodating space of the holder shell, and the support assembly comprises a main support and an auxiliary support, and the auxiliary support is arranged on both sides of the tracker.
[0044] Further, a bracket bearing is arranged on the auxiliary support, and a protruding column is arranged on both sides of the tracker, and the protruding column is rotationally fixedly connected to the bracket bearing.
[0045] Further, the AI camera is arranged on one side of the main support, and a power supply is arranged on the other side of the main support.
[0046] Further, a switch button is arranged on the holder shell.
[0047] In a possible implementation, the AI camera is further adapted to intelligently track and shoot the jersey number, and specifically comprises:
[0048] obtaining a first human body picture containing a key target;
[0049] inputting the first human body picture into a target detection model, identifying a number region in the first human body picture and performing cutting, and obtaining a first number picture;
[0050] inputting the first number picture into a number classification model, and obtaining first jersey number data;
[0051] matching the key target in a real-time shooting picture according to the first jersey number data and performing tracking shooting.
[0052] Optionally, matching the key target in a real-time shooting picture according to the first jersey number data and performing tracking shooting comprises:
[0053] inputting continuous real-time shooting pictures into a target detection model, sequentially identifying number regions in the real-time shooting pictures and performing cutting, and obtaining a second number picture queue;
[0054] sequentially performing number classification on second number pictures in the second number picture queue through the number classification model, and obtaining a second jersey number data queue;
[0055] find second shirt number data matching the first shirt number data in the second shirt number data queue, if there is second shirt number data matching the first shirt number data, track and shoot the first human target corresponding to the second shirt number data through a target tracking algorithm, if there is no second shirt number data matching the first shirt number data, continue to identify and match.
[0056] Optionally, tracking and shooting the first human target corresponding to the second shirt number data through the target tracking algorithm comprises:
[0057] predicting the movement of the first human target through the target tracking algorithm to obtain a next predicted position of the first human target;
[0058] automatically adjusting the pitch angle and the yaw angle of the shooting holder according to the next predicted position to continuously obtain a second human picture containing the first human target.
[0059] Optionally, the second human picture continuously obtained is sequentially identified and classified through the target detection model and the number classification model to obtain a third shirt number data queue.
[0060] the third shirt number data in the third shirt number data queue and the second shirt number data are sequentially matched, if the third shirt number data in the third shirt number data queue and the second shirt number data are consistent, the tracking and shooting is continued, if there is third shirt number data inconsistent with the second shirt number data in the third shirt number data queue, the tracking and shooting fails, and the key target is matched in the real-time shooting picture for tracking and shooting again.
[0061] In a possible implementation, the continuous real-time shooting picture is input into the target detection model, the number area in the real-time shooting picture is sequentially identified and cut to obtain a second number picture queue, and the method further comprises:
[0062] When the real-time shooting picture is a long-distance shooting, the real-time shooting picture is cut into picture blocks, and each picture block is enlarged by a preset magnification;
[0063] Each enlarged picture block is identified and cut for the number area through the target detection model to obtain the second number picture queue.
[0064] Optionally, if the fourth shirt number data on the second human target in the real-time shooting picture is detected through the target detection model and the number classification model during the tracking and shooting, and the fourth shirt number data is consistent with the second shirt number data, the tracking target of the tracking and shooting is transferred to the second human target.
[0065] Optionally, the first human body picture and the first shirt number data are supplemented to the training data of the target detection model and the number classification model on the server side for model iteration.
[0066] In a second aspect of the application, an intelligent tracking and shooting method is provided. The method comprises:
[0067] obtaining a first picture containing a key target captured by an AI camera in a shooting gimbal, inputting the first picture into a target detection model to obtain a key target position;
[0068] converting the key target position into a first physical coordinate according to the distortion parameters and the rotation angle of the AI camera;
[0069] intelligently tracking and shooting the key target according to the first physical coordinate and a multi-target tracking algorithm.
[0070] Optionally, the key target includes but is not limited to a person and a ball.
[0071] In a possible implementation, in addition to intelligently tracking and shooting the key target, moving pictures of the key target are continuously obtained and the picture region of interest is analyzed, and a rotation instruction is sent to the shooting gimbal according to the movement of the picture region of interest
[0072] Optionally, the selection method of the picture region of interest includes but is not limited to user-defined selection and algorithm automatic selection. In the user-defined selection, the user can manually smear or outline the range of the picture region of interest in the interactive interface.
[0073] In a possible implementation, intelligently tracking and shooting the key target according to the first physical coordinate and the multi-target tracking algorithm comprises:
[0074] predicting the movement of the key target according to the first physical coordinate by the multi-target tracking algorithm to obtain a second physical coordinate of the key target;
[0075] automatically adjusting the pitch angle and the yaw angle of the shooting gimbal according to the second physical coordinate to continuously obtain a second picture containing the key target.
[0076] Optionally, the method further comprises:
[0077] extracting facial features of the first picture by a face recognition and face skeleton point detection model and a face feature extraction model to obtain first facial feature data;
[0078] extracting human body features of the first picture by a human body feature extraction model to obtain first human body feature data;
[0079] The second picture is input into the human feature extraction model in sequence for feature extraction, second human feature data is obtained, third human feature data similar to the first human feature data is screened from the second human feature data, and a third picture corresponding to the third human feature data is captured through intelligent target tracking;
[0080] The third picture is subjected to face feature extraction through the face recognition and face skeleton point detection model and the face feature extraction model, second face feature data is obtained, and the third picture is subjected to secondary confirmation according to the first face feature data and the second face feature data.
[0081] Optionally, the first picture is subjected to face feature extraction through the face recognition and face skeleton point detection model and the face feature extraction model, and first face feature data is obtained, including:
[0082] The first picture is subjected to face positioning and alignment through the face recognition and face skeleton point detection model, and first face data and first face skeleton point data are obtained;
[0083] The first face data and the first face skeleton point data are input into the face feature extraction model, and the first face feature data is obtained.
[0084] Optionally, the third picture is subjected to secondary confirmation according to the first face feature data and the second face feature data, including:
[0085] The first face feature data and the second face feature data are subjected to similarity matching;
[0086] If the first face feature data and the second face feature data are similar, intelligent target tracking capturing is continued;
[0087] If the first face feature data and the second face feature data are not similar, the third picture is re-screened.
[0088] Optionally, the method further includes:
[0089] When the key target is captured through intelligent target tracking, the third picture is randomly extracted, the third picture is input into the human feature extraction model for feature extraction, and fourth human feature data is obtained;
[0090] The fourth human feature data is subjected to similarity matching with a target feature queue;
[0091] If the fourth human feature data is similar to the target feature queue, intelligent target tracking capturing is continued;
[0092] If the fourth human feature data is not similar to the target feature queue, intelligent target tracking capturing fails, intelligent target tracking capturing is paused, and the third picture is reselected;
[0093] The target feature queue is a feature sequence set of all third human body feature data collected historically.
[0094] In a possible implementation, the shooting gimbal is in communication connection with the mobile terminal, and the method further includes:
[0095] obtaining an operation instruction of the user on the mobile terminal;
[0096] adjusting the pitch angle and the yaw angle of the shooting gimbal according to the operation instruction, and zooming the selected picture in the operation instruction in the mobile terminal.
[0097] In a possible implementation, the shooting gimbal further includes a display screen, which can directly display the execution result of the shooting gimbal after receiving the driving instruction.
[0098] In a possible implementation, the AI camera is used to capture the specified action of the user as the start of the operation instruction, and the AI camera starts to recognize the operation instruction of the user only when the user completes the action instruction in the display screen within a specified time.
[0099] In a third aspect of the present application, an electronic device is provided. The electronic device includes a memory and a processor, the memory having a computer program stored thereon, and the processor implementing the method as described above when executing the program.
[0100] In a fourth aspect of the present application, a computer readable storage medium is provided, having a computer program stored thereon, the program being executed by a processor to implement the method according to the second aspect of the present application.
[0101] The advantages of the present application are as follows:
[0102] The intelligent tracking shooting gimbal provided by the embodiments of the present application has the advantages of simple structure, convenient use, high control stability and high accuracy, and the like, by arranging the tracker and the driving mechanism for driving the tracker to rotate in the accommodation space of the gimbal shell, realizing 360-degree horizontal rotation of the tracker by the first driving mechanism, and realizing 90-degree vertical tilting rotation of the tracker by the second driving mechanism, thereby realizing effective and accurate omnidirectional tracking of the shooting gimbal. Further, by obtaining a first picture containing a key target captured by the AI camera in the shooting gimbal, inputting the first picture into a target detection model, and obtaining the position of the key target, the position of the key target is converted into a first physical coordinate according to the distortion parameter and the rotation angle of the AI camera, and the key target is intelligently tracked and shot according to the first physical coordinate and a multi-target tracking algorithm, thereby realizing efficient intelligent tracking and shooting, improving the efficiency of event shooting, and reducing the cost of shooting.
[0103] Further, the intelligent tracking and shooting gimbal of the embodiments of the present application is also suitable for tracking and shooting the jersey number. Specifically, a first human body picture containing a key target is obtained, the first human body picture is input into a target detection model, a number area in the first human body picture is recognized and cropped, and a first number picture is obtained. Then, the first number picture is input into a number classification model, and first jersey number data is obtained. Finally, the key target is matched and tracked in a real-time shooting picture according to the first jersey number data, and the real-time recognition of the jersey number is realized, and a basis is provided for the tracking and shooting of subsequent games.
[0104] It should be understood that the content described in the summary section is not intended to limit the key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0105] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent by describing in detail the following embodiments thereof with reference to the attached drawings in which:
[0106] Figure 1 a flowchart of an intelligent tracking and shooting method according to an embodiment of the present application;
[0107] Figure 2 a flowchart of confirming a key target for intelligent tracking and shooting according to an embodiment of the present application;
[0108] Figure 3 a structural schematic diagram of interaction between a mobile terminal and a shooting gimbal according to an embodiment of the present application;
[0109] Figure 4 a block diagram of an intelligent tracking and shooting gimbal device according to an embodiment of the present application;
[0110] Figure 5 a flowchart of an intelligent tracking and shooting method based on a jersey number according to another embodiment of the present application;
[0111] Figure 6 a schematic diagram of an intelligent tracking and shooting method based on a jersey number according to another embodiment of the present application;
[0112] Figure 7 a schematic diagram of jersey number recognition according to another embodiment of the present application;
[0113] Figure 8 a flowchart of offline training of an end model according to another embodiment of the present application;
[0114] Figure 9Block diagram of the smart tracking shooting platform based on the shirt number according to another embodiment of the present application;
[0115] Figure 10 Schematic diagram of the three-dimensional structure of the shooting gimbal according to another embodiment of the present application;
[0116] Figure 11 Schematic diagram of the three-dimensional structure of the shooting gimbal after removing the gimbal shell according to another embodiment of the present application;
[0117] Figure 12 Schematic diagram of the three-dimensional structure of the tracking shell according to another embodiment of the present application;
[0118] Figure 13 Schematic diagram of the structure of the terminal device or server suitable for implementing the embodiments of the present application.
[0119] Reference signs:
[0120] 1, gimbal base; 2, gimbal shell; 3, tracker; 31, AI camera; 32, tracking shell;
[0121] 41, first driving mechanism; 42, second driving mechanism; 411, bearing seat; 51, main support; 52, auxiliary support;
[0122] 6, clamping structure; 7, switch button; 8, power supply; 9, control mainboard. DETAILED DESCRIPTION
[0123] To make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0124] The term "and / or" in this document is only to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the front and rear associated objects.
[0125] In the present application, unless otherwise explicitly specified and limited, the terms "connection", "fixation" and the like should be understood broadly, for example, "connection" can be fixed connection, or detachable connection, or one body, unless otherwise explicitly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0126] In addition, the technical solutions among various embodiments of the present application can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize the combination, and when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist and is not within the scope of the present application.
[0127] Embodiment one
[0128] Figure 1 A flowchart of an intelligent tracking shooting method according to an embodiment of the present application. Referring to Figure 1 , the method comprises:
[0129] Obtaining a first picture containing a key target shot by an AI camera in a shooting gimbal, inputting the first picture into a target detection model, and obtaining the position of the key target.
[0130] Among them, the key target includes but is not limited to people and balls in sports events, and the selection of the key target can be customized by the user in the mobile terminal (including but not limited to smart phones, tablet devices) through touch, voice control and gesture control and other ways, and then the key target information is sent to the shooting gimbal by the mobile terminal.
[0131] The target detection model is used to identify and locate the category and position information of the key target from the image and video stream. The target detection model used in the present application includes but is not limited to the YOLO (You Only Look Once) model. By dynamically dividing the first picture image containing the key target into a variable density grid of 52x52 to 104x104, automatically adjusting the grid granularity according to the initial size of the target, ensuring the detection recall rate of small targets (such as 10x10 pixels), and then predicting multiple bounding boxes and class probabilities for each grid, an end-to-end real-time detection is realized.
[0132] In this embodiment, the target detection model is used to realize accurate identification of the key target, has anti-interference ability in complex scenes, and effectively overcomes the missed detection problem of traditional visual detection.
[0133] According to the distortion parameters and rotation angles of the AI camera, the position of the key target is converted into the first physical coordinates.
[0134] In one possible implementation, the position of the key target is (u, v). First, the key target position is de-distorted, and the calculation formula of the de-distorted coordinates (u', v') is as follows:
[0135]
[0136]
[0137] uc = u d (1 + k1r 2 +k2r 4 +k3r 6 )+2p1u d v d +p2(r 2 +2u d 2 ),
[0138] v c = v d (1 + k1r 2 +k2r 4 +k3r 6 )+p1(r 2 +2v d 2 )+2p2u d v d ,
[0139] u' = f x u c +c x , v' = f y v c +c y ,
[0140] wherein (c x , c y ) is the center point coordinate of the picture, f x , f y is the focal length of the AI camera, k1, k2, k3 is the radial distortion coefficient, p1, p2 is the tangential distortion coefficient. Then, the de-distorted coordinates (u', v') are normalized and converted to obtain the normalized coordinates:
[0141]
[0142] Finally, the normalized coordinates (x, y) are rotated to the first physical coordinate system direction, and the first physical coordinates (x dir , y dir , z dir ) are calculated as follows:
[0143]
[0144] wherein R is a preset rotation matrix mapped to the first physical coordinate system direction.
[0145] In this embodiment, through the double processing of distortion correction and spatial coordinate mapping, an accurate physical space coordinate system is established, and the positioning deviation caused by the distortion of the AI camera is eliminated, providing standardized space data input for subsequent intelligent tracking and shooting.
[0146] According to the first physical coordinates and the multi-target tracking algorithm, the key target is intelligently tracked and photographed.
[0147] In a possible implementation, in addition to intelligently tracking and photographing the key target, moving pictures of the key target are continuously acquired and the picture region of interest is analyzed, and a rotation instruction is sent to the photographing gimbal according to the movement of the picture region of interest. The intelligent control rotation instruction of the photographing gimbal is generated based on the analysis of the region of interest, and the tracking purpose is achieved without the need for panoramic splicing or dynamic cropping, thereby reducing the cost of event photography while ensuring high-resolution photography.
[0148] The region of interest (ROI) is a local region that needs to be processed by outlining the processed image in a box, a circle, an ellipse, or an irregular polygon. The selection method of the region of interest includes but is not limited to user-defined selection and algorithm automatic selection. In the user-defined selection, the user can manually smear or outline the range of the region of interest in the interactive interface. In the algorithm automatic selection, the target detection model outputs an initial bounding box of the key target after the user selects the key target, and the initial bounding box is the initial ROI region. For example, the user selects a running athlete in the picture, and the target detection model generates a rectangular box covering the whole body of the athlete, which is the initial ROI region. Then, the motion of the key target is inferred by analyzing the spatiotemporal gradient changes between adjacent picture frames containing the key target by using the optical flow method, and a new ROI region is obtained to cover the position that the key target may reach in the next frame, so as to avoid the key target moving out of the detection range. The optical flow method aims to improve the efficiency and accuracy of motion estimation by local calculation. First, feature points such as corner points or texture salient points are extracted in the initial ROI region. Then, based on the gradient equation constructed based on the brightness constancy assumption and the extracted initial ROI region feature points, the key feature points of the new ROI region are calculated, so as to obtain the new ROI region.
[0149] Further, a rotation instruction is sent to the photographing gimbal according to the movement of the picture region of interest, and an AI camera rotation instruction combining the position change analysis of the region of interest and the physical space coordinate conversion is generated. First, the pixel offset of the ROI center point (x c ,y c ) of the current picture frame and the initial ROI region center point (x0, y0) is calculated, and the calculation formula is: Δx = x c -x0, Δy = y c -y0. Then, the pixel offset is converted into the angle that needs to be adjusted by the photographing gimbal according to the current pitch angle and yaw angle of the camera inside the photographing gimbal, and the calculation formula is as follows:
[0150]
[0151] wherein, △φ is the calculated yaw angle that needs to be adjusted, △θ is the calculated pitch angle that needs to be adjusted, s is the pixel size of the AI camera, f is the focal length of the AI camera, and φ is the initial yaw angle of the shooting gimbal at present.
[0152] In this embodiment, the multi-target tracking algorithm combines key target trajectory prediction and motion compensation organically, ensuring tracking continuity in complex scenes.
[0153] Optionally, the key target is intelligently tracked and photographed according to the first physical coordinates and the multi-target tracking algorithm, comprising:
[0154] The movement of the key target is predicted according to the first physical coordinates by the multi-target tracking algorithm, and the second physical coordinates of the key target are obtained.
[0155] The pitch angle and the yaw angle of the shooting gimbal are automatically adjusted according to the second physical coordinates, and the second picture containing the key target is continuously obtained.
[0156] The multi-target tracking algorithm (Simple Online and Realtime Tracking, SORT) is a high-efficiency multi-target tracking algorithm based on motion modeling and data association. Its core idea is to realize continuous tracking of target identity across frames by fusing target motion prediction in the time dimension and detection box position information in the spatial dimension. First, the position and bounding box of the key target detected by the target detection model in each frame are obtained. Then, the Kalman filter is used to predict the bounding box position of the key target in the next frame picture according to its historical trajectory (i.e. the first physical coordinates). Finally, the Hungarian algorithm is introduced to establish the association relationship between the picture frame containing the key target at present and the picture frame containing the key target previously, and the key target in the current frame and the key target in the previous frame are matched by minimizing the association cost, so that the bounding box of each key target is associated with at most one predicted bounding box, and thus the second physical coordinates of the key target are obtained.
[0157] In one possible implementation, the first physical coordinates are (x1, y1), the second physical coordinates predicted by the multi-target tracking algorithm are (x2, y2), and the yaw angle △φ that the shooting gimbal needs to adjust is:
[0158]
[0159] wherein, D is the target horizontal parameter. The pitch angle △θ that the shooting gimbal needs to adjust is:
[0160]
[0161] In addition, the rotation speed of the shooting gimbal can be controlled by the PID controller, and a smooth rotation speed instruction is generated according to the pitch angle difference or the yaw angle difference to avoid shaking of the shooting gimbal. The calculation formula of the smooth rotation speed u(t) is as follows:
[0162]
[0163] wherein e(t) is the pitch angle difference or the yaw angle difference, K p is a hyperparameter with a set value of 0.8, K i is a hyperparameter with a set value of 0.2, and K d is a hyperparameter with a set value of 0.1.
[0164] In this embodiment, through the deep cooperation of the multi-target tracking algorithm and the shooting gimbal control in the shooting gimbal, high-robustness intelligent tracking in a complex event shooting scene is realized.
[0165] Optionally, the method further comprises:
[0166] The first picture is subjected to face feature extraction by the face recognition and face skeleton point detection model and the face feature extraction model, and first face feature data is obtained;
[0167] The first picture is subjected to human feature extraction by the human feature extraction model, and first human feature data is obtained;
[0168] The second picture is sequentially input into the human feature extraction model for feature extraction, second human feature data is obtained, third human feature data similar to the first human feature data is screened from the second human feature data, and a third picture corresponding to the third human feature data is subjected to intelligent target tracking and shooting.
[0169] The third picture is subjected to face feature extraction by the face recognition and face skeleton point detection model and the face feature extraction model, second face feature data is obtained, and the third picture is subjected to secondary confirmation according to the first face feature data and the second face feature data.
[0170] Figure 2 A flowchart of intelligent tracking and shooting key target confirmation according to the embodiments of the present application is shown in Figure 2 .
[0171] The face recognition and face skeleton point detection model adopts a lightweight model RetinaFace-MobileNet to quickly locate the face boundary box of the key target in the captured picture, and then based on the MediaPipe Face Mesh model, 468 face key points (including but not limited to eyelids, corners of the mouth, and the tip of the nose) are extracted to construct the face topological structure of the key target. The face feature extraction model includes but is not limited to arcface, OpenFace, and Dlib, which are used to extract the face feature data in the picture. The human feature extraction model includes but is not limited to CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), and LSTM (Long Short-Term Memory), which are used to extract the human feature data in the picture. In the present application, when the camera holder is intelligently tracking and capturing the key target, the similarity of the human feature data is compared to confirm the correctness of the tracking of the key target. At the same time, the camera holder synchronizes the captured picture of the key target to the mobile terminal, and the mobile terminal compares the similarity of the face feature data to confirm the correctness of the tracking and capturing of the key target, thereby realizing the matching of the key target in the intelligent terminal and the camera holder.
[0172] In the present embodiment, the combination of face and human dual biometric features improves the accuracy of key target locking and tracking, and avoids interference such as shielding and dressing. In addition, the human feature data preliminary screening and the face feature data secondary matching reduce the false tracking rate of the capturing.
[0173] Optionally, the face feature extraction model is used to extract the face feature data of the first picture through face recognition and face skeleton point detection model, and the first face feature data is obtained, including:
[0174] The face recognition and face skeleton point detection model are used to locate and align the face of the first picture, and the first face data and the first face skeleton point data are obtained.
[0175] The first face data and the first face skeleton point data are input into the face feature extraction model to obtain the first face feature data.
[0176] In a possible implementation, if the face recognition and face skeleton point detection model cannot detect the first face data and the first face skeleton point data in the first picture, the detection is continued in the subsequent frames of the first picture until the first face data and the first face skeleton point data of the key target are obtained, and the first face data and the first face skeleton point data are saved in the mobile terminal.
[0177] In this embodiment, geometric alignment is achieved through face recognition and face skeleton point detection model, which eliminates the interference of posture difference on face feature extraction, and improves the discriminability of face feature data extracted by the face feature extraction model.
[0178] Optionally, the secondary confirmation of the third picture according to the first face feature data and the second face feature data comprises:
[0179] The first face feature data and the second face feature data are subjected to similarity matching.
[0180] If the first face feature data and the second face feature data are similar, the intelligent target tracking shooting is continued.
[0181] If the first face feature data and the second face feature data are not similar, the third picture is re-screened.
[0182] For example, the first face feature data is [0.12, 0.34, 0.56,... 0.78], the second face feature data is [0.13, 0.35, 0.57,... 0.79], the similarity matching value between the first face feature data and the second face feature data is calculated by using the Euclidean distance, and the similarity threshold of the face feature data is set to 0.5. If the first face feature data and the second face feature data are not similar, the tracking and shooting of the AI camera key target fail, and the human body feature data needs to be compared again in the shooting picture to re-screen the third picture.
[0183] In this embodiment, the precise confirmation of the key target is realized through the similarity matching of the face feature data, and the continuity of the shooting and tracking under the complex event scene is ensured.
[0184] Optionally, the method further comprises:
[0185] When the key target is subjected to intelligent target tracking shooting, the third picture is randomly extracted, the third picture is input into the human body feature extraction model for feature extraction, and the fourth human body feature data is obtained.
[0186] The fourth human body feature data is subjected to similarity matching with the target feature queue.
[0187] If the fourth human body feature data is similar to the target feature queue, the intelligent target tracking shooting is continued.
[0188] If the fourth human body feature data is not similar to the target feature queue, the intelligent target tracking shooting fails, the intelligent target tracking shooting is paused, and the third picture is reselected.
[0189] The target feature queue is a feature sequence set of all third human body feature data collected historically.
[0190] For example, the fourth human feature data is [0.1, 0.6, 0.8,..., 0.9], the target feature queue is [[0.1, 0.3, 0.5,..., 0.2],..., [0.2, 0.4, 0.7,..., 0.3]], the similarity values between the fourth human feature data and each feature data in the target feature queue are calculated in turn by using the Euclidean distance, and then all the similarity values are averaged to obtain the similarity value 0.3 between the fourth human feature data and the target feature queue. Therefore, the fourth human feature data is not similar to the target feature queue, the tracking and shooting of the AI camera key target fails, and the human feature data needs to be compared again in the shooting screen to reselect the third screen.
[0191] In the embodiment, the fourth human feature data is matched with the target feature queue by random sampling, so that the AI camera can continuously track the correct key target.
[0192] Optionally, the shooting gimbal is in communication connection with the mobile terminal, and the method further comprises:
[0193] obtaining an operation instruction of the user on the mobile terminal;
[0194] adjusting the pitch angle and the yaw angle of the shooting gimbal according to the operation instruction, and zooming the screen selected in the operation instruction in the mobile terminal.
[0195] The operation instruction of the user includes but is not limited to text input, voice input and gesture control. If the user selects a specific point in the shooting screen to zoom in, the shooting gimbal automatically adjusts the focal length of the shooting. If the user rotates the screen by left and right sliding in the display screen, the shooting gimbal automatically adjusts the pitch angle and the yaw angle at a constant speed to provide the user with the required screen.
[0196] In the embodiment, the pitch angle and the yaw angle of the shooting gimbal are adjusted according to the operation instruction of the user, so that real-time user interaction is realized and the user experience is improved.
[0197] Figure 3 The structure schematic diagram of the interaction between the intelligent terminal and the shooting gimbal according to the embodiment of the application is shown in Figure 3
[0198] Real-time communication between the mobile terminal and the shooting gimbal, and then the real-time intelligent tracking and shooting of the key target is realized. The mobile terminal includes but is not limited to a smart phone and a camera. The shooting gimbal integrates an NPU chip, a display screen and an AI camera. In the NPU chip, a variety of machine learning models and algorithms such as a target detection model, a face recognition and face skeleton point detection model, a face feature extraction model, a human feature extraction model, a multi-target tracking algorithm and a position coordinate conversion algorithm can be integrated, which can realize real-time scene analysis and avoid the burden and delay of cloud processing. The AI camera adopts a wide-angle lens and an AISoC (AI System on Chip) chip, which can be used for intelligent tracking and shooting of events, and can also be used for collecting and integrating image data captured by the mobile terminal or another camera in the shooting gimbal. The display screen can directly display the execution result captured by the shooting gimbal after receiving the driving instruction, and it is co-directional with the AI camera, so that the user can directly control the shooting gimbal through the display screen and then interact.
[0199] In a possible implementation, in order to enable the shooting gimbal to directly interact with the user and correctly identify the operation instruction of the user, the AI camera captures the specified action of the user as the start of the operation instruction, so as to prevent accidental interruption or accidental start of shooting due to misrecognition of the action or gesture of the user in the normal shooting tracking scene. For example, an action instruction and a specified completion time are displayed in the display screen of the shooting gimbal, the action instruction includes but is not limited to opening the palm, and the AI camera identifies that the user completes the action instruction in the display screen within the specified time, and then starts to identify the operation instruction of the user.
[0200] According to the embodiments of the present disclosure, the following technical effects are achieved:
[0201] 1) Real-time automatic intelligent tracking control of the AI camera in the shooting gimbal is realized, and the tracking and shooting efficiency of the key target is significantly improved.
[0202] 2) Multi-dimensional perception data and various machine learning algorithms are fused, which ensures the accuracy of key target positioning and the high resolution of event shooting.
[0203] 3) A dynamic feedback adjustment mechanism is established, and through multiple authentications, the intelligent tracking and shooting of the key target by the AI camera is ensured to be effective.
[0204] Embodiment Two
[0205] This embodiment is an intelligent tracking and shooting gimbal corresponding to the intelligent tracking and shooting method described in Embodiment One, like Figure 4 The block diagram of the intelligent tracking and shooting gimbal device of this embodiment is shown, Figure 10 、 Figure 11A three-dimensional structural schematic diagram of the shooting gimbal is shown; as Figure 4 shown comprising:
[0206] An AI camera is used to acquire a first picture containing a key target, input the first picture into a target detection model, and acquire a key target position; the key target position is converted into a first physical coordinate according to the distortion parameters and the rotation angle of the AI camera; the key target is intelligently tracked according to the first physical coordinate and a multi-target tracking algorithm; a rotation instruction is sent to the driving mechanism;
[0207] A clamping structure is used to clamp and fix the camera device.
[0208] A driving mechanism is used to receive the rotation instruction and drive the shooting gimbal to rotate.
[0209] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described modules can refer to the corresponding process in the foregoing method embodiment one, which will not be repeated here.
[0210] For the specific structure of the intelligent tracking shooting gimbal, please refer to Figure 10 Figure 11 The intelligent tracking shooting gimbal provided by the embodiment, the shooting gimbal comprises a gimbal base 1, a gimbal shell 2 having an accommodating space, a tracker 3, and a driving mechanism for driving the tracker 3 to rotate; the tracker and the driving mechanism are both arranged in the accommodating space of the gimbal shell 2; the driving mechanism comprises a first driving mechanism 41 and a second driving mechanism 42, the first driving mechanism 41 is drivingly connected to the gimbal base 1, and realizes 360-degree horizontal rotation of the tracker 3; the second driving mechanism 42 is drivingly connected to the tracker 3, and realizes 90-degree vertical tilting rotation of the tracker 3. Further, a clamping structure 6 for clamping the camera device is fixedly arranged above the tracker 3; a switch button 7 is arranged on the gimbal shell 2.
[0211] The use principle of the shooting holder is as follows: first, the switch button 7 on the holder shell 2 is started to press, and the tracker 3 of the shooting holder is waited to complete the reset self-checking; then, the clamping structure 6 fixedly connected above the tracker 3 clamps the camera device, such as a smart terminal mobile phone, the botgo application APP installed on the camera device is opened and clicked to add the equipment, the application will be automatically connected with the shooting holder, the clamping installation makes the camera of the camera device and the AI camera of the shooting holder face the same direction, under the action of the tracking information obtained by the AI camera of the shooting holder, the first driving mechanism 41 combined with the second driving mechanism 42 drives the AI camera of the shooting holder to realize the full-range tracking of horizontal 360-degree rotation and vertical 90-degree tilting rotation; when the AI camera tracks the movement, the camera of the camera device can face different directions, that is, the shooting angle of the camera of the camera device is changed, so that the camera of the camera device is always aimed at the changed shooting target, the accurate tracking function of the shooting holder is ensured, and the beneficial effects of simple design structure, stable operation and convenience are achieved.
[0212] When the camera of the camera device changes the shooting angle following the movement of the AI camera, the situation that the moving shooting target is located outside the collection range of the camera can be effectively prevented, the shooting angle of the camera is changed following the movement of the shooting target, the shooting target is always located within the collection range of the camera of the camera device, and the camera of the camera device is used to shoot the shooting target to collect static images or dynamic images. For example, when the shooting target is a dancing athlete, the smart phone can effectively collect images of the dancing athlete in motion. For another example, when the shooting target is a yoga athlete, the smart phone can effectively collect images of the yoga athlete in motion. For another example, when the shooting target is a football player, the smart phone can effectively collect dynamic images of the football and the player in motion. Therefore, the camera device is fixedly arranged on the tracker 3 of the shooting holder, the movement of the tracker 3 is controlled, the shooting angle of the camera of the camera device is changed, the dynamic shooting target is ensured to be located within the collection range of the camera, and the camera of the camera device is used to shoot the shooting target to realize effective and accurate tracking of the shooting target and image collection.
[0213] By using the technical scheme of the embodiment, the tracker 3 and the driving mechanism for driving the tracker 3 to rotate are arranged in the accommodation space of the holder shell, the horizontal 360-degree rotation of the tracker is realized by driving the first driving mechanism 41 connected to the holder base 1, the vertical 90-degree tilting rotation of the tracker is realized by driving the second driving mechanism 42 connected to the tracker 3, the effective and accurate full-range tracking of the shooting holder is realized, and the shooting holder has the advantages of simple structure, convenient use, high control stability and high accuracy.
[0214] As a preferred implementation manner, as shown in Figure 11 , Figure 12 , the tracker 3 of the embodiment comprises an AI camera 31 and a tracking housing 32 for fixedly mounting the AI camera 31, one side of the tracking housing 32 is provided with an arc-shaped rack 321, the second driving mechanism 42 is provided with a driving gear matched with the arc-shaped rack 321, the driving motor of the second driving mechanism 42 drives the driving gear to rotate, thereby driving the arc-shaped rack 321 to move, and then realizing the accurate and effective rotation of the tracker in the vertical direction.
[0215] It should be noted that the AI camera 31 of the embodiment is the tracking sensor in the application number 202423143844X.
[0216] Further, as shown in Figure 11 , the middle part of the holder base 1 is provided with a bearing seat 411, and the driver of the first driving mechanism 41 is rotationally fixedly connected to the bearing seat 411. The driver is a driving motor, and the output end of the driving motor is fixedly connected to the bearing seat 411. According to the role of the AI camera acquiring tracking target information, the driving motor of the first driving mechanism 41 is controlled to rotate, thereby realizing the accurate and effective rotation of the AI camera in the horizontal direction.
[0217] As a preferred implementation manner, as shown in Figure 11 , a control mainboard 9 is arranged between the driver of the first driving mechanism 41 and the bearing seat 411, the control mainboard 9 is fixedly connected to the driver of the first driving mechanism 41, and the control mainboard 9 and the driver of the first driving mechanism 41 are both fixedly connected to the holder housing 2, having the advantages of reasonable and compact structure design.
[0218] As a preferred implementation manner, as shown in Figure 11 , a support assembly is further arranged in the accommodating space of the holder housing, the support assembly is fixedly connected to the holder housing, the support assembly comprises a main support 51 and an auxiliary support 52, the auxiliary support 52 is arranged on both sides of the tracker, a bracket bearing is arranged on the auxiliary support 52, protruding columns 322 are arranged on both sides of the tracking housing 32 of the tracker, the protruding columns 322 are rotationally fixedly connected to the bracket bearings of the auxiliary support 52, and the driving motor of the second driving mechanism 42 is fixed to the auxiliary support 52, thereby ensuring the accurate and effective rotation of the tracker in the vertical direction.
[0219] Further, as shown in Figure 11 , the tracker 3 and the second driving mechanism 42 are arranged on one side of the main support 51, and a power supply 8 is arranged on the other side of the main support 51 to realize the power supply of the shooting holder.
[0220] Embodiment three
[0221] This embodiment is based on embodiment two, as a preferred embodiment, the intelligent tracking and shooting gimbal of this embodiment is also suitable for intelligent tracking of the jersey number, for specific tracking and shooting method, please refer to Figure 5 The flowchart of the intelligent tracking and shooting method based on the jersey number is shown in the figure, which includes:
[0222] Obtain a first human body picture containing a key target.
[0223] Among them, the first human body picture is a front and / or back photo containing the jersey number of the key target, and the mobile terminal includes but is not limited to mobile phone, tablet and camera. For the acquisition method of the first human body picture, the present application does not make any limitation, which can be actively uploaded by the user, or the user can specify the local picture by point selection, frame selection and other ways in real-time shooting picture.
[0224] Input the first human body picture into the target detection model, identify the number area in the first human body picture and cut it, and obtain the first number picture.
[0225] Among them, the target detection model is used to identify and locate the position information of the number area in the first human body picture from the image. In one of the embodiments, the target detection model includes an input layer, a Backbone network layer, a Neck feature fusion layer and a prediction layer. Among them, the input layer is used for pre-processing the input image to meet the needs of model training and prediction. The Backbone network layer is used for feature extraction, which uses the depth optimized C2f module as the basic unit to reduce the network size while improving the performance. The Neck feature fusion layer is used to fuse the feature maps from different stages of the Backbone network layer to enhance the feature representation ability. The prediction layer includes SPPF(Spatial Pyramid Pooling Fast) module, PAA(Probabilistic Anchor Assignment) module and PAN(Path Aggregation Network) module, Head module, SPPF module is used to splice the feature maps of different scales together to improve the detection ability of the model to the image target of different sizes, PAA module is used to intelligently assign anchor box to optimize the selection of positive and negative samples and improve the training effect of the model, PAN module is used to aggregate the features of different levels, and the expression ability of the feature map is enhanced from bottom to top and from top to bottom. The Head module is used for the final target detection prediction, and the position information of the final number area in the first human body picture is output to accurately cut the number area.
[0226] The first number picture is input into the number classification model to obtain first shirt number data.
[0227] The number classification model is used to identify and classify the number in the first number picture.
[0228] The number recognition function is usually completed by an OCR (Optical Character Recognition) technology, but the OCR technology has poor effect in a dynamic picture, and cannot recognize accurate numbers when the shirt is wrinkled. Considering that the text content involved in the shirt number recognition scenario may only include the numbers 0-9, using the OCR technology wastes computing power and may further reduce the accuracy by recognizing results other than numbers. Therefore, in this embodiment, the shirt number recognition process is divided into two steps, S102 and S103, to realize number region recognition and number classification, which can improve the accuracy of shirt number recognition.
[0229] In this application, the specific selection of the number region recognition and number classification model is not limited: the YOLOv8 (You Only Look Once version 8) model can be used to integrate the number region recognition and number classification; or a combination of a recognition model and a classification model can be used, such as a convolutional neural network, ViT (Vision Transformer), etc. for the recognition model, and MobileNet, EfficientNet-Lite, etc. large model or traditional vector machine, random forest, etc. traditional machine learning model for the classification model.
[0230] The first shirt number data is matched with a key target in a real-time shooting picture and tracked and shot.
[0231] Figure 6 The flowchart of the intelligent tracking and shooting method based on the shirt number according to the embodiments of the application is shown in FIG. 1. Figure 6
[0232] First, a first human body picture containing a key target is obtained, and a number region in the first human body picture is identified by a target detection model. If there is no number region in the first human body picture, the shirt number recognition is directly ended, the user is prompted that the recognition fails, and the user is required to upload or specify the first human body picture again. If there is a number region in the first human body picture, the number region is cropped and then input into the number classification model to obtain the final first shirt number data as the basis for determining the key target.
[0233] To explain more intuitively, Figure 7 The schematic diagram of the shirt number recognition according to the embodiments of the application is shown in FIG. 2.Figure 7 As shown:
[0234] The two-stage image recognition algorithm is adopted to extract the shirt number data. First, the number region in the picture is recognized and cropped through the target classification model, and then the number in the number region is recognized by using the number classification model, and the number "17" in the number region is output, which provides a basis for subsequent tracking shooting of the player with the number 17.
[0235] In this embodiment, the key target consistent with the first shirt number data is found and matched in the real-time shooting picture, that is, the key target, and the tracking shooting is performed, so as to realize the intelligent tracking of the key target and improve the user experience.
[0236] Optionally, the key target is matched and tracked in the real-time shooting picture according to the first shirt number data, and the tracking shooting comprises:
[0237] The continuous real-time shooting pictures are input into the target detection model, the number regions in the real-time shooting pictures are recognized and cut in sequence, and the second number picture queue is obtained;
[0238] The second number picture in the second number picture queue is classified by the number classification model in sequence, and the second shirt number data queue is obtained;
[0239] The second shirt number data matching the first shirt number data is searched in the second shirt number data queue, if the second shirt number data matching the first shirt number data exists, the first human body target corresponding to the second shirt number data is tracked and shot by the target tracking algorithm, if the second shirt number data matching the first shirt number data does not exist, the identification and matching are continued.
[0240] For example, the first shirt number data is "7", and the second shirt number data queue of a real-time shooting picture is ["1", "3", "5", "7"] (it is explained that there are four players in the real-time shooting picture at this time, and the shirt numbers of the players are 1, 3, 5 and 7). "7" is matched successfully in the foregoing queue, therefore, the player with the number "7" in the real-time shooting picture is determined as the key target, and the human body target corresponding to the number "7" in the real-time shooting picture is tracked and shot by the target tracking algorithm.
[0241] Further, in the tracking and shooting process of the first human target corresponding to the second jersey number data, the human feature extraction model can be used to extract the human features of the first human target to obtain human feature data of the first human target. Then, the real-time shooting picture is input into the human feature extraction model for feature extraction to obtain human feature data of the real-time shooting picture. The human target similar to the first human target in the human feature data of the real-time shooting picture is selected for tracking and shooting. The human feature extraction model includes but is not limited to CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network) and LSTM (Long Short-Term Memory), which is used to extract the human feature data in the picture.
[0242] Preferably, since the sizes of the plurality of second number pictures may have a large difference, in order to help the model converge, the second number picture can be adjusted to a preset size, such as 64*64 pixels, before being input into the number classification model. In addition, in the subsequent tracking and shooting process, in order to solve the problem of unclear picture caused by long distance shooting, the shooting picture can be split and enlarged before being input into the target detection model, so as to adjust the resolution of the shooting picture and then input it into the target detection model and the number classification model for recognition. The recognition result is spliced according to the original position, so as to solve the problem of unclear number in the human picture in long distance shooting.
[0243] In this embodiment, real-time number recognition is performed in the real-time shooting picture to find the first human target with the same jersey number as the first jersey number data, which ensures the accuracy of number recognition and provides a basis for subsequent tracking and shooting.
[0244] Optionally, the first human target in the second number picture is tracked and shot by using a target tracking algorithm, including:
[0245] The movement of the first human target is predicted by using the target tracking algorithm to obtain the next predicted position of the first human target.
[0246] The pitch angle and the yaw angle of the shooting holder are automatically adjusted according to the next predicted position to continuously obtain the second human picture containing the first human target.
[0247] The specific selection of the target tracking algorithm is not specifically limited in the present application. Generally, mature target tracking algorithms such as MHT (Multiple Hypothesis Tracking) algorithm, MOT (Multiple Object Tracking) algorithm, SORT (Simple Online and Realtime Tracking) algorithm, etc. can all achieve the technical solution of the present application. In one possible implementation, the SORT algorithm is used to realize target tracking. SORT is a high-efficiency target tracking algorithm based on motion modeling and data association, and its core idea is to realize continuous tracking of target identity across frames by fusing target motion prediction in the time dimension and bounding box position information in the spatial dimension. First, the Kalman filter is used to predict the bounding box position of the first human body target in the next frame according to the historical trajectory of the first human body target. Then, the Hungarian algorithm is introduced to establish the association relationship between the current frame containing the first human body target and the previous frame containing the first human body target, and the first human body target in the current frame and the first human body target in the previous frame are matched by minimizing the association cost, so as to ensure that the bounding box of each first human body target is associated with at most one predicted bounding box, and thus the next predicted position of the first human body target is obtained.
[0248] In one possible implementation, the initial position of the first human body target is (x1, y1), the next predicted position of the first human body target predicted by the target tracking algorithm is (x2, y2), and the yaw angle Δφ that needs to be adjusted by the shooting gimbal is:
[0249]
[0250] where D is the target horizontal parameter. The pitch angle Δθ that needs to be adjusted by the shooting gimbal is:
[0251]
[0252] In addition, the rotation speed of the shooting gimbal can be controlled by a PID controller, and a smooth rotation speed command is generated according to the pitch angle difference or the yaw angle difference to avoid shaking of the shooting gimbal. The calculation formula of the smooth rotation speed u(t) is as follows:
[0253]
[0254] where e(t) is the pitch angle difference or the yaw angle difference, K p is a hyperparameter with a set value of 0.8, K i is a hyperparameter with a set value of 0.2, and K d is a hyperparameter with a set value of 0.1.
[0255] In this embodiment, the tracking and shooting of the first human target is realized through a target tracking algorithm, thereby improving the robustness in a complex event shooting scene.
[0256] Optionally, the method further comprises:
[0257] The continuously acquired second human body pictures are sequentially identified and classified by the target detection model and the number classification model to obtain a third shirt number data queue;
[0258] The third shirt number data in the third shirt number data queue and the second shirt number data are sequentially matched, if the third shirt number data in the third shirt number data queue and the second shirt number data are consistent, the tracking and shooting is continued, if there is third shirt number data in the third shirt number data queue that is inconsistent with the second shirt number data, the tracking and shooting fails, and the key target is matched in the real-time shooting picture for tracking and shooting again.
[0259] For example, the second shirt number data is "7", and the third shirt number data queue obtained by identifying the continuously acquired second human body pictures through the target detection model and the number classification model is ["7", "7", "7", "5"], in the tracking and shooting process, there is third shirt number data in the third shirt number data queue that is inconsistent with the second shirt number data, the first human target is lost, the tracking fails, and the key target player with shirt number "7" needs to be found and matched again as the first human target in the real-time shooting picture for tracking and shooting again, so as to avoid continuing to track and shoot the player with shirt number "5" as the player with shirt number "7".
[0260] In this embodiment, the third shirt number data is matched with the second shirt number data, which effectively prevents the tracking and shooting target from being replaced by mistake in a complex scene, is a real-time error correction mechanism, and improves the accuracy and robustness of tracking and shooting.
[0261] Optionally, the continuous real-time shooting pictures are input into the target detection model, the number regions in the real-time shooting pictures are sequentially identified and cut, and a second number picture queue is obtained, and the method further comprises:
[0262] When the real-time shooting picture is a long-distance shooting, the real-time shooting picture is divided into picture blocks, and each picture block is enlarged by a preset magnification;
[0263] Each enlarged picture block is identified and cut by the target detection model to obtain the second number picture queue.
[0264] The preset magnification includes but is not limited to 1 times, 2 times, 3 times and 4 times.
[0265] In this embodiment, the problem of small numbers in the human body picture in long-distance shooting is solved by cutting and magnifying.
[0266] Optionally, the method further comprises:
[0267] If the fourth jersey number data on the second human body target is consistent with the second jersey number data in the real-time shooting picture detected by the target detection model and the number classification model during the tracking shooting process, the tracking target of the tracking shooting is transferred to the second human body target.
[0268] For example, the second jersey number data is "7", and the fourth jersey number data on the second human body target is also "7" in the real-time shooting picture detected by the target detection model and the number classification model during the tracking shooting process. In this case, the original tracking target is lost, resulting in tracking failure. Therefore, the tracking target needs to be transferred to the second human body target, and the tracking shooting is restarted, so as to avoid continuing to track the wrong tracking target as the "7" player.
[0269] In this embodiment, when the fourth jersey number data on the second human body target is consistent with the second jersey number data, the tracking target is transferred in time to ensure that the human body target corresponding to the number is continuously tracked, instead of losing the target or incorrectly tracking an unrelated person.
[0270] Optionally, the method further comprises:
[0271] The first human body picture and the first jersey number data are supplemented to the training data of the target detection model and the number classification model on the server side for model iteration.
[0272] Figure 8 The flowchart of the offline training of the server-side model according to the embodiments of the present application is shown in FIG. 1. Figure 8
[0273] The offline training of the real-time inference model (i.e., the target detection model and the number classification model) is deployed on the server side, and the training data includes but is not limited to basketball games, football games, volleyball games, and rugby games. When identifying the jersey number, the pre-trained real-time inference model (i.e., the target detection model and the number classification model) is pulled to the mobile terminal for running, so as to realize efficient real-time identification of the jersey number. In this process, the first human body picture and the first jersey number data after the jersey number identification are collected and returned to the server side as new training data to be supplemented to the offline training set of the model. Through this data return mechanism, the model can continuously learn new scene features and identification modes, thereby being iteratively optimized.
[0274] In this embodiment, by supplementing the first human body picture and the first jersey number data into the offline data on the server side, the generalization ability and recognition accuracy of the model are continuously improved, and the model performance is optimized.
[0275] According to the embodiments of the present application, the following technical effects are achieved:
[0276] 1) The real-time accurate identification of the jersey number is realized through the target detection model and the number classification model, which provides a basis for subsequent tracking and shooting of the game.
[0277] 2) In the case of long-distance shooting, the problem of inaccurate identification of small jersey numbers in long-distance shooting is effectively solved through the methods of segmentation, magnification and parallel inference identification.
[0278] 3) After identifying the jersey number, the identification resources of online inference are used as the data source for model iteration to construct an efficient data backflow mechanism.
[0279] Embodiment Four
[0280] This embodiment is a device embodiment corresponding to the method for intelligent tracking and shooting based on the jersey number in Embodiment Three, which further illustrates the scheme described in the present application.
[0281] Figure 9 A block diagram of an intelligent tracking and shooting platform based on the jersey number according to an embodiment of the present application is shown, as shown in Figure 9 includes:
[0282] The acquisition module 501 is configured to acquire a first human body picture containing a key target.
[0283] The identification module 502 is configured to input the first human body picture into a target detection model, identify the number area in the first human body picture and perform cutting, and acquire a first number picture.
[0284] The classification module 503 is configured to input the first number picture into a number classification model and acquire first jersey number data.
[0285] The tracking and shooting module 504 is configured to match the key target in a real-time shooting picture according to the first jersey number data and perform tracking and shooting.
[0286] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described modules can refer to the corresponding process in the foregoing method embodiment three, which will not be repeated here.
[0287] Figure 13 A structural schematic diagram of a terminal device or a server suitable for implementing the embodiments of the present application is shown.
[0288] AsFigure 13 As shown, the terminal device or server includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes in accordance with a program stored in a read only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the terminal device or server are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0289] Connected to the I / O interface 605 are an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as necessary. A removable media 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 610 as necessary, so that a computer program read therefrom is installed into the storage section 608 as necessary.
[0290] In particular, according to embodiments of the present application, the above method flow steps can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable media 611. When the computer program is executed by the central processing unit (CPU) 601, the above-described functions defined in the system of the present application are performed.
[0291] It should be noted that the computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium or a combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium can include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer-readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a computer-readable storage medium and / or a computer-readable transmission medium. In the present application, a computer-readable transmission medium can include any computer-readable medium that is not a computer-readable storage medium. In the present application, a computer-readable storage medium can be any computer-readable medium excluding propagating signals per se. The computer-readable storage medium of the present application can be a computer program product.
[0292] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functional processes, and operations that can be implemented in systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0293] The units or modules described in the embodiments of the present application can be implemented in the form of software or in the form of hardware. The units or modules described can also be arranged in a processor. In some cases, the names of the units or modules do not constitute a limitation on the units or modules themselves.
[0294] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the present application.
[0295] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application described in the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above application concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features applied in the present application (but not limited to) having similar functions.
Claims
1. An intelligent tracking shooting gimbal, characterized in that, The pan-tilt head comprises a pan-tilt base, a pan-tilt shell, a tracker and a driving mechanism, wherein the tracker and the driving mechanism are arranged in the pan-tilt shell; The tracker comprises an AI camera, which is configured to acquire a first image containing a key target, input the first image into a target detection model, acquire a key target position, convert the key target position into a first physical coordinate according to a distortion parameter and a rotation angle of the AI camera, intelligently track the key target according to the first physical coordinate and a multi-target tracking algorithm, and send a rotation instruction to the driving mechanism; The driving mechanism is configured to receive the rotation instruction and drive the shooting pan-tilt head to rotate; the driving mechanism comprises a first driving mechanism and a second driving mechanism, the first driving mechanism is drivingly connected to the pan-tilt base to realize 360-degree horizontal rotation of the tracker, and the second driving mechanism is drivingly connected to the tracker to realize 90-degree vertical tilting rotation of the tracker.
2. The intelligent tracking photographic gimbal of claim 1, wherein, The intelligent tracking shooting of the key target according to the first physical coordinate and the multi-target tracking algorithm comprises: predicting movement of the key target according to the first physical coordinate by the multi-target tracking algorithm to acquire a second physical coordinate of the key target; automatically adjusting a pitch angle and a yaw angle of the shooting pan-tilt head according to the second physical coordinate to continuously acquire a second image containing the key target.
3. The intelligent tracking photographic head according to claim 2, characterized in that, It also comprises a face recognition and face skeleton point detection model, a face feature extraction model and a human body feature extraction model; The face recognition and face skeleton point detection model and the face feature extraction model are used to extract face features from the first image to acquire first face feature data; The human body feature extraction model is used to extract human body features from the first image to acquire first human body feature data; The second image is sequentially input into the human body feature extraction model for feature extraction to acquire second human body feature data, third human body feature data similar to the first human body feature data is screened from the second human body feature data, and a third image corresponding to the third human body feature data is intelligently tracked and shot; The face recognition and face skeleton point detection model and the face feature extraction model are used to extract face features from the third image to acquire second face feature data, and the first face feature data and the second face feature data are used to perform secondary confirmation on the third image.
4. The intelligent tracking and photographing gimbal according to claim 3, characterized in that, The face recognition and face skeleton point detection model and the face feature extraction model are used to extract face features from the first image to acquire first face feature data, comprising: The face recognition and face skeleton point detection model is used to locate and align the face in the first image to acquire first face data and first face skeleton point data; The first face data and the first face skeleton point data are input into the face feature extraction model to acquire the first face feature data.
5. The intelligent tracking and photographing gimbal according to claim 3, characterized in that, The first face feature data and the second face feature data are used to perform secondary confirmation on the third image, comprising: The first facial feature data and the second facial feature data are similarity matched; If the first facial feature data and the second facial feature data are similar, the intelligent target tracking shooting is continued; If the first facial feature data and the second facial feature data are not similar, the third picture is re-screened.
6. The intelligent tracking shooting gimbal according to claim 3, wherein, when the intelligent target tracking shooting is performed on the key target, the third picture is randomly extracted, the third picture is input into the human feature extraction model for feature extraction, and fourth human feature data is obtained; the fourth human feature data is similarity matched with a target feature queue; if the fourth human feature data is similar to the target feature queue, the intelligent target tracking shooting is continued; if the fourth human feature data is not similar to the target feature queue, the intelligent target tracking shooting fails, the intelligent target tracking shooting is paused, and the third picture is reselected; the target feature queue is a feature sequence set of all third human feature data collected historically.
7. The intelligent tracking and photographing gimbal according to claim 3, characterized in that, The shooting gimbal is in communication connection with a mobile terminal, and specifically: an operation instruction of a user on the mobile terminal is obtained; an elevation angle and a yaw angle of the shooting gimbal are adjusted according to the operation instruction, and a picture selected in the operation instruction is zoomed in the mobile terminal. 8.The intelligent tracking and photographing gimbal of claim 1, wherein, A clamping structure for clamping a camera is fixedly arranged above the tracker. 9.The intelligent tracking and photographing gimbal of claim 1, wherein, The tracker further comprises a tracking housing for fixedly mounting the AI camera, one side of the tracking housing is provided with an arc-shaped rack, and the second driving mechanism is provided with a driving gear matched with the arc-shaped rack. 10.The intelligent tracking and photographing gimbal according to claim 1, wherein, A bearing seat is arranged in the middle part of the gimbal base, and the driver of the first driving mechanism is rotatably connected to the bearing seat. 11.The intelligent tracking and photographing gimbal of claim 10, wherein, A control mainboard is arranged between the driver of the first driving mechanism and the bearing seat, and the control mainboard and the driver of the first driving mechanism are fixedly connected to the gimbal housing. 12.The intelligent tracking and photographing gimbal according to claim 1, wherein, A support assembly is further arranged in the gimbal housing, the support assembly is fixedly connected to the gimbal housing, and the support assembly comprises a main support and an auxiliary support, and the auxiliary support is arranged on both sides of the tracker.
13. The intelligent tracking gimbal of claim 12, wherein, A bracket bearing is arranged on the auxiliary support, and a protruding column is arranged on both sides of the tracker, and the protruding column is rotatably connected to the bracket bearing.
14. The intelligent tracking, photographing gimbal according to claim 1, wherein, The AI camera is further adapted to intelligently track and shoot a jersey number, and specifically comprises: a first human picture containing a key target is obtained; the first human picture is input into a target detection model, a number area in the first human picture is recognized and cropped, and a first number picture is obtained; the first number picture is input into a number classification model, and first jersey number data is obtained; the key target is matched and tracked in a real-time shooting picture according to the first jersey number data.
15. The intelligent tracking gimbal according to claim 14, wherein, The matching and tracking of the key target in the real-time shooting picture according to the first jersey number data comprises: Input continuous real-time shooting pictures into the target detection model, identify and cut the number area in the real-time shooting pictures in sequence, and obtain a second number picture queue; Classify the second number pictures in the second number picture queue in sequence through the number classification model, and obtain a second shirt number data queue; Find second shirt number data matching the first shirt number data in the second shirt number data queue, if there is second shirt number data matching the first shirt number data, track and shoot the first human body target corresponding to the second shirt number data through the target tracking algorithm, if there is no second shirt number data matching the first shirt number data, continue to identify and match.
16. The intelligent tracking gimbal according to claim 15, wherein, The tracking and shooting of the first human body target corresponding to the second shirt number data through the target tracking algorithm comprises: Predict the movement of the first human body target through the target tracking algorithm, and obtain the next predicted position of the first human body target; According to the next predicted position, automatically adjust the pitch angle and yaw angle of the shooting cloud platform, and continuously obtain the second human body picture containing the first human body target.
17. The intelligent tracking shooting cloud platform according to claim 16, wherein: The second human body picture continuously obtained is identified and classified in sequence through the target detection model and the number classification model, and a third shirt number data queue is obtained; The third shirt number data in the third shirt number data queue and the second shirt number data are matched in sequence, if the third shirt number data in the third shirt number data queue and the second shirt number data are consistent, the tracking and shooting is continued, if there is third shirt number data inconsistent with the second shirt number data in the third shirt number data queue, the tracking and shooting fails, and the key target is matched in the real-time shooting picture again to perform the tracking and shooting.
18. The intelligent tracking gimbal of claim 15, wherein, The input of the continuous real-time shooting picture into the target detection model, the identification and cutting of the number area in the real-time shooting picture in sequence, and the obtaining of the second number picture queue further comprise: When the real-time shooting picture is a long-distance shooting, the real-time shooting picture is divided into picture blocks, and each picture block is enlarged by a preset magnification; Each enlarged picture block is identified and cut by the target detection model to obtain the second number picture queue.
19. The intelligent tracking shooting cloud platform according to claim 17, wherein: If the second human body target is detected in the real-time shooting picture through the target detection model and the number classification model during the tracking and shooting, and the fourth shirt number data on the second human body target is consistent with the second shirt number data, the tracking target of the tracking and shooting is transferred to the second human body target.
20. The intelligent tracking gimbal of claim 14, wherein, Further comprising: The first human body picture and the first shirt number data are supplemented to the training data of the target detection model and the number classification model on the server side for model iteration.
21. An intelligent tracking photographing method, characterized by, Comprise: Acquiring a first picture containing a key target captured by an AI camera in a shooting holder, inputting the first picture into a target detection model to acquire a key target position; Converting the key target position into a first physical coordinate according to a distortion parameter and a rotation angle of the AI camera; Intelligently tracking and shooting the key target according to the first physical coordinate and a multi-target tracking algorithm.
22. The intelligent tracking photography method of claim 21, wherein, The intelligent tracking and shooting of the key target according to the first physical coordinate and the multi-target tracking algorithm comprises: Predicting the movement of the key target according to the first physical coordinate by the multi-target tracking algorithm to acquire a second physical coordinate of the key target; Automatically adjusting the pitch angle and the yaw angle of the shooting holder according to the second physical coordinate to continuously acquire a second picture containing the key target.
23. The intelligent tracking photography method of claim 22, wherein, The method further comprises: Extracting facial features of the first picture by a face recognition and face skeleton point detection model and a face feature extraction model to acquire first facial feature data; Extracting human body features of the first picture by a human body feature extraction model to acquire first human body feature data; Inputting the second picture into the human body feature extraction model in sequence to extract features and acquire second human body feature data, screening third human body feature data similar to the first human body feature data from the second human body feature data, and intelligently tracking and shooting a third picture corresponding to the third human body feature data; Extracting facial features of the third picture by the face recognition and face skeleton point detection model and the face feature extraction model to acquire second facial feature data, and performing secondary confirmation on the third picture according to the first facial feature data and the second facial feature data.
24. The intelligent tracking photography method of claim 23, wherein, The extraction of the facial features of the first picture by the face recognition and face skeleton point detection model and the face feature extraction model to acquire the first facial feature data comprises: Positioning and aligning the face of the first picture by the face recognition and face skeleton point detection model to acquire first face data and first face skeleton point data; Inputting the first face data and the first face skeleton point data into the face feature extraction model to acquire the first facial feature data.
25. The intelligent tracking photography method of claim 23, wherein, The secondary confirmation of the third picture according to the first facial feature data and the second facial feature data comprises: Matching the similarity of the first facial feature data and the second facial feature data; If the first facial feature data and the second facial feature data are similar, the intelligent target tracking and shooting is continued; If the first facial feature data and the second facial feature data are not similar, the third picture is re-screened.
26. The intelligent tracking photography method of claim 23, wherein, The method further comprises: When the intelligent target tracking and shooting of the key target is performed, randomly extracting the third picture, inputting the third picture into the human body feature extraction model to extract features, and acquiring fourth human body feature data; Matching the similarity of the fourth human body feature data and a target feature queue; If the fourth human body feature data is similar to the target feature queue, the intelligent target tracking shooting is continued; If the fourth human body feature data is not similar to the target feature queue, the intelligent target tracking shooting fails, the intelligent target tracking shooting is paused, and the third picture is reselected; The target feature queue is a feature sequence set of all third human body feature data collected historically.
27. The intelligent tracking photography method of claim 23, wherein, The shooting holder is in communication connection with a mobile terminal, and the method further comprises: obtaining an operation instruction of a user on the mobile terminal; adjusting the pitch angle and the yaw angle of the shooting holder according to the operation instruction, and zooming the picture selected in the operation instruction in the mobile terminal.
28. An electronic device, comprising a memory and a processor, a computer program is stored on the memory, characterized in that, The processor executes the computer program to implement the method in any one of claims 21-27.
29. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method in any one of claims 21-27.
Citation Information
Cited By
Cooperative working method and device of user terminal and pan-tilt camera, pan-tilt camera and medium
CN121531236A