Image capture method for image capture system, handheld gimbal, and unmanned aerial vehicle
By introducing a multi-target tracking start command into the shooting system, the posture and parameters of the shooting device are automatically adjusted, solving the problems of image offset and high computational resource consumption in group tracking scenarios, and achieving efficient multi-target tracking shooting effects.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-02
AI Technical Summary
Existing technologies cannot effectively track multiple target objects in group tracking scenarios such as multi-person dances, group photos, and stage performances, resulting in image shift and excessive consumption of computing resources, which affects the shooting experience.
A shooting system is provided that, by acquiring a multi-target tracking start command, identifies and tracks multiple target objects, and automatically adjusts the posture and shooting parameters of the shooting device based on the tracking information, so that the multiple target objects are kept in the captured image. The system includes the introduction of a first tracking mode and a second tracking mode to meet different needs.
It effectively reduces image offset and computational resource consumption, improves the multi-target tracking shooting effect in group tracking scenarios, supports real-time tracking shooting at 30FPS and above, and enhances the level of intelligent control.
Smart Images

Figure CN2024122916_02042026_PF_FP_ABST
Abstract
Description
Shooting method of shooting system, handheld gimbal and unmanned aerial vehicle TECHNICAL FIELD
[0001] The present application relates to the field of shooting control, in particular to a shooting method of a shooting system, a handheld gimbal and an unmanned aerial vehicle. BACKGROUND
[0002] With the continuous development of image processing technology, in order to realize the effect of tracking shooting of the shooting device, it is often necessary to track the target object in the shooting picture. Through tracking the target object, the position of the target object in the shooting picture can be updated in real time, thereby improving the display effect of the shooting picture.
[0003] In related technologies, picture tracking is often based on a single target tracking algorithm, but it cannot meet the group tracking scene involving multiple targets such as multi-person dance, group photograph and stage performance, thereby affecting the shooting experience.
[0004] SUMMARY
[0005] Therefore, the embodiments of the present application provide a shooting method of a shooting system, a handheld gimbal and an unmanned aerial vehicle, aiming to improve the tracking and shooting effect of the shooting system in a group tracking scene.
[0006] The technical solutions of the embodiments of the present application are as follows:
[0007] The embodiments of the present application provide a shooting method of a shooting system, the shooting system comprising a gimbal and a shooting device carried on the gimbal; the shooting method comprising:
[0008] obtaining a tracking start instruction indicating starting multi-target tracking;
[0009] based on the tracking start instruction, identifying multiple target objects in a captured image of the shooting device and tracking the multiple target objects in the captured image;
[0010] based on the tracking information of the multiple target objects, automatically adjusting the posture of the shooting device and / or the shooting parameter of the shooting device, so that the multiple target objects are kept in the captured image.
[0011] The embodiments of the present application also provide a handheld gimbal, comprising:
[0012] a handheld part;
[0013] a gimbal arranged on the handheld part;
[0014] a shooting device carried on the gimbal;
[0015] a memory for storing a computer program;
[0016] The processor is configured to, when the computer program is executed:
[0017] obtain a tracking start instruction indicating starting multi-target tracking;
[0018] identify a plurality of target objects in a captured image of the photographing device based on the tracking start instruction, and track the plurality of target objects in the captured image;
[0019] automatically adjust a posture of the photographing device and / or a photographing parameter of the photographing device based on tracking information of the plurality of target objects, so that the plurality of target objects are kept in the captured image.
[0020] The embodiments of the present application further provide a UAV, comprising:
[0021] a body;
[0022] a power system arranged in the body, the power system being configured to provide power for the UAV;
[0023] a gimbal connected to the body;
[0024] a photographing device carried on the gimbal;
[0025] a memory configured to store a computer program;
[0026] a processor configured to, when the computer program is executed:
[0027] obtain a tracking start instruction indicating starting multi-target tracking;
[0028] identify a plurality of target objects in a captured image of the photographing device based on the tracking start instruction, and track the plurality of target objects in the captured image;
[0029] automatically adjust a posture of the photographing device and / or a photographing parameter of the photographing device based on tracking information of the plurality of target objects, so that the plurality of target objects are kept in the captured image.
[0030] The technical scheme provided in the embodiments of the present application comprises: obtaining a tracking starting instruction indicating starting multi-target tracking; identifying a plurality of target objects in a captured image of a shooting device based on the tracking starting instruction, and tracking the plurality of target objects in the captured image; and automatically adjusting a posture of the shooting device and / or a shooting parameter of the shooting device based on tracking information of the plurality of target objects, so that the plurality of target objects are kept in the captured image. In this way, tracking shooting can be started based on the obtained tracking starting instruction, and in a group tracking scene such as multi-person dancing, group photographing, stage performance, etc., the posture of the shooting device and / or the shooting parameter of the shooting device are automatically adjusted based on the tracking information of the plurality of target objects in the captured image of the shooting device, so that the plurality of target objects are kept in the captured image, thereby effectively improving the tracking shooting effect of multi-target tracking in the group tracking scene. BRIEF DESCRIPTION OF DRAWINGS
[0031] FIG. 1 is a flowchart of a shooting method according to an embodiment of the present application;
[0032] FIG. 2 is a schematic diagram of an interface for generating a tracking starting instruction by a shooting device according to an application embodiment of the present application;
[0033] FIG. 3 is a schematic diagram of a position of a central region of a captured image according to an application embodiment of the present application;
[0034] FIG. 4 is another schematic diagram of an interface for generating a tracking starting instruction by a shooting device according to an application embodiment of the present application;
[0035] FIG. 5 is a schematic diagram of a structure of a handheld gimbal according to an embodiment of the present application;
[0036] FIG. 6 is a schematic diagram of a structure of a drone according to an embodiment of the present application. DETAILED DESCRIPTION
[0037] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing specific embodiments only and is not intended to be limiting of the application.
[0039] In the scenes involving multiple people, such as group dancing, group photo, stage performance, etc., users usually want the picture to contain as many target objects as possible and to be placed in the center of the crowd, rather than just placing a single person in the center of the picture, in order to achieve better composition. To meet this demand, the related technology adopts a group tracking scheme that can analyze the scene. By analyzing the scene of the video, the main characters in the video can be identified, and the head frames of these characters are generated, and then the picture is aligned to the middle position of the head frames. However, this scheme has the following obvious defects in actual use:
[0040] (1) Picture offset problem: Since the group tracking does not set a clear main character object, when some people in the picture move, the camera picture may deviate from the center position following their movement. Since this process is random, it brings a bad experience to the user.
[0041] (2) Delay and redundancy problem: The above scene analysis consumes certain computing resources, resulting in a certain delay in group tracking, which can only run at a speed of 20 FPS (Frames Per Second). This is not friendly enough for some scenes that require real-time performance. At the same time, in shooting scenes such as group photo, multi-person dance performance, etc., scene analysis is often redundant. The photographer usually cleans up the scene before shooting to ensure that all objects in the picture are the main subjects to be shot. In this case, scene analysis not only increases the computing burden, but also does not bring actual added value.
[0042] Based on this, in various embodiments of the present application, a real-time and reliable multi-target tracking shooting method is provided, which can realize group tracking of multiple targets and reduce the interference of tracking delay and random movement of individual objects.
[0043] The embodiment of the present application provides a photographing method of a photographing system, the photographing system comprising a holder and a photographing device carried on the holder. It can be understood that the holder can be used as a support platform of a motion camera, a mobile phone and the like photographing device, and the attitude of the photographing device is flexibly adjusted while the photographing device is stabilized. Common holders include two-axis holders and three-axis holders. The two-axis holder can be a two-axis holder comprising a yaw axis, a pitch axis, a two-axis holder comprising a yaw axis and a roll axis, or a two-axis holder comprising a pitch axis and a roll axis. The three-axis holder can comprise a yaw axis, a roll axis and a pitch axis. The movement of each axis of the holder is driven by a motor. On the one hand, the photographing device can be controlled to rotate around different axes by the motor, so as to adjust the attitude of the photographing device according to the needs of a user; on the other hand, the change of the attitude of the photographing device caused by shaking can be eliminated by reversing the rotation of the motor. The holder can be installed on a handheld part to be used as a handheld photographing device, or can be installed on a carrier such as a drone or a vehicle to realize more rich scene photographing.
[0044] As shown in FIG. 1, the photographing method of the embodiment of the present application comprises:
[0045] In step 101, a tracking start instruction indicating starting multi-target tracking is acquired.
[0046] Here, the photographing system can acquire the tracking start instruction indicating starting multi-target tracking based on human-computer interaction operation, so as to start the multi-target tracking function based on the needs of photographing, and then realize the setting of multi-target tracking photographing in a specific scene.
[0047] In step 102, based on the tracking start instruction, a plurality of target objects in a captured image of the photographing device are identified, and the plurality of target objects in the captured image are tracked.
[0048] Here, the photographing system can identify a plurality of target objects in a captured image of the photographing device based on the tracking start instruction, and track the plurality of target objects in the captured image based on a tracking algorithm. The target objects can be persons, animals, bionic robots, drones, cars and the like, which are not limited in the embodiment of the present application. The photographing device can be a photosensitive imaging device or an electronic device with a photosensitive device, for example, a CCD (Charge Coupled Device, charge coupled device) camera, a mobile phone with a photographing function, a video recording device and the like electronic device.
[0049] In step 103, based on the tracking information of the plurality of target objects, the attitude of the photographing device and / or the photographing parameter of the photographing device is automatically adjusted, so that the plurality of target objects are kept in the captured image.
[0050] Here, the photographing system can automatically adjust the posture of the photographing device and / or photographing parameters of the photographing device based on the tracking information of the plurality of target objects, where the adjustment of the posture of the photographing device can be achieved based on the rotation of each rotation axis of the gimbal, and the photographing parameters of the photographing device can include, but are not limited to, focal length, optical zoom parameter, digital zoom parameter, shutter speed, aperture parameter, and sensitivity, to meet the photographing framing requirement, so that the plurality of target objects are kept in the captured image, and the tracking photographing effect of multi-target tracking in a group tracking scene is improved.
[0051] It can be understood that the method of the embodiments of the present application can start tracking photographing based on the obtained tracking start instruction, thereby eliminating the scene analysis of the photographed picture, saving the calculation resources consumed by scene analysis in a group tracking scene, reducing the photographing delay of group tracking, and effectively improving the tracking photographing effect of multi-target tracking in a group tracking scene. In addition, the introduction of the tracking start instruction can reasonably control whether the tracking photographing function is started based on user demand, has high intelligence, and can better meet the photographing requirements of different photographing scenes.
[0052] Exemplarily, the tracking start instruction has a tracking mode, and the type of the tracking mode includes a first tracking mode and a second tracking mode, where in the first tracking mode, the plurality of target objects at least includes a main character object; and in the second tracking mode, the plurality of target objects does not include a main character object.
[0053] In an embodiment, the tracking start instruction can identify the tracking mode, where the type of the tracking mode includes a first tracking mode and a second tracking mode. The first tracking mode, i.e., the tracking mode with a main character, can specify at least one of the plurality of target objects as a main character object, so that the position of the main character object in the captured image can be given priority to meet the photographing requirement of group tracking that needs to highlight a specific object. The second tracking mode, i.e., the tracking mode without a main character, supports the captured image to be composed according to the distribution of the overall group object, which can meet the photographing requirement of ordinary group tracking.
[0054] It can be understood that through the setting of the tracking mode and the introduction of the first tracking mode and the second tracking mode, the purpose of mode switching according to the photographing requirement of group tracking can be met, and in the first tracking mode, the picture composition of the captured image can be based on the determined main character object, which can effectively avoid the problem that the picture of group tracking is deviated by the movement of individual objects in the group. In addition, compared with scene analysis of the picture, the consumption of calculation resources and the tracking delay can be reduced, and a faster group tracking speed can be met, for example, a speed of more than 30 FPS can be supported for tracking photographing.
[0055] In the first tracking mode, the plurality of target objects are kept in the captured image, including that the main character object is kept in a preset region of the captured image and the plurality of target objects are kept in the captured image.
[0056] It should be noted that in the first tracking mode, in order to meet the requirement that the main character object is highlighted in the captured image, the main character object is preferentially kept in the preset region of the captured image and the plurality of target objects are kept in the captured image, thereby realizing both group tracking and highlighting of the main character object, which is suitable for stage performance or video shooting scenes where the main character is highlighted.
[0057] In an example, the preset region is a central region of the captured image, that is, the main character object is kept in the central region of the captured image, which can effectively improve the picture composition effect of the main character in group tracking.
[0058] In an example, the central region is a central sub-grid region selected after the captured image is divided according to a 3x3 grid. Here, the captured image can be divided into a 3x3 grid according to a three-equal proportion along the width direction and the height direction, and the central grid is selected as the central sub-grid region.
[0059] In the first tracking mode, the overall display position of the plurality of target objects is adjusted based on the display position of the main character object, so that the main character object is kept in the preset region of the captured image and the center of the plurality of target objects is close to or located at the center of the captured image. Here, since the plurality of target objects are adjusted in picture based on the overall display position, the center of the plurality of target objects refers to the center of the overall layout of the plurality of target objects.
[0060] It can be understood that in the first tracking mode, the display position of the main character object is preferentially tracked, and the overall display position of the plurality of target objects is adjusted based on the display position of the main character object, so that the center of the plurality of target objects is close to or at the center of the captured image while the main character object is kept in the preset region of the captured image, that is, the center of the plurality of target objects coincides with the center of the captured image as much as possible. In this way, the tracking and shooting requirements of group tracking and highlighting of the main character can be considered, and the picture of the captured image can be prevented from being deviated due to the movement of individual objects in the group.
[0061] In the first tracking mode, the tracking information includes a first group frame and a main character frame, wherein the first group frame is adjusted based on the position of the main character frame, so that the main character frame is located in the preset region while the center of the first group frame is close to or located at the center of the captured image, the first group frame is determined based on the plurality of target objects, and the main character frame is determined based on the main character object.
[0062] Here, the first group frame can be constructed based on the distribution of the head regions of the plurality of target objects, for example, can be a square frame surrounding the head regions of the plurality of target objects. The main character frame can be constructed based on the position of the head region of the main character object, for example, can be a square frame surrounding the head region of the main character object. Wherein, tracking the plurality of target objects in the captured image can be understood as multi-target tracking of the plurality of target objects in at least two adjacent frames captured by the shooting device, so that the position adjustment of the main character frame is prior to the first group frame, and the center of the first group frame coincides with the center of the captured image as much as possible while the main character frame is located in the preset area of the captured image. Here, the center coincidence can be understood as the centers of the two being the same or within a set deviation accuracy range.
[0063] Exemplarily, the center of the first group frame is determined based on the deviation amount between the main character frame and the preset area, so that the deviation amount between the center of the first group frame and the center of the captured image is minimized.
[0064] It can be understood that in the first tracking mode, the center position of the first group frame can be adjusted based on the deviation amount between the main character frame and the preset area of the captured image, so that the deviation amount between the center of the first group frame and the center of the captured image is minimized while the main character frame is located in the preset area, that is, the center of the first group frame is made to coincide with the center of the captured image as much as possible while ensuring that the main character frame is located in the preset area of the captured image. In this way, the display effect of highlighting the main character can be achieved while optimizing the framing of the captured image and improving the shooting effect of group tracking.
[0065] Exemplarily, in the second tracking mode, the tracking information includes a second group frame, wherein the center of the second group frame coincides with the center of the captured image, and the second group frame is determined based on the plurality of target objects.
[0066] Here, in the second tracking mode, no positioning requirement is set for each target object in the captured image, allowing the framing to be based on the group distribution of the plurality of target objects, which is suitable for group photos or dance scenes. Wherein, the second group frame can be constructed based on the distribution of the head regions of the plurality of target objects, for example, can be a square frame surrounding the head regions of the plurality of target objects. The center of the second group frame coincides with the center of the captured image to achieve tracking and shooting framing for multi-target tracking, which can be understood as the centers of the two being the same or within a set deviation accuracy range.
[0067] Exemplarily, the tracking start instruction indicating the start of multi-target tracking includes:
[0068] generate the tracking start instruction based on a first interaction operation of selecting a region on the image captured by the photographing device.
[0069] In an embodiment, the photographing device has a display screen displaying the captured image, which can be a touch screen, and the user can select a region on the captured image based on a touch operation on the touch screen, and the photographing device can generate the tracking start instruction in response to the touch operation (i.e., the first interaction operation).
[0070] Illustratively, the generating of the tracking start instruction based on the first interaction operation of selecting a region on the image captured by the photographing device includes:
[0071] If it is determined based on the first interaction operation that the selected region corresponds to at least two target objects, the tracking start instruction is generated.
[0072] Here, the photographing device determines based on the first interaction operation that the selected region corresponds to at least two target objects, indicating that the group needs to be tracked, and generates the tracking start instruction.
[0073] Illustratively, the photographing device determines based on the first interaction operation that the selected region corresponds to at least two target objects, generates the tracking start instruction, and by default starts the first tracking mode and selects the main character object, thereby further improving the intelligent level of group tracking photography.
[0074] Illustratively, the selecting of the main character object includes:
[0075] Selecting a target object as the main character object based on the selected region and the positions of the at least two target objects.
[0076] Here, the photographing device can select a target object as the main character object based on the selected region and the positions of the target objects in the captured image. The strategy for selecting the main character object can be pre-set based on the relative relationship between the selected region and the positions of the target objects, thereby realizing automatic selection of the main character object, effectively saving operation steps, and improving the intelligent level of tracking photography.
[0077] Illustratively, the selecting of the main character object based on the selected region and the positions of the at least two target objects includes: selecting a target object closest to the center of the selected region from the at least two target objects as the main character object. In this way, the target object corresponding to the center of the selected region can be selected as the main character object. In other application examples, the target object corresponding to the top-left corner, the top-right corner, the bottom-left corner, or the bottom-right corner of the selected region can be selected as the main character object, which is not limited in the embodiments of the present application.
[0078] Illustratively, the photographing device determines, based on the first interactive operation, that the selected region corresponds to at least two target objects, generates a tracking start instruction, and starts the second tracking mode by default, i.e., enters the group tracking mode without a main character, thereby further improving the intelligent level of group tracking shooting control.
[0079] Illustratively, the photographing device generates the tracking start instruction based on the second interactive operation of voice collection and / or image collection.
[0080] The photographing device generates the tracking start instruction based on the second interactive operation of voice collection and / or image collection.
[0081] In an embodiment, to facilitate remote user control, the photographing device can generate the tracking start instruction based on the second interactive operation of voice collection and / or image collection. In this way, the interactive experience in a group tracking shooting scenario can be further improved, making control simple and fast.
[0082] Illustratively, the photographing device generates the tracking start instruction based on the second interactive operation of voice collection and / or image collection, including:
[0083] The photographing device generates the tracking start instruction based on the second interactive operation determining that at least one of a set voice instruction and a set gesture action is collected.
[0084] In an application example, the set voice instruction can be "start group tracking shooting" or similar voice instructions. The set gesture action can be a preconfigured gesture action for generating a tracking start instruction, for example, when an "O" or "V" or "P" shaped gesture action is made to determine that a target object exists in a captured image, the tracking start instruction is generated. The form of the set voice instruction and the set gesture action is not limited in the embodiments of the present application, and can be personalized according to needs by those skilled in the art.
[0085] Illustratively, the photographing device generates the tracking start instruction based on the second interactive operation, and starts the first tracking mode by default, and selects the main character object, thereby further improving the intelligent level of group tracking shooting control.
[0086] Illustratively, the photographing device generates the tracking start instruction based on the second interactive operation, and starts the first tracking mode by default, and selects the main character object, thereby further improving the intelligent level of group tracking shooting control.
[0087] The photographing device determines the main character object based on an information source of the second interactive operation.
[0088] Here, the photographing device can determine the main character object based on an information source of the collected second operation, thereby realizing automatic selection of the main character object, effectively saving operation steps, and improving the intelligent level of tracking shooting.
[0089] Illustratively, the photographing device generates the tracking start instruction based on the second interactive operation of voice collection and / or image collection, including:
[0090] The main character object is determined based on a target object corresponding to the collected setting voice instruction or setting gesture action.
[0091] For example, in the gesture interaction mode, the main character object can be set as a target object making the setting gesture action or an object closest to the target object making the setting gesture action; in the voice interaction mode, the main character object can be set as a target object issuing the setting voice command or an object closest to the target object issuing the setting voice command.
[0092] Exemplarily, the photographing device generates a tracking start instruction based on the second interaction operation, and starts the second tracking mode by default, i.e., enters the group tracking mode without a main character, thereby further improving the intelligent level of group tracking shooting control.
[0093] Exemplarily, in order to realize intelligent switching of the tracking mode in the photographing process, the method further includes:
[0094] Switching the tracking mode of the tracking start instruction based on a switching instruction for switching the tracking mode in the photographing process.
[0095] Here, the photographing device can generate a switching instruction for switching the tracking mode based on a touch operation of a user on a touch screen, or a key operation on the photographing device, or a touch or key operation on a gimbal base, or a collected gesture or voice interaction of a user remotely, and switch the tracking mode of the tracking start instruction, and the configuration of the switching instruction can be defined based on individual needs, which is not limited in the embodiments of the present application.
[0096] Exemplarily, in order to realize automatic exiting of the current tracking mode in the photographing process, the method further includes:
[0097] Exitting the current tracking mode based on an exiting instruction indicating exiting the tracking mode.
[0098] Here, the photographing device can generate an exiting instruction for exiting the tracking mode based on a touch operation of a user on a touch screen, or a key operation on the photographing device, or a touch or key operation on a gimbal base, or a collected gesture or voice interaction of a user remotely, and the configuration of the exiting instruction can be defined based on individual needs, which is not limited in the embodiments of the present application.
[0099] Exemplarily, the tracking information includes a group frame indicating display positions of the plurality of target objects, and the automatically adjusting the posture of the photographing device and / or the photographing parameter of the photographing device, so that the plurality of target objects are kept in the image captured by the photographing device, includes:
[0100] based on the center position and / or the size of the group box, automatically adjusting a pose of the photographing device and / or a photographing parameter of the photographing device, so that the plurality of target objects are kept in the image captured by the photographing device.
[0101] Here, based on tracking the plurality of target objects in the captured image, the photographing device can generate a group box indicating the display positions of the plurality of target objects, and based on the center position and / or the size of the group box, automatically adjust the pose of the photographing device and / or the photographing parameter, so that the plurality of target objects are kept in the image captured by the photographing device, thereby effectively guaranteeing the picture composition effect of the captured image in the group tracking shooting scene. The tracking of the plurality of target objects in the captured image can be understood as multi-target tracking of the target objects in at least two adjacent frames of captured images shot by the photographing device. The mechanism adopted by the generated group box based on different tracking modes is different. In the first tracking mode, the generated group box preferentially ensures that the main character object is located in the set region of the captured image, and the center of the group box is close to or coincides with the center of the captured image. In the second tracking mode, the group box is constructed based on the distribution of the plurality of target objects, so that the center of the group box coincides with the center of the captured image.
[0102] It can be understood that the photographing system can determine the pose adjustment amount of the photographing device based on the position deviation amount of the center positions of the group box of the previous frame captured image and the group box of the current frame captured image, and realize the pose adjustment of the photographing device by controlling the displacement of the holder, so that the picture effect of the captured image of the photographing device meets the needs of group tracking. The photographing system can also determine the adjustment amount of the focal length based on the size deviation amount of the sizes of the group box of the previous frame captured image and the group box of the current frame captured image, and adjust the focal length of the photographing device, so that the picture effect of the captured image of the photographing device meets the needs of group tracking.
[0103] Exemplarily, the identifying the plurality of target objects in the captured image of the photographing device comprises:
[0104] detecting the body and the head of the plurality of target objects in the captured image of the photographing device, and identifying the plurality of target objects in the captured image of the photographing device.
[0105] Here, by introducing the detection of the body and the head, the plurality of target objects in the captured image of the photographing device can be identified based on the detection results of the body and the head of the target objects, so that the head region of the plurality of target objects in the captured image of the photographing device can be more accurately identified, thereby improving the accuracy of the tracking information of group tracking, effectively improving the multi-target tracking quality in the group tracking scene, and thereby improving the shooting effect of the photographing device in the group tracking scene.
[0106] Exemplarily, the detecting a plurality of target objects in the captured image by the photographing device includes:
[0107] performing target detection on the captured image based on the pre-trained target detector to obtain a first detection result corresponding to the body of the target object and a second detection result corresponding to the head of the target object; wherein a confidence threshold of the first detection result is greater than a confidence threshold of the second detection result;
[0108] matching the obtained first detection result and second detection result based on the positional relationship between the body and the head of the target object, and obtaining the body region parameter and the head region parameter of the identified target object based on the matched first detection result and second detection result.
[0109] It should be noted that the pre-trained target detector supports simultaneous detection of the body and the head of the target object, that is, supports detecting the body and the head of the target object as detection targets. Since head detection is more difficult than body detection, in order to improve the accuracy of detection and avoid missing detection, in the embodiment of the present application, the confidence thresholds of body detection and head detection are set respectively, wherein the confidence threshold of the body is greater than the confidence threshold of the head, so that the recall rate of head detection can be improved; in addition, the positional relationship between the body and the head of the target object is introduced, so that the detected body and head can be matched, so that the accuracy of head detection can be improved on the basis of ensuring high recall rate of head detection, thereby providing guarantee for tracking the head and / or body of the target object.
[0110] Exemplarily, the tracking the plurality of target objects in the captured image includes:
[0111] In the first tracking mode, a single-target tracking algorithm of tracking the body is adopted for a main character object in the plurality of target objects, and a multi-target tracking algorithm of tracking the head is adopted for other target objects except the main character object, to generate a first group box of group tracking;
[0112] In the second tracking mode, a multi-target tracking algorithm of tracking the head is adopted for the plurality of target objects, to generate a second group box of group tracking.
[0113] It should be noted that in the first tracking mode, since there is a main character object, in order to improve the accuracy of tracking the main character object and ensure the main character tracking effect of the captured image in the group tracking scene, the single-target tracking algorithm of tracking the body is adopted for the main character object, and the multi-target tracking algorithm of tracking the head is adopted for other target objects except the main character object, which can reduce the complexity of the tracking algorithm as much as possible on the basis of meeting the tracking precision of the main character object, thereby reducing the consumption of computing resources, and meeting the demand of fast tracking and shooting.
[0114] In the second tracking mode, since there is no main character object, a multi-target tracking algorithm for tracking heads can be used for all target objects, so as to reduce the complexity of the tracking algorithm as much as possible, thereby reducing the consumption of computing resources and meeting the demand for fast tracking and shooting.
[0115] For example, the multi-target tracking algorithm for tracking heads can use a SORT (Simple Online and Realtime Tracking) algorithm. For example, the head targets in adjacent frames can be matched based on the IoU (Intersection over Union) between the head targets in adjacent multi-frame captured images, so as to realize continuous tracking of the head targets. In this way, stable head detection results of multiple target objects can be generated, and the influence of missed or false head targets on the group box can be prevented. Since this algorithm does not need to perform complex feature extraction and model training, it has high computational efficiency and real-time performance.
[0116] For example, the multi-target tracking algorithm for tracking heads can use a SORT (Simple Online and Realtime Tracking) algorithm. For example, the head targets in adjacent frames can be matched based on the IoU (Intersection over Union) between the head targets in adjacent multi-frame captured images, so as to realize continuous tracking of the head targets. In this way, stable head detection results of multiple target objects can be generated, and the influence of missed or false head targets on the group box can be prevented. Since this algorithm does not need to perform complex feature extraction and model training, it has high computational efficiency and real-time performance.
[0117] The single-target tracking algorithm for tracking bodies is used for the main character object to obtain head region parameters of the main character object.
[0118] The multi-target tracking algorithm for tracking heads is used for the other target objects other than the main character object to obtain head region parameters of the other target objects.
[0119] The first group box is generated based on the head region parameters of the main character object and the head region parameters of the other target objects.
[0120] It can be understood that in the first tracking mode, the first group box is generated based on the head region parameters of each target object, that is, based on the head positions of each target object. The head position of the main character object is given priority over the head positions of the other target objects, that is, the center position of the first group box is adjusted based on the movement of the head position of the main character object, so that the head position of the main character object is in a preset region of the captured image, and the center of the first group box is as close as possible to the center of the captured image.
[0121] For example, the first group box is generated based on the head region parameters of the main character object and the head region parameters of the other target objects, including:
[0122] generate an initial group frame and a head frame of the main character object based on the head region parameter of the main character object and the head region parameter of the other target objects, wherein the initial group frame corresponds to the head region of all the target objects, and the head frame of the main character object corresponds to the head region of the main character object;
[0123] determine a first deviation amount based on a center of the initial group frame and a center of the captured image, and determine a second deviation amount based on the head frame of the main character object and a preset region of the captured image;
[0124] correct a deviation amount of the initial group frame based on the first deviation amount and the second deviation amount;
[0125] offset correct the initial group frame based on the deviation amount to obtain the first group frame.
[0126] It can be understood that the initial group frame is adjusted based on the position of the head frame of the main character object, wherein the head frame of the main character object needs to be located in the preset region, so that the main character object is highlighted in the preset region of the captured image. Based on this, the first deviation amount of the center of the initial group frame and the center of the captured image can be determined, and the second deviation amount of the head frame of the main character object and the preset region of the captured image can be determined. Then, based on the first deviation amount and the second deviation amount, the initial group frame is corrected so that the initial group frame coincides with the center of the captured image as much as possible, to obtain the center of the corrected group frame. Then, based on the center, the group frame surrounding all the target objects is constructed to obtain the first group frame in the first tracking mode.
[0127] Exemplarily, the single-target tracking algorithm for tracking the body of the main character object is adopted to obtain the head region parameter of the main character object, including:
[0128] based on the body region parameter of the main character object, a single-target tracking algorithm is adopted to track the body region;
[0129] based on the head region matched with the tracked body region, the head region parameter of the main character object is obtained.
[0130] It can be understood that, since the single-target tracking algorithm for tracking the body is adopted, and the head region of the main character object is determined based on the tracked body region, the tracking accuracy of the main character object can be effectively improved, and thus the tracking effect of the main character object is improved.
[0131] Exemplarily, the single-target tracking algorithm for tracking the body of the main character object is adopted to obtain the head region parameter of the main character object, including:
[0132] For the captured image of the non-key frame, the IoU (intersection over union) is calculated based on the identified body region parameters of the target object and the body region parameters of the main character object of the previous frame, and the body region parameters and the re-identification feature (Re-identification, ReID) used for body region tracking are updated based on the IoU, wherein the re-identification feature represents the image feature of the corresponding body region, that is, the re-identification feature of the main character object is constructed based on the image feature of the body region of the main character object.
[0133] For the captured image of the key frame, the re-identification feature is extracted based on the identified body region parameters of the target object, and is compared with the updated re-identification feature of the main character object of the previous frame, and the body region parameters and the re-identification feature used for body region tracking are updated based on the comparison result.
[0134] Here, in order to improve the tracking efficiency of the single target tracking algorithm, the captured image is distinguished into key frames and non-key frames in the embodiment of the application, wherein for the captured image of the non-key frame, the body region parameters and the re-identification feature used for body region tracking are updated based on the IoU. Exemplarily, for the captured image of the non-key frame, all the detected body regions are matched with the body region of the main character object of the previous captured image through the IoU, if the IoU is greater than a set threshold, the body region corresponding to the IoU greater than the set threshold generates the updated body region parameters, and the image feature of the corresponding body region is taken as the updated re-identification feature. For the captured image of the key frame, the re-identification features of all the detected body regions are extracted, and are compared with the updated re-identification feature of the main character object of the previous captured image, the body region with the most similar re-identification feature is selected to generate the updated body region parameters, and the image feature of the body region is taken as the updated re-identification feature.
[0135] Exemplarily, the shooting system sets the captured image of the shooting device as the captured image of the key frame based on a set frequency. The set frequency can be reasonably set based on the tracking accuracy of the single target tracking algorithm, for example, for the captured image of the shooting device, one frame is set as a key frame every 30 frames, it can be understood that the other frames except the key frames are non-key frames.
[0136] The shooting method of the embodiment of the application will be further described in detail in combination with an application embodiment.
[0137] In the application embodiment, the shooting system includes a gimbal and a shooting device mounted on the gimbal, the shooting device can be a mobile phone with a camera function, the mobile phone is in communication connection with the gimbal, and the posture of the mobile phone can be adjusted by controlling the gimbal. The shooting method of the application embodiment will be described below by taking a person as the object of tracking and shooting.
[0138] In the application embodiment, the tracking start instruction in the shooting method run by the mobile phone has a tracking mode, which specifically includes:
[0139] (1) There is a main role mode (corresponding to the first tracking mode described above): In this mode, the position of the main character is given priority to ensure that the main character is always within the preset area of the captured image, thereby avoiding picture deviation caused by the movement of individual objects in the group. This mode is suitable for scenes where a particular character needs to be highlighted, such as stage performances or video shooting of main characters.
[0140] (2) No main character mode (corresponding to the second tracking mode described above): In this mode, no position requirements are set for all objects in the captured image, allowing the captured image to be composed according to the distribution of the entire group. This mode is suitable for group photos or dance scenes where all characters are important objects.
[0141] It can be understood that the user can switch between the above two modes according to actual needs. Through this real-time and switchable group tracking scheme, the problems of high latency and picture deviation in the related art during group tracking can be effectively solved, and the user's shooting experience in a multi-person scene can be improved.
[0142] The following exemplary describes the generation of a tracking start instruction by the mobile phone based on the user's interactive operation to start the group tracking function.
[0143] As shown in FIG. 2, the mobile phone displays a captured image based on the display screen after opening the shooting software, and the user can frame select multiple heads in the captured image based on touch operation on the display screen to initiate a group tracking request. The mobile phone generates a tracking start instruction in response to the multiple heads in the captured image being framed selected, and by default starts the first tracking mode, taking the head closest to the center of the framed area as the main character. Among them, the single target tracking algorithm for tracking human bodies is used for tracking the main character, and the multi-target tracking algorithm for tracking human heads is used for tracking the remaining characters.
[0144] In the application embodiment, in order to perform single target tracking of human body and multi-target tracking of human head under the condition of ensuring network inference speed, a target detector capable of detecting human body and human head at the same time is trained. Since it is more difficult to detect human head than human body, in order to improve the recall rate of human head, different confidence filtering threshold strategies are adopted for the detection results. Specifically, for the detection result of human head, a lower confidence filtering threshold (for example, 0.1) is set to ensure that as many human heads as possible are detected; for the detection result of human body, a higher confidence filtering threshold (for example, 0.25) is set to ensure the accuracy of human body detection. Subsequently, using the prior knowledge of the position relationship between human head and human body, for example, human head is located above human body, each human head frame and human body frame detected is matched. This matching strategy not only ensures the high recall rate of human head detection, but also ensures the accuracy of human head detection, thereby realizing reliable simultaneous tracking of human body and human head.
[0145] It should be noted that in the application embodiment, the single target tracking algorithm for tracking human body combines ReID feature extraction and target detection technology. For the captured image of a non-key frame, all human body frames detected are matched with the human body frame (i.e. tracking frame) of the main character in the last frame through IoU, if the IoU is greater than 0.45, the tracking frame is updated, and the ReID feature in the tracking frame is extracted as the ReID feature of the main character. For the captured image of a key frame, the ReID features in all detected human body frames are compared with the ReID feature in the human body frame (i.e. tracking frame) of the main character in the last frame, and the human body frame corresponding to the most similar feature is selected as the new tracking frame, and the ReID feature of the main character is updated.
[0146] Exemplarily, the multi-target tracking algorithm for tracking human head adopts the SORT algorithm for human head multi-target tracking based on IoU. The multi-target tracking is mainly to generate stable human head to prevent missing or misdetecting human head.
[0147] In the first tracking mode, the minimum bounding rectangle of all human head frames in the captured image is taken as the initial group frame. In order to prevent the group frame from being deviated by other characters except the main character, in the application embodiment, the human head frame of the main character is limited in a preset area of the captured image.
[0148] Exemplarily, the preset area is the central area of the captured image. As shown in FIG. 3, the central area is the central sub-grid area selected after the captured image is divided into a 3×3 grid. Here, the captured image can be divided into a 3×3 grid by being equally divided along the width direction and the height direction, and the central grid is selected as the central sub-grid area.
[0149] In the application embodiment, the offset of the holder can be calculated based on the center of the group frame and the center of the captured image. The offset of the holder can be simulated first, and the center of the group frame can be moved to the center of the captured image. At this time, the head frame of the main character will also be offset. If the position of the head frame of the main character after the offset exceeds the preset area of the captured image, the head frame of the main character needs to be moved back to the preset area to ensure that the head frame of the main character is always within the preset area. According to the offset of the head frame of the main character after the movement back, the position of the center of the current group frame is deduced. Finally, the center of the group frame is spread outward until the group frame contains all the head frames in the picture, forming a new group frame, that is, the tracking information in the first tracking mode. The mobile phone can control the offset of the holder based on the tracking information, for example, according to the position offset between the center of the group frame of the current frame captured image and the center of the group frame of the last frame captured image, the pose adjustment amount of the holder is controlled, so as to realize the effect of group tracking shooting in the main character mode.
[0150] Exemplarily, as shown in FIG. 4, after the mobile phone generates the tracking starting instruction in response to the multiple head frames being selected on the captured image, the mobile phone can also output a text prompt box and related auxiliary marks, for example, output a text prompt box of “group composition has been started, click or swipe to switch the tracking object”, and display the head frames of the characters on the application interface. The user can switch the tracking object, that is, switch the main character, based on the interactive operation such as clicking or swiping. In addition, the user can also exit or close the group composition (that is, exit the current tracking mode) based on the interactive operation, or switch the tracking mode of the group composition, for example, switch from the default first tracking mode to the second tracking mode.
[0151] Exemplarily, the user can cancel the main character group tracking mode by clicking the cancel button on the main character tracking box, and enter the non-main character tracking mode (corresponding to the second tracking mode described above). The non-main character tracking mode no longer has a main character, so there is no single target tracking of the main character, and only the multi-target tracking algorithm based on the IoU of the head frame is used to realize the tracking of the group head frame.
[0152] It should be noted that in the second tracking mode, the group frame can be constructed based on the distribution of the group head frame, and the center of the group frame can be moved to coincide with the center of the captured image to obtain the tracking information of the multi-target tracking. The mobile phone can control the offset of the holder based on the tracking information, for example, since the center of the group frame in the second tracking mode coincides with the center of the captured image, the pose adjustment amount of the holder can be controlled according to the position offset between the center of the group frame of the current frame captured image and the center of the captured image, so as to realize the effect of group tracking shooting in the non-main character mode described above.
[0153] It can be understood that the shooting method of the application embodiment has the following technical advantages:
[0154] (1) Optimize the framing effect of the captured image based on a group frame
[0155] The application embodiment can automatically generate a group frame containing the entire group in a multi-person scene. By controlling the pose of the gimbal and adjusting the posture of the shooting device through the group frame, the picture of the captured image can be adjusted in real time to ensure that all target objects are captured well in the picture, thereby achieving a high-quality framing effect. This is particularly important for applications such as group photos and multi-person dances that require a whole scene to be displayed.
[0156] (2) Introduce a main character object to ensure picture stability of the captured image
[0157] The application embodiment introduces a main character mode. When this mode is turned on, the main character object will always remain within the preset area of the capture. This not only ensures the core position of the main character in the picture, but also prevents the movement of other non-main character objects from causing the lens to shift, thereby maintaining the stability of the picture. This function is very useful in scenarios where a specific person needs to be highlighted, such as stage performances or conference recordings.
[0158] (3) Low latency supports high real-time applications
[0159] Based on the optimization of the tracking algorithm, the tracking latency of the multi-target tracking is low, close to the performance of some single-target trackers. This means that users can apply this group tracker to scenarios with high real-time requirements, such as real-time live streaming and fast dancing. The small latency not only improves user experience, but also ensures that the tracking effect can respond to environmental changes in time in rapidly changing scenarios.
[0160] To implement the method of the application embodiment, the application embodiment also provides a handheld gimbal. The handheld gimbal comprises a handheld part, a gimbal, a shooting device, a memory, and a processor, wherein the gimbal is arranged on the handheld part; the shooting device is mounted on the gimbal; the memory is used to store a computer program; and the processor is used to execute the computer program to perform the aforementioned shooting method of the application embodiment.
[0161] FIG. 5 is a structural schematic diagram of a handheld gimbal according to an embodiment of the application. Referring to FIG. 5, the handheld gimbal comprises a handheld part 501, a gimbal 502 arranged on the handheld part 501, and a shooting device 503 fixed to the gimbal 502, wherein the shooting device 503 can be a mobile phone with a camera function. The handheld gimbal further comprises a memory storing a computer program and a processor configured to execute the computer program to perform the aforementioned shooting method of the application embodiment, thereby achieving the tracking shooting effect of multi-target tracking of the mobile phone in a group tracking scenario.
[0162] To implement the method of the embodiments of the present application, the embodiments of the present application further provide a UAV. The UAV comprises a body, a power system, a holder, a shooting device, a memory and a processor, wherein the power system is arranged on the body and is configured to provide power for the UAV; the holder is connected to the body; the shooting device is mounted on the holder; the memory is configured to store a computer program; and the processor is configured to execute the computer program to implement the shooting method of the embodiments of the present application.
[0163] FIG. 6 is a structural schematic diagram of a UAV according to an embodiment of the present application. Referring to FIG. 6, the UAV comprises a body 601, a holder 602 arranged on the body 601, and a shooting device 603 fixed to the holder 602. The UAV further comprises a power system 604 arranged on the body 601, and the power system 604 is configured to provide power for the UAV. The shooting device 603 can be a CCD camera. The UAV further comprises a memory configured to store a computer program and a processor. The processor is configured to execute the computer program to implement the aforementioned shooting method of the embodiments of the present application, thereby achieving the tracking and shooting effect of the CCD camera of the UAV in the multi-target tracking scene.
[0164] The aforementioned processor can be implemented by one or more of an application specific integrated circuit (ASIC), a DSP, a programmable logic device (PLD), a complex programmable logic device (CPLD), a field programmable logic gate array (FPGA), a general purpose processor, a controller, a micro controller unit (MCU), a microprocessor, or other electronic elements, to execute the aforementioned shooting method.
[0165] Exemplarily, the processor is configured to:
[0166] obtain a tracking start instruction indicating starting of the multi-target tracking;
[0167] identify a plurality of target objects in a captured image of the shooting device based on the tracking start instruction, and track the plurality of target objects in the captured image;
[0168] automatically adjust a posture of the shooting device and / or a shooting parameter of the shooting device based on tracking information of the plurality of target objects, so that the plurality of target objects are kept in the captured image.
[0169] Exemplarily, the tracking start instruction has a tracking mode, and a type of the tracking mode comprises: a first tracking mode and a second tracking mode, wherein in the first tracking mode, the plurality of target objects at least comprises a main character object; and in the second tracking mode, the plurality of target objects does not comprise a main character object.
[0170] Exemplarily, in the first tracking mode, the plurality of target objects are kept in the captured image, comprising:
[0171] The main character object is kept in a preset area of the captured image, and the plurality of target objects are kept in the captured image.
[0172] Exemplarily, the preset area is a central area of the captured image.
[0173] Exemplarily, the central area is a center sub-grid area selected after the captured image is divided according to a 3*3 grid.
[0174] Exemplarily, in the first tracking mode, the overall display position of the plurality of target objects is adjusted based on the display position of the main character object, so that the main character object is kept in the preset area of the captured image, and the center of the plurality of target objects is close to or located at the center of the captured image.
[0175] Exemplarily, in the first tracking mode, the tracking information comprises: a first group frame and a main character frame, wherein the first group frame is adjusted based on the position of the main character frame, so that the main character frame is located in the preset area while the center of the first group frame is close to or located at the center of the captured image, the first group frame is determined based on the plurality of target objects, and the main character frame is determined based on the main character object.
[0176] Exemplarily, the center of the first group frame is determined based on the deviation amount between the main character frame and the preset area, so that the deviation amount between the center of the first group frame and the center of the captured image is minimized.
[0177] Exemplarily, in the second tracking mode, the tracking information comprises a second group frame, wherein the center of the second group frame coincides with the center of the captured image, and the second group frame is determined based on the plurality of target objects.
[0178] Exemplarily, the tracking start instruction for starting the multi-target tracking comprises:
[0179] The tracking start instruction is generated based on a first interaction operation of selecting an area on the captured image of the shooting device.
[0180] Exemplarily, the first interaction operation of selecting a region on the image captured by the photographing device generates the tracking start instruction, including:
[0181] Based on the first interaction operation, if it is determined that the selected region corresponds to at least two target objects, the tracking start instruction is generated.
[0182] Exemplarily, the processor is configured to:
[0183] Based on the tracking start instruction, the first tracking mode is started, and the main character object is selected.
[0184] Exemplarily, the selection of the main character object includes:
[0185] Based on the selected region and the positions of the at least two target objects, one target object is selected as the main character object.
[0186] Exemplarily, the selection of the main character object based on the selected region and the positions of the at least two target objects includes:
[0187] The target object closest to the center of the selected region is selected from the at least two target objects as the main character object.
[0188] Exemplarily, the processor is configured to:
[0189] Based on the tracking start instruction, the second tracking mode is started.
[0190] Exemplarily, the tracking start instruction indicating the start of multi-target tracking includes:
[0191] The second interaction operation based on voice collection and / or image collection generates the tracking start instruction.
[0192] Exemplarily, the second interaction operation based on voice collection and / or image collection generates the tracking start instruction, including:
[0193] Based on the second interaction operation, if it is determined that at least one of a set voice instruction and a set gesture action is collected, the tracking start instruction is generated.
[0194] Exemplarily, the processor is configured to:
[0195] Based on the tracking start instruction, the first tracking mode is started, and the main character object is selected.
[0196] Exemplarily, the selection of the main character object includes:
[0197] determine a main character object based on the second interactive operation information source.
[0198] Exemplarily, the determining a main character object based on the second interactive operation information source comprises:
[0199] determine a main character object based on a target object corresponding to the collected setting voice instruction or setting gesture action.
[0200] Exemplarily, the processor is configured to:
[0201] start the second tracking mode based on the tracking start instruction.
[0202] Exemplarily, the processor is configured to:
[0203] switch the tracking mode of the tracking start instruction based on a switching instruction for switching tracking modes in a shooting process.
[0204] Exemplarily, the processor is configured to:
[0205] exit the current tracking mode based on an exit instruction indicating to exit the tracking mode.
[0206] Exemplarily, the tracking information comprises a group box indicating display positions of the plurality of target objects, and the automatically adjusting the posture of the shooting device and / or the shooting parameter of the shooting device, so that the plurality of target objects are kept in the image captured by the shooting device, comprises:
[0207] automatically adjust the posture of the shooting device and / or the shooting parameter of the shooting device based on a center position and / or a size of the group box, so that the plurality of target objects are kept in the image captured by the shooting device.
[0208] Exemplarily, the identifying the plurality of target objects in the image captured by the shooting device comprises:
[0209] detect bodies and heads of the plurality of target objects in the image captured by the shooting device to identify the plurality of target objects in the image captured by the shooting device.
[0210] Exemplarily, the detecting bodies and heads of the plurality of target objects in the image captured by the shooting device to identify the plurality of target objects in the image captured by the shooting device comprises:
[0211] perform target detection on the captured image based on a pre-trained target detector to obtain a first detection result corresponding to a body of a target object and a second detection result corresponding to a head of the target object; wherein a confidence threshold of the first detection result is greater than a confidence threshold of the second detection result.
[0212] Based on the positional relationship between the body and the head of the target object, the obtained first detection result and the second detection result are matched, and based on the matched first detection result and the second detection result, the body region parameter and the head region parameter of the recognized target object are obtained.
[0213] Exemplarily, the tracking of the plurality of target objects in the captured image comprises:
[0214] In the first tracking mode, a single-target tracking algorithm for tracking a body is adopted for a main character object in the plurality of target objects, and a multi-target tracking algorithm for tracking a head is adopted for other target objects except the main character object, to generate a first group box of group tracking.
[0215] In the second tracking mode, the multi-target tracking algorithm for tracking a head is adopted for the plurality of target objects, to generate a second group box of group tracking.
[0216] Exemplarily, the single-target tracking algorithm for tracking a body is adopted for the main character object in the plurality of target objects, and the multi-target tracking algorithm for tracking a head is adopted for other target objects except the main character object, to generate the first group box of group tracking, which comprises:
[0217] The single-target tracking algorithm for tracking a body is adopted for the main character object, to obtain a head region parameter of the main character object.
[0218] The multi-target tracking algorithm for tracking a head is adopted for other target objects except the main character object, to obtain a head region parameter of the other target objects.
[0219] Based on the head region parameter of the main character object and the head region parameter of the other target objects, the first group box is generated.
[0220] Exemplarily, the first group box is generated based on the head region parameter of the main character object and the head region parameter of the other target objects, which comprises:
[0221] Based on the head region parameter of the main character object and the head region parameter of the other target objects, an initial group box and a head box of the main character object are generated, wherein the initial group box corresponds to the head region of all target objects, and the head box of the main character object corresponds to the head region of the main character object.
[0222] A first deviation amount is determined based on the center of the initial group box and the center of the captured image, and a second deviation amount is determined based on the head box of the main character object and a preset region of the captured image.
[0223] Based on the first deviation amount and the second deviation amount, the offset amount of the initial group box is corrected.
[0224] Offset correction is performed on the initial group frame based on the offset amount, to obtain the first group frame.
[0225] Exemplarily, the single-target tracking algorithm for tracking the body of the main character object is adopted to obtain the head region parameter of the main character object, including:
[0226] Based on the body region parameter of the main character object, a single-target tracking algorithm is adopted to track the body region.
[0227] Based on the head region matched with the tracked body region, the head region parameter of the main character object is obtained.
[0228] Exemplarily, based on the body region parameter of the main character object, the single-target tracking algorithm is adopted to track the body region, including:
[0229] For the captured image of the non-key frame, IoU (intersection over union) is calculated based on the identified body region parameter of the target object and the body region parameter of the main character object of the previous frame, and the body region parameter and the re-identification feature used for body region tracking are updated based on the IoU, wherein the re-identification feature represents the image feature of the corresponding body region.
[0230] For the captured image of the key frame, the re-identification feature is extracted based on the identified body region parameter of the target object, and is compared with the updated re-identification feature of the main character object of the previous frame, and the body region parameter and the re-identification feature used for body region tracking are updated based on the comparison result.
[0231] The processor is configured to:
[0232] Based on a set frequency, the captured image of the shooting device is set as the captured image of the key frame.
[0233] It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM). The magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory described in the embodiments of the present application is intended to include, but not limited to, these and any other suitable types of memory.
[0234] In the example embodiments, the embodiments of the present application further provide a computer storage medium, which can be specifically a computer readable storage medium, for example, a memory for storing a computer program, and the computer program can be executed by a processor to complete the steps described in the embodiments of the present application. The computer readable storage medium can be a ROM, a PROM, an EPROM, an EEPROM, a Flash Memory, a magnetic surface memory, an optical disc, or a CD-ROM memory, etc.
[0235] In the example embodiments, the embodiments of the present application further provide a computer program product, which includes a computer program, and the computer program can be executed by a processor of a shooting system to complete the steps described in the embodiments of the present application.
[0236] It should be noted that "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0237] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0238] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A photographing method of a photographing system, characterized by, The photographing system comprises a holder and a photographing device carried on the holder; and the photographing method comprises the following steps: obtaining a tracking starting instruction indicating starting multi-target tracking; based on the tracking starting instruction, identifying multiple target objects in a captured image of the photographing device, and tracking the multiple target objects in the captured image; based on tracking information of the multiple target objects, automatically adjusting a posture of the photographing device and / or a photographing parameter of the photographing device, so that the multiple target objects are kept in the captured image.
2. The method of claim 1, wherein, The tracking starting instruction has a tracking mode, and the type of the tracking mode comprises a first tracking mode and a second tracking mode, wherein in the first tracking mode, the multiple target objects at least include a main character object; and in the second tracking mode, the multiple target objects do not include a main character object.
3. The method of claim 2, wherein, In the first tracking mode, keeping the multiple target objects in the captured image comprises: the main character object is kept in a preset area of the captured image, and the multiple target objects are kept in the captured image.
4. The method of claim 3, wherein, The preset area is a central area of the captured image.
5. The method of claim 4, wherein, The central area is a center sub-grid area selected after the captured image is divided according to a 3*3 grid.
6. The method of claim 3, wherein, In the first tracking mode, the overall display position of the multiple target objects is adjusted based on the display position of the main character object, so that the main character object is kept in the preset area of the captured image, and the center of the multiple target objects is close to or located at the center of the captured image.
7. The method of claim 6, wherein, In the first tracking mode, the tracking information comprises a first group frame and a main character frame, wherein the first group frame is adjusted based on the position of the main character frame, so that the main character frame is located in the preset area, and the center of the first group frame is close to or located at the center of the captured image, the first group frame is determined based on the multiple target objects, and the main character frame is determined based on the main character object.
8. The method of claim 7, wherein, The center of the first group frame is determined based on the deviation amount between the main character frame and the preset area, so that the deviation amount between the center of the first group frame and the center of the captured image is minimized.
9. The method of claim 2, wherein, In the second tracking mode, the tracking information comprises a second group frame, wherein the center of the second group frame coincides with the center of the captured image, and the second group frame is determined based on the multiple target objects.
10. The method of claim 2, wherein, The method further comprises: based on a first interaction operation of selecting an area on the captured image of the photographing device, generating the tracking starting instruction.
11. The method of claim 10, wherein, The method further comprises: based on the first interaction operation, determining that the selected area corresponds to at least two target objects, and then generating the tracking starting instruction.
12. The method of claim 11, wherein, The method further comprises: based on the tracking starting instruction, starting the first tracking mode, and selecting the main character object.
13. The method of claim 12, wherein, The method further comprises: based on the selected area and the positions of the at least two target objects, selecting one target object as the main character object.
14. The method of claim 13, wherein, The selecting a target object as the main character object based on the selected area and the positions of the at least two target objects comprises: selecting a target object closest to the center of the selected area as the main character object from the at least two target objects.
15. The method of claim 11, wherein, The method further comprises: starting the second tracking mode based on the tracking starting instruction.
16. The method of claim 2, wherein, The obtaining the tracking starting instruction indicating starting multi-target tracking comprises: generating the tracking starting instruction based on a second interactive operation of voice collection and / or image collection.
17. The method of claim 16, wherein, The generating the tracking starting instruction based on the second interactive operation of voice collection and / or image collection comprises: generating the tracking starting instruction based on determining that at least one of a set voice instruction and a set gesture action is collected based on the second interactive operation.
18. The method of claim 17, wherein, The method further comprises: starting the first tracking mode and selecting the main character object based on the tracking starting instruction.
19. The method of claim 18, wherein, The selecting the main character object comprises: determining the main character object based on an information source of the second interactive operation.
20. The method of claim 19, the determining the main character object based on the information source of the second interactive operation comprises: determining the main character object based on a target object corresponding to the set voice instruction or the set gesture action collected. The method further comprises:
21. The method of claim 17, wherein, starting the second tracking mode based on the tracking starting instruction. The method further comprises:
22. The method of claim 2, wherein, switching the tracking mode of the tracking starting instruction based on a switching instruction for switching the tracking mode in a shooting process. The method further comprises:
23. The method of claim 2, wherein, quitting the current tracking mode based on a quitting instruction indicating quitting the tracking mode. The tracking information comprises a group box indicating display positions of the plurality of target objects, and the automatically adjusting the posture of the shooting device and / or the shooting parameter of the shooting device so that the plurality of target objects are kept in the image captured by the shooting device comprises:
24. The method of claim 1, wherein, automatically adjusting the posture of the shooting device and / or the shooting parameter of the shooting device based on a center position and / or a size of the group box so that the plurality of target objects are kept in the image captured by the shooting device. The identifying the plurality of target objects in the image captured by the shooting device comprises:
25. The method of claim 1, wherein, detecting bodies and heads of the plurality of target objects in the image captured by the shooting device to identify the plurality of target objects in the image captured by the shooting device. The detecting the bodies and the heads of the plurality of target objects in the image captured by the shooting device to identify the plurality of target objects in the image captured by the shooting device comprises:
26. The method of claim 25, wherein, performing target detection on the captured image based on a pre-trained target detector to obtain a first detection result corresponding to a body of a target object and a second detection result corresponding to a head of the target object, wherein a confidence threshold of the first detection result is greater than a confidence threshold of the second detection result; matching the obtained first detection result and second detection result based on a positional relationship between the body and the head of the target object, and obtaining body region parameters and head region parameters of the identified target object based on the matched first detection result and second detection result. The tracking the plurality of target objects in the captured image comprises:
27. The method of claim 2, wherein, In the first tracking mode, a single-target tracking algorithm for tracking a body is used for a main character object in the plurality of target objects, and a multi-target tracking algorithm for tracking a head is used for other target objects except the main character object, to generate a first group box of group tracking; In the second tracking mode, the multi-target tracking algorithm for tracking a head is used for the plurality of target objects, to generate a second group box of group tracking.
28. The method of claim 27, wherein, The single-target tracking algorithm for tracking a body is used for the main character object, to obtain a head region parameter of the main character object; The multi-target tracking algorithm for tracking a head is used for the other target objects except the main character object, to obtain a head region parameter of the other target objects; The first group box is generated based on the head region parameter of the main character object and the head region parameter of the other target objects. The first group box is generated based on the head region parameter of the main character object and the head region parameter of the other target objects, including:
29. The method of claim 28, wherein, The initial group box corresponding to the head regions of all target objects and the head box of the main character object corresponding to the head region of the main character object are generated based on the head region parameter of the main character object and the head region parameter of the other target objects; A first deviation amount is determined based on the center of the initial group box and the center of the captured image, and a second deviation amount is determined based on the head box of the main character object and a preset region of the captured image; The offset amount of the initial group box is corrected based on the first deviation amount and the second deviation amount; The initial group box is offset corrected based on the offset amount, to obtain the first group box. The single-target tracking algorithm for tracking a body is used for the main character object, to obtain the head region parameter of the main character object, including: The single-target tracking algorithm is used for body region tracking based on the body region parameter of the main character object; 30. The method of claim 28, wherein, The head region parameter of the main character object is obtained based on the head region matched with the tracked body region. The single-target tracking algorithm is used for body region tracking based on the body region parameter of the main character object, including: For a captured image of a non-key frame, an intersection over union (IoU) is calculated based on the identified body region parameter of the target object and the body region parameter of the main character object of a previous frame, and the body region parameter and a re-identification feature used for body region tracking are updated based on the IoU, where the re-identification feature represents an image feature of the corresponding body region; 31. The method of claim 30, wherein, For a captured image of a key frame, a re-identification feature is extracted based on the identified body region parameter of the target object, and is compared with the updated re-identification feature of the main character object of a previous frame, and the body region parameter and the re-identification feature used for body region tracking are updated based on a comparison result. The method further includes: 32. The method of claim 31, wherein, setting a capture image of a photographing device as the key frame based on a set frequency.
33. A handheld gimbal, comprising: Comprise: A handheld part; A holder, provided on the handheld part; A photographing device, carried on the holder; A memory, for storing a computer program; A processor, configured to, when running the computer program: Obtain a tracking start instruction indicating starting multi-target tracking; Based on the tracking start instruction, identify a plurality of target objects in a capture image of the photographing device, and track the plurality of target objects in the capture image; Based on tracking information of the plurality of target objects, automatically adjust the posture of the photographing device and / or the photographing parameters of the photographing device, so that the plurality of target objects remain in the capture image.
34. The handheld gimbal of claim 33, wherein, The tracking start instruction has a tracking mode, and the type of the tracking mode includes: a first tracking mode and a second tracking mode, wherein in the first tracking mode, the plurality of target objects at least includes a main character object; in the second tracking mode, the plurality of target objects does not include a main character object.
35. The handheld gimbal of claim 34, wherein, In the first tracking mode, the plurality of target objects remaining in the capture image includes: The main character object remains in a preset area of the capture image and the plurality of target objects remain in the capture image.
36. The handheld gimbal of claim 35, wherein, The preset area is the central area of the capture image.
37. The handheld holder according to claim 36, wherein the central area is a central sub-grid area selected after the capture image is divided according to a 3x3 grid.
38. The handheld gimbal of claim 35, wherein, In the first tracking mode, the overall display position of the plurality of target objects is adjusted based on the display position of the main character object, so that the main character object remains in the preset area of the capture image and the center of the plurality of target objects is close to or located at the center of the capture image.
39. The handheld gimbal of claim 38, wherein, In the first tracking mode, the tracking information includes: a first group frame and a main character frame, wherein the first group frame is adjusted based on the position of the main character frame, so that the main character frame is located in the preset area while the center of the first group frame is close to or located at the center of the capture image, the first group frame is determined based on the plurality of target objects, and the main character frame is determined based on the main character object.
40. The handheld gimbal of claim 39, wherein, The center of the first group frame is determined based on the deviation amount between the main character frame and the preset area, so that the deviation amount between the center of the first group frame and the center of the capture image is minimized.
41. The handheld gimbal of claim 34, wherein, In the second tracking mode, the tracking information includes a second group frame, wherein the center of the second group frame coincides with the center of the capture image, and the second group frame is determined based on the plurality of target objects.
42. The handheld gimbal of claim 34, wherein, The tracking start instruction indicating starting multi-target tracking includes: Based on a first interactive operation of selecting an area on the capture image of the photographing device, the tracking start instruction is generated.
43. The handheld gimbal of claim 42, wherein, The tracking start instruction is generated based on the first interactive operation of selecting an area on the capture image of the photographing device, including: If it is determined that the selected area corresponds to at least two target objects based on the first interactive operation, the tracking start instruction is generated.
44. The handheld gimbal of claim 43, wherein, The processor is configured to: start the first tracking mode based on the tracking start instruction, and select the main character object.
45. The handheld gimbal of claim 44, wherein, The selecting the main character object comprises: selecting a target object as the main character object based on the selected area and the positions of the at least two target objects.
46. The handheld gimbal of claim 45, wherein, The selecting a target object as the main character object based on the selected area and the positions of the at least two target objects comprises: selecting a target object closest to the center of the selected area from the at least two target objects as the main character object.
47. The handheld gimbal of claim 43, wherein, The processor is configured to: start the second tracking mode based on the tracking start instruction.
48. The handheld gimbal of claim 34, wherein, The tracking start instruction indicating starting multi-target tracking comprises: generating the tracking start instruction based on a second interactive operation of voice collection and / or image collection.
49. The handheld gimbal of claim 48, wherein, The generating the tracking start instruction based on the second interactive operation of voice collection and / or image collection comprises: generating the tracking start instruction based on determining that at least one of a set voice instruction and a set gesture action is collected based on the second interactive operation.
50. The handheld gimbal of claim 49, wherein, The processor is configured to: start the first tracking mode based on the tracking start instruction, and select the main character object. The selecting the main character object comprises:
51. The handheld gimbal of claim 50, wherein, determining the main character object based on an information source of the second interactive operation.
52. The handheld gimbal of claim 51, wherein the determining the main character object based on the information source of the second interactive operation comprises: determining the main character object based on a target object corresponding to the set voice instruction or the set gesture action collected. The processor is configured to:
53. The handheld gimbal of claim 49, wherein, start the second tracking mode based on the tracking start instruction. The processor is configured to:
54. The handheld gimbal of claim 34, wherein, switch the tracking mode of the tracking start instruction based on a switching instruction for switching the tracking mode in a shooting process. The processor is configured to:
55. The handheld gimbal of claim 34, wherein, exit the current tracking mode based on an exit instruction indicating exiting the tracking mode. The tracking information comprises a group box indicating display positions of the plurality of target objects, and the automatically adjusting the posture of the shooting device and / or the shooting parameter of the shooting device so that the plurality of target objects are kept in the image captured by the shooting device comprises:
56. The handheld gimbal of claim 33, wherein, automatically adjusting the posture of the shooting device and / or the shooting parameter of the shooting device based on a center position and / or a size of the group box so that the plurality of target objects are kept in the image captured by the shooting device. The identifying the plurality of target objects in the image captured by the shooting device comprises:
57. The handheld gimbal of claim 33, wherein, detecting bodies and heads of the plurality of target objects in the image captured by the shooting device to identify the plurality of target objects in the image captured by the shooting device. The detecting the bodies and the heads of the plurality of target objects in the image captured by the shooting device to identify the plurality of target objects in the image captured by the shooting device comprises:
58. The handheld gimbal of claim 57, wherein, performing target detection on the captured image based on a pre-trained target detector to obtain a first detection result corresponding to a body of a target object and a second detection result corresponding to a head of the target object, wherein a confidence threshold of the first detection result is greater than a confidence threshold of the second detection result. The first detection result and the second detection result are matched based on a positional relationship between a body and a head of the target object, and body region parameters and head region parameters of the recognized target object are obtained based on the matched first detection result and second detection result.
59. The handheld gimbal of claim 34, wherein, The tracking of the plurality of target objects in the captured image comprises: In the first tracking mode, a single-target tracking algorithm for tracking a body is used for a main character object in the plurality of target objects, and a multi-target tracking algorithm for tracking a head is used for other target objects except the main character object, to generate a first group box of group tracking; In the second tracking mode, the multi-target tracking algorithm for tracking a head is used for the plurality of target objects, to generate a second group box of group tracking.
60. The handheld gimbal of claim 59, wherein, The single-target tracking algorithm for tracking a body is used for the main character object to obtain head region parameters of the main character object; The multi-target tracking algorithm for tracking a head is used for the other target objects to obtain head region parameters of the other target objects; The first group box is generated based on the head region parameters of the main character object and the head region parameters of the other target objects. The first group box is generated based on the head region parameters of the main character object and the head region parameters of the other target objects, comprising: An initial group box corresponding to head regions of all target objects and a head box of the main character object corresponding to a head region of the main character object are generated based on the head region parameters of the main character object and the head region parameters of the other target objects; 61. The handheld gimbal of claim 60, wherein, A first deviation amount is determined based on a center of the initial group box and a center of the captured image, and a second deviation amount is determined based on the head box of the main character object and a preset region of the captured image; The offset amount of the initial group box is corrected based on the first deviation amount and the second deviation amount; The initial group box is offset corrected based on the offset amount, to obtain the first group box. The single-target tracking algorithm for tracking a body is used for the main character object to obtain head region parameters of the main character object, comprising: The single-target tracking algorithm is used for body region tracking based on the body region parameters of the main character object; 62. The handheld gimbal of claim 60, wherein, The head region parameters of the main character object are obtained based on a head region matched with the tracked body region. The single-target tracking algorithm is used for body region tracking based on the body region parameters of the main character object, comprising: For a captured image of a non-key frame, an intersection over union (IoU) is calculated based on the recognized body region parameters of the target object and body region parameters of the main character object of a previous frame, and body region parameters and re-identification features used for body region tracking are updated based on the IoU, wherein the re-identification features represent image features of the corresponding body region. 63. The handheld gimbal of claim 62, wherein, For the captured image of the key frame, re-identification features are extracted based on the identified body region parameters of the target objects, and are compared with the updated re-identification features of the main character object of the previous frame, and the body region parameters and the re-identification features for body region tracking are updated based on the comparison result.
64. The handheld gimbal of claim 63, wherein, The processor is configured to: set the captured image of the shooting device as the captured image of the key frame based on a set frequency.
65. A drone, comprising: comprise: a machine body; a power system provided in the machine body, the power system being configured to provide power for the unmanned aerial vehicle; a gimbal connected to the machine body; a shooting device mounted on the gimbal; a memory configured to store a computer program; a processor configured to, when running the computer program: obtain a tracking start instruction indicating starting multi-target tracking; identify a plurality of target objects in the captured image of the shooting device based on the tracking start instruction, and track the plurality of target objects in the captured image; automatically adjust the posture of the shooting device and / or the shooting parameters of the shooting device based on the tracking information of the plurality of target objects, so that the plurality of target objects are kept in the captured image The tracking start instruction has a tracking mode, and the type of the tracking mode includes: a first tracking mode and a second tracking mode, wherein in the first tracking mode, the plurality of target objects at least includes a main character object; in the second tracking mode, the plurality of target objects does not include a main character object.
66. The drone of claim 65, wherein, In the first tracking mode, the plurality of target objects kept in the captured image includes:
67. The drone of claim 66, wherein, the main character object is kept in a preset area of the captured image and the plurality of target objects are kept in the captured image. The preset area is the central area of the captured image.
68. The drone of claim 67, wherein, 69. The unmanned aerial vehicle of claim 68, wherein the central area is a center sub-grid area selected after the captured image is divided according to a 3x3 grid. In the first tracking mode, the overall display position of the plurality of target objects is adjusted based on the display position of the main character object, so that the main character object is kept in the preset area of the captured image and the center of the plurality of target objects is close to or located at the center of the captured image.
70. The drone of claim 67, wherein, In the first tracking mode, the tracking information includes a first group frame and a main character frame, wherein the first group frame is adjusted based on the position of the main character frame, so that the main character frame is located in the preset area while the center of the first group frame is close to or located at the center of the captured image, the first group frame is determined based on the plurality of target objects, and the main character frame is determined based on the main character object.
71. The drone of claim 70, wherein, The center of the first group frame is determined based on the deviation amount between the main character frame and the preset area, so that the deviation amount between the center of the first group frame and the center of the captured image is minimized.
72. The drone of claim 71, wherein, In the second tracking mode, the tracking information includes a second group frame, wherein the center of the second group frame coincides with the center of the captured image, and the second group frame is determined based on the plurality of target objects.
73. The drone of claim 66, wherein, The tracking start instruction indicating starting multi-target tracking comprises:
74. The drone of claim 66, wherein, The tracking start instruction is generated based on a first interaction operation of selecting a region on an image captured by the photographing device.
75. The drone of claim 74, wherein, The tracking start instruction is generated based on a first interaction operation of selecting a region on an image captured by the photographing device, and includes: The tracking start instruction is generated based on a determination that the selected region corresponds to at least two target objects based on the first interaction operation.
76. The drone of claim 75, wherein, The processor is configured to: The first tracking mode is started based on the tracking start instruction, and the main character object is selected.
77. The drone of claim 76, wherein, The main character object is selected, and includes: A target object is selected as the main character object based on the selected region and positions of the at least two target objects.
78. The drone of claim 77, wherein, The main character object is selected based on the selected region and positions of the at least two target objects, and includes: A target object closest to a center of the selected region is selected as the main character object from the at least two target objects.
79. The drone of claim 75, wherein, The processor is configured to: The second tracking mode is started based on the tracking start instruction.
80. The drone of claim 66, wherein, The tracking start instruction is obtained, and includes: The tracking start instruction is generated based on a second interaction operation of voice collection and / or image collection.
81. The drone of claim 80, wherein, The tracking start instruction is generated based on a second interaction operation of voice collection and / or image collection, and includes: The tracking start instruction is generated based on a determination that at least one of a set voice instruction and a set gesture action is collected based on the second interaction operation.
82. The drone of claim 81, wherein, The processor is configured to: The first tracking mode is started based on the tracking start instruction, and the main character object is selected.
83. The drone of claim 82, wherein, The main character object is selected, and includes: The main character object is determined based on an information source of the second interaction operation.
84. The unmanned aerial vehicle of claim 83, wherein the main character object is determined based on the information source of the second interaction operation, and includes: The main character object is determined based on a target object corresponding to the set voice instruction or the set gesture action collected.
85. The drone of claim 81, wherein, The processor is configured to: The second tracking mode is started based on the tracking start instruction.
86. The drone of claim 66, wherein, The processor is configured to: The tracking mode of the tracking start instruction is switched based on a switching instruction for switching the tracking mode in a photographing process.
87. The drone of claim 66, wherein, The processor is configured to: A current tracking mode is exited based on an exit instruction indicating to exit the tracking mode.
88. The drone of claim 65, wherein, The tracking information includes a group box indicating display positions of the plurality of target objects, and the posture of the photographing device and / or the photographing parameter of the photographing device is automatically adjusted such that the plurality of target objects are kept in the image captured by the photographing device, and includes: The posture of the photographing device and / or the photographing parameter of the photographing device is automatically adjusted based on a center position and / or a size of the group box such that the plurality of target objects are kept in the image captured by the photographing device.
89. The drone of claim 65, wherein, The plurality of target objects in the image captured by the photographing device are identified, and includes: The plurality of target objects in the image captured by the photographing device are identified by detecting bodies and heads of the plurality of target objects in the image captured by the photographing device.
90. The drone of claim 89, wherein, The detecting a plurality of target objects in a captured image by a shooting device, and identifying the plurality of target objects in the captured image, comprises: performing target detection on the captured image based on a pre-trained target detector to obtain a first detection result corresponding to a body of a target object and a second detection result corresponding to a head of the target object; wherein a confidence threshold of the first detection result is greater than a confidence threshold of the second detection result; matching the obtained first detection result and second detection result based on a positional relationship between the body and the head of the target object, and obtaining a body region parameter and a head region parameter of the identified target object based on the matched first detection result and second detection result. The tracking the plurality of target objects in the captured image, comprises:
91. The drone of claim 66, wherein, in the first tracking mode, adopting a single-target tracking algorithm for tracking a body for a main character object in the plurality of target objects, and adopting a multi-target tracking algorithm for tracking a head for other target objects except the main character object, to generate a first group box of group tracking; in the second tracking mode, adopting the multi-target tracking algorithm for tracking the head for the plurality of target objects, to generate a second group box of group tracking. The adopting the single-target tracking algorithm for tracking the body for the main character object in the plurality of target objects, and adopting the multi-target tracking algorithm for tracking the head for other target objects except the main character object, to generate the first group box of group tracking, comprises:
92. The drone of claim 91, wherein, adopting the single-target tracking algorithm for tracking the body for the main character object to obtain a head region parameter of the main character object; adopting the multi-target tracking algorithm for tracking the head for the other target objects to obtain a head region parameter of the other target objects; generating the first group box based on the head region parameter of the main character object and the head region parameter of the other target objects. The generating the first group box based on the head region parameter of the main character object and the head region parameter of the other target objects, comprises:
93. The drone of claim 92, wherein, generating an initial group box and a head box of the main character object based on the head region parameter of the main character object and the head region parameter of the other target objects, wherein the initial group box corresponds to head regions of all target objects, and the head box of the main character object corresponds to a head region of the main character object; determining a first deviation amount based on a center of the initial group box and a center of the captured image, and determining a second deviation amount based on the head box of the main character object and a preset region of the captured image; correcting a deviation amount of the initial group box based on the first deviation amount and the second deviation amount; performing deviation correction on the initial group box based on the deviation amount to obtain the first group box. The adopting the single-target tracking algorithm for tracking the body for the main character object to obtain the head region parameter of the main character object, comprises:
94. The drone of claim 92, wherein, adopting the single-target tracking algorithm to track a body region based on a body region parameter of the main character object; obtaining the head region parameter of the main character object based on a head region matched with the tracked body region. 95. The drone of claim 94, wherein, The single-target tracking algorithm is used for tracking the body region based on the body region parameter of the main character object, and includes: For the captured image of the non-key frame, an intersection over union IoU is calculated based on the identified body region parameter of the target object and the body region parameter of the main character object of the previous frame, and the body region parameter and the re-identification feature used for body region tracking are updated based on the IoU, wherein the re-identification feature represents the image feature of the corresponding body region; For the captured image of the key frame, the re-identification feature is extracted based on the identified body region parameter of the target object, and is compared with the updated re-identification feature of the main character object of the previous frame, and the body region parameter and the re-identification feature used for body region tracking are updated based on the comparison result.
96. The drone of claim 95, wherein, The processor is configured to: set the captured image of the shooting device as the captured image of the key frame based on a set frequency.
Citation Information
Patent Citations
Multi-person following shooting method and device, equipment and storage medium
CN110232706A
Handheld gimbal and photographing control method thereof
CN111316630A
Method of measuring adhesiveness for secondary battery
KR1020250147452A
Photographing method, photographic device, terminal device and storage medium
WO2023065125A1
Cited By
Gimbal multi-target tracking method for live broadcast scene, gimbal and medium
CN122294001A
Gimbal multi-target tracking method for live broadcast scene, gimbal and medium
CN122294001B