Automatic focusing method and device of shooting equipment, electronic equipment and storage medium

By tracking and detecting the face of a preset target person, and using head, shoulder, or body features to select focus parameters, the problem of inaccurate focus when the person turns around is solved, achieving a stable and continuous autofocus effect.

CN121357409APending Publication Date: 2026-01-16ARASHI VISION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511370135.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing shooting equipment is prone to losing focus when a person turns around because it cannot detect the face, resulting in inaccurate focus and affecting the shooting effect.

Method used

By tracking a preset target person, determining its location and size in the video frame, and using head, shoulder, or body features for face detection, different focus parameters are selected based on the detection results to achieve stable tracking and accurate focus.

Benefits of technology

It achieves stable and continuous tracking when the subject turns around, improving focusing accuracy and ensuring stable imaging results and clear image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121357409A_ABST
    Figure CN121357409A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic focusing method and device of shooting equipment, electronic equipment and a storage medium. The method comprises the following steps: performing target tracking on a preset target person to obtain an area position and a first size of a tracking target corresponding to the preset target person in a current video frame; determining a face detection area in the current video frame according to the area position and the first size; performing face detection on the face detection area to obtain a detection result; selecting a corresponding focusing parameter according to a detection result, and focusing a shooting device which collects the current video frame based on the focusing parameter; wherein when the detection results are different, the selected focusing parameters are different. By adopting the method, the focusing accuracy can be improved, stable and continuous automatic focusing tracking is realized, the imaging effect is stable, and the image quality is clear.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image shooting, and in particular to an automatic focusing method and device of a shooting device, an electronic device, and a storage medium. BACKGROUND

[0002] With the development of image shooting technology, electronic devices have penetrated into our lives, and more and more people use various shooting devices to take photos, record videos, and live broadcast to record their lives. At present, most of the shooting devices on the market generally have an automatic focusing function. In the application scenario with a person, the commonly adopted technical solution is to use a "face-first focusing" mode, that is, when a face is detected in the frame, the face area is focused first, so that the imaging frame has clear quality and clear subject.

[0003] However, when the person turns around, the above method may lose the focus target because the face cannot be detected, which may cause the camera to focus inaccurately, and thus the shooting frame may be poor. SUMMARY

[0004] Therefore, it is necessary to provide an automatic focusing method and device of a shooting device, an electronic device, and a storage medium, which can improve the focusing accuracy.

[0005] In a first aspect, the present application provides an automatic focusing method of a shooting device. The method comprises:

[0006] tracking a preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in a current video frame, the tracking target including any one of a human body, a head, or a head and shoulder part of the preset target person;

[0007] determining a face detection region in the current video frame according to the region position and the first size;

[0008] performing face detection on the face detection region to obtain a detection result;

[0009] selecting a corresponding focusing parameter according to the detection result, and focusing a shooting device collecting the current video frame based on the focusing parameter; wherein the selected focusing parameter is different when the detection result is different.

[0010] In some embodiments, the current video frame contains multiple persons; before tracking the preset target person, the method further comprises:

[0011] determining a center position of a first video frame;

[0012] detect a coverage area of each of the tracking targets corresponding to the characters in the first video frame;

[0013] calculate a center point of the coverage area;

[0014] calculate a distance between the center position and the center point corresponding to each of the tracking targets;

[0015] mark a character corresponding to a minimum distance in the distances as the preset target character.

[0016] In some embodiments, the current video frame contains multiple characters; before the target tracking of the preset target character, the method further comprises:

[0017] detect a coverage area of each of the tracking targets corresponding to the characters in the first video frame; mark a character corresponding to a coverage area with a maximum first dimension in the coverage areas as the preset target character; or,

[0018] in response to an input target selection instruction, mark a character designated by the target selection instruction in the first video frame as the preset target character.

[0019] In some embodiments, the focus parameter is a parameter for focusing on a head feature, a face feature, or a human eye feature;

[0020] selecting a corresponding focus parameter according to the detection result and focusing a shooting device for collecting the current video frame based on the focus parameter comprises:

[0021] selecting a focus parameter for focusing on a head feature, a face feature, or a human eye feature based on the detection result, and focusing a shooting device for collecting the current video frame according to the head feature, the face feature, or the human eye feature.

[0022] In some embodiments, focusing a shooting device for collecting the current video frame according to the head feature, the face feature, or the human eye feature comprises:

[0023] when the face detection area does not contain a face feature of the preset target character, focusing a shooting device for collecting the current video frame according to the head feature;

[0024] when the face detection area contains a face feature of the preset target character, detecting whether the face feature includes a human eye feature; if not, focusing a shooting device for collecting the current video frame according to the face feature; if yes, focusing a shooting device for collecting the current video frame according to the human eye feature in the face feature.

[0025] In some embodiments, the method further comprises:

[0026] When the preset target person is not at the preset position in the current video frame, adjusting the shooting device according to the preset target person so that the preset target person is at the preset position in the current video frame.

[0027] In some embodiments, the determining the face detection region in the current video frame according to the region position and the first size comprises:

[0028] When the tracking target is the head or the head-shoulder part of the preset target person, in the current video frame, a head region or a head-shoulder region corresponding to the region position and the first size is taken as the face detection region.

[0029] When the tracking target is the body of the preset target person, in the current video frame, a target region is selected according to the region position, and the target region is taken as the face detection region; wherein the first size is greater than a second size of the target region.

[0030] In a second aspect, the present application further provides an automatic focusing device of a shooting device. The device comprises:

[0031] a target tracking module, configured to track a preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in a current video frame, the tracking target comprising any one of a body, a head or a head-shoulder part of the preset target person;

[0032] a face detection region determining module, configured to determine a face detection region in the current video frame according to the region position and the first size;

[0033] a face detection module, configured to perform face detection on the face detection region to obtain a detection result;

[0034] a focusing module, configured to select a corresponding focusing parameter according to the detection result, and perform focusing on a shooting device collecting the current video frame based on the focusing parameter; wherein when the detection result is different, the selected focusing parameter is different.

[0035] In a third aspect, the present application further provides an electronic device. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the automatic focusing method of the shooting device in the first aspect when executing the computer program.

[0036] In a fourth aspect, the present application provides a computer readable storage medium. The computer readable storage medium has a computer program stored thereon, and the computer program, when executed by a processor, implements the automatic focusing method of the photographing device in the first aspect.

[0037] In a fifth aspect, the present application provides a computer program product. The computer program product comprises a computer program, and the computer program, when executed by a processor, implements the automatic focusing method of the photographing device in the first aspect.

[0038] The automatic focusing method of the photographing device, the device, the electronic device and the storage medium described above can obtain the region position and the first size of the tracking target corresponding to the preset target person in the current video frame by tracking the preset target person, the tracking target comprising any one of the human body, the head or the head-shoulder part of the preset target person, then determine the face detection region in the current video frame according to the region position and the first size, perform face detection on the face detection region to obtain a detection result, and finally select different focusing parameters according to different detection results, and focus the photographing device collecting the current video frame based on the selected focusing parameters. In the above scheme, the human body / head / shoulder of the tracking target is tracked, compared with the traditional way of directly tracking the face, even if the preset target person turns around and faces away from the photographing device, stable tracking of the target can be realized, so that the photographing device will not lose the focus target when photographing. Further, based on the region position and the first size of the tracking target, the face detection region can be accurately determined based on the tracking of the target, so that automatic and accurate focusing can be realized based on the detection result of the face detection region. In addition, different focusing parameters are selected for different detection results, which further improves the accuracy of focusing. Therefore, the above method not only realizes stable and continuous tracking focusing, but also improves the accuracy of focusing, so that the imaging effect is stable and the picture quality is clear. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 Flowchart of the automatic focusing method of the photographing device in some embodiments;

[0040] Figure 2 Flowchart of the automatic focusing method of the photographing device in some embodiments;

[0041] Figure 3 Flowchart of the automatic focusing method of the photographing device in some embodiments;

[0042] Figure 4 Flowchart of the face detection region determination step in some embodiments;

[0043] Figure 5A flowchart of an auto-focusing method of a photographing device in some embodiments;

[0044] Figure 6 A structural block diagram of an auto-focusing device of a photographing device in some embodiments;

[0045] Figure 7 An internal structural diagram of an electronic device in some embodiments. DETAILED DESCRIPTION

[0046] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0047] In some embodiments, as shown in FIG. 1, an auto-focusing method of a photographing device is provided, and the present embodiment is exemplified by the method applied to a photographing device. The photographing device includes a camera, a smart phone, a tablet computer, a personal computer, and other electronic devices that can take photos. Figure 1

[0048] The auto-focusing method of the photographing device in the present embodiment includes the following steps:

[0049] In step 102, a target tracking is performed on the preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in a current video frame.

[0050] The tracking target can refer to a tracking part when the target tracking is performed on the preset target person, and the tracking target includes any one of a human body, a head, or a head-shoulder part of the preset target person. The region position can refer to a position of a region where the tracking target is located in the current video frame. The first size can refer to a size of the region where the tracking target is located in the current video frame.

[0051] In some embodiments, since the photographing device can capture a side face and a front face of the preset target person during the turning of the preset target person, the features shown by the side face and the front face are different, and when the preset target person turns to face the back of the photographing device, the face of the preset target person disappears. Therefore, when the target tracking is performed on the preset target person, using the aforementioned tracking target for tracking is more stable than directly using the face for tracking.

[0052] ​In some embodiments, when target tracking is performed on the preset target person, an external rectangular frame is used to frame the tracking target, and the framed region is taken as the region of the tracking target in the current video frame (i.e., the region where the tracking target is located). For example, when target tracking is performed on the head of the preset target person, an external rectangular frame is used to frame the head, and the framed region is taken as the region of the head in the current video frame.

[0053] In some embodiments, a preset target tracking algorithm is used to perform target tracking on the preset target person, to obtain the region position and the first size of the tracking target corresponding to the preset target person in the current video frame.

[0054] In some embodiments, the preset target person can be tracked by using a DCF (Discriminative Correlation Filter) and other filter-based trackers, or by using a SiamRPN (Siamese Region Proposal Network) and other CNN (Convolutional Neural Networks)-based trackers, or by using other trackers, which are not limited herein.

[0055] In some embodiments, before step 102, the shooting device can first detect the tracking target of each person, specifically: in response to a target tracking instruction, the second video frame is detected to obtain the tracking target. For example, the user makes an OK gesture to the shooting device, thereby triggering the shooting device to detect the second video frame collected in real time, thereby detecting the head, head-shoulder region or human body of the person.

[0056] Specifically, the target tracking instruction can be an instruction for instructing the shooting device to initiate target tracking on the person, for example, after the user makes an OK gesture, the shooting device generates the target tracking instruction.

[0057] In some embodiments, the target tracking instruction described above can be generated by the shooting device according to the user triggering a specific button, or can be generated by the shooting device according to the user's gesture or human body posture, or can be generated by the shooting device according to the user's instruction to the shooting device through a control device (such as a remote control device), or can be automatically generated by the shooting device after being powered on for a period of time (such as 10s), or can be generated by the shooting device when detecting that a person enters a specified position (such as the center position) of the shooting picture, which is not limited herein.

[0058] In some embodiments, the second video frame is captured in real time by the shooting device, and the second video frame is captured before the current video frame.

[0059] The method for detecting the second video frame can include a common detection method, a hand-crafted feature-based detection method, or a convolutional neural network technology-based detection method. The hand-crafted feature-based detection method includes but is not limited to a template matching method, a key point matching method, and a key feature method. The convolutional neural network technology-based detection method includes but is not limited to a YOLO (You Only Look Once Detector), an SSD (Single Shot MultiBox Detector), an R-CNN (Region-based Convolutional Neural Networks), and a Mask R-CNN (Mask Region-based Convolutional Neural Networks).

[0060] In step 104, a face detection region is determined in the current video frame according to the region position and the first size.

[0061] The face detection region can refer to a region where a head or a head-shoulder part of the preset target person is located.

[0062] In some embodiments, according to the region position and the first size obtained in the foregoing steps, the face detection region is determined in the current video frame from the tracking target. For example, when the tracking target is a head, a region corresponding to a circumscribed rectangular frame of the head can be taken as the face detection region. For example, when the tracking target is a head-shoulder part, a region corresponding to a circumscribed rectangular frame of the head-shoulder part can be taken as the face detection region. For another example, when the tracking target is a human body, one third or one fourth of an upper part of a circumscribed rectangular frame of the human body can be taken as the face detection region.

[0063] In step 106, face detection is performed on the face detection region to obtain a detection result.

[0064] The detection result is used to indicate whether a face of a person (i.e., a human face) exists in the face detection region, i.e., the detection result is used to indicate whether a facial feature of the preset target person exists in the face detection region.

[0065] In some embodiments, a pre-set face detection algorithm is used to perform face detection on the face detection region to obtain the detection result.

[0066] In some embodiments, the face detection algorithm can adopt a common detection method, a hand-crafted feature-based detection method, or a detection method based on a convolutional neural network technology. Among them, the hand-crafted feature-based detection method includes but is not limited to a template matching method, a key point matching method, and a key feature method. The detection method based on the convolutional neural network technology includes but is not limited to YOLO, SSD, R-CNN, and Mask R-CNN.

[0067] It should be noted that when the face detection is performed on the face detection region, if multiple faces are detected, the largest face is selected from the multiple faces as the detection result; or the face closest to the center position of the current video frame is selected from the multiple faces as the detection result; or the face with the highest confidence is selected from the multiple faces as the detection result. Among them, the confidence can be output by the detector corresponding to the aforementioned detection method.

[0068] In step 108, a corresponding focusing parameter is selected according to the detection result, and the shooting device for collecting the current video frame is focused based on the focusing parameter; wherein when the detection result is different, the selected focusing parameter is different.

[0069] Among them, the focusing parameter can refer to a parameter for the shooting device to focus on a specific feature selected. For example, the focusing parameter can be a parameter for focusing on a head feature, a face feature, or an eye feature.

[0070] The detection result can be a first detection result or a second detection result, the first detection result indicating that the face detection region has a face feature, and the second detection result indicating that the face detection region does not have a face feature. Since the pre-set target person can turn his head, when the face detection is performed on the face detection region, the face feature can be detected or not detected, and thus when the detection result is different, the selected focusing parameter is different.

[0071] In some embodiments, by selecting a focusing parameter corresponding to the detection result according to the detection result, and focusing the shooting device based on the focusing parameter, the shooting device will not lose the focusing target, thereby improving the accuracy of focusing.

[0072] In some embodiments, after determining the detection result, the shooting device can be focused by using a contrast focusing method or a phase focusing method.

[0073] Contrast Detection Auto Focus (CDAF), which searches for the position of the lens with the maximum contrast of the focusing area as the focus point by moving the lens back and forth. For example, when focusing on a face area, the contrast is calculated according to the pixels of the face area, and then the moving direction of the lens of the photographing device is determined according to the calculated contrast. When the position with the maximum contrast is found in repeated movement, the focusing of the lens of the photographing device is completed. When searching, a common search method is a hill climbing search algorithm.

[0074] Phase Detection Auto Focus (PDAF), which reserves PDAF pixel points on the photosensitive element to specifically perform phase detection, calculates the offset value of focusing through the phase difference, and then quickly moves the lens to the target position according to the offset value, thereby achieving accurate focusing. Compared with contrast focusing, phase focusing does not require repeated movement of the lens of the photographing device, and the focusing speed is fast. For example, when focusing on a face area, the phase difference is calculated according to the PDAF pixel points corresponding to the face area, and then the offset value of focusing is determined according to the phase difference. Then, the lens of the photographing device is moved according to the offset value to achieve focusing.

[0075] It should be noted that the focusing speed of phase focusing is faster than that of contrast focusing, and the success rate of contrast focusing is higher than that of phase focusing. The specific focusing method is not limited in the present application, and a suitable focusing method can be selected according to the actual situation.

[0076] The automatic focusing method of the photographing device in the embodiment of the present application performs target tracking on the preset target person, obtains the area position and the first size of the tracking target corresponding to the preset target person in the current video frame, then determines the face detection area in the current video frame according to the area position and the first size, performs face detection on the face detection area, obtains the detection result, and finally selects different focusing parameters according to different detection results, and focuses the photographing device for collecting the current video frame based on the selected focusing parameters. Therefore, when the photographing device is photographing, the focusing target will not be lost, thereby improving the accuracy of focusing, realizing stable and continuous automatic focusing tracking, making the imaging effect stable, the picture clear, and the subject clear, and improving the user experience.

[0077] As shown in FIG. 1, Figure 2 In some embodiments, the current video frame contains multiple persons, and before step 102, the automatic focusing method of the photographing device further includes:

[0078] Step 202, determining the center position of the first video frame.

[0079] The first video frame is collected by a shooting device and is collected before the current video frame. The center position can refer to the center of a picture corresponding to the first video frame collected by the shooting device.

[0080] For example, the center position can be the body center of the first video frame. For example, the picture corresponding to the first video frame is placed in a Cartesian coordinate system, and after the coordinates of the picture boundary are determined, the center position of the first video frame is determined according to the coordinates of the picture boundary.

[0081] Step 204, detecting the coverage area of the tracking target corresponding to each character in the first video frame.

[0082] The coverage area can refer to the area occupied by the circumscribed rectangular frame of the tracking target in the first video frame. For example, when the tracking target is a head, the coverage area can refer to the area covered by the circumscribed rectangular frame of the head in the first video frame.

[0083] The coverage area of the circumscribed rectangular frame of the tracking target of each character in the first video frame is detected to obtain the coverage area of the tracking target corresponding to each character.

[0084] Step 206, calculating the center point of the coverage area.

[0085] The center point of the coverage area can refer to the center position of the circumscribed rectangular frame of the tracking target.

[0086] For example, when the center point of the coverage area is the center position of the circumscribed rectangular frame of the tracking target, the center point of the coverage area can be determined according to the four corners of the circumscribed rectangular frame.

[0087] Step 208, calculating the distance between the center position and the center point corresponding to each tracking target.

[0088] The distance between the center position and the center point corresponding to each tracking target can be Euclidean distance, Manhattan distance, or other distances, which are not specifically limited in the present application. For example, when Euclidean distance is used, the distance between the center position and the center point corresponding to the tracking target is calculated by the calculation formula of Euclidean distance.

[0089] Step 210, marking the character corresponding to the minimum distance in the plurality of distances as a preset target character.

[0090] In some embodiments, the character corresponding to the minimum distance in the plurality of distances calculated in step 208 is marked as a preset target character for target tracking.

[0091] In some embodiments, before step 102, the current video frame contains multiple people, and the autofocus method of the shooting device further includes: detecting the coverage area of ​​the tracking target corresponding to each person in the first video frame; marking the person corresponding to the first largest coverage area in each coverage area as a preset target person; or, in response to the input target selection instruction, marking the person specified by the target selection instruction in the first video frame as a preset target person.

[0092] Specifically, in this embodiment, the coverage area can refer to the area covered by the tracking target in the first video frame. After detecting the coverage area of ​​the tracking target corresponding to each person in the first video frame, the person corresponding to the largest first-size coverage area in each coverage area is marked as the preset target person for target tracking.

[0093] In some embodiments, the first video frame is displayed on the shooting device, and the user inputs a target selection command on the shooting device so that the shooting device uses the person specified by the user's target selection command as the preset target person and then performs target tracking on the preset target person.

[0094] In some embodiments, the focus parameters are parameters used to represent focusing using head features, facial features, or eye features. Step 108 includes, but is not limited to, the following steps: selecting focus parameters for focusing using head features, facial features, or eye features based on the detection results, and focusing the shooting device that captures the current video frame based on the head features, facial features, or eye features.

[0095] Specifically, in this embodiment, depending on the detection results, a feature is selected from head features, facial features, or human eye features as a focus parameter, and the shooting device capturing the current video frame is focused based on the selected focus parameter.

[0096] like Figure 3 As shown, in some embodiments, the step "focusing the capturing device that captures the current video frame based on head features, facial features, or eye features" includes, but is not limited to, the following steps:

[0097] Step 302: When the face detection area does not contain the facial features of the preset target person, focus the shooting device that captures the current video frame based on the head features.

[0098] In some embodiments, when the detection result indicates that the face detection area does not include the facial features of the preset target person, it means that in the current video frame, the preset target person is facing away from the shooting device. In this case, the shooting device capturing the current video frame is focused on the head features, that is, the shooting device is focused on the head of the preset target person.

[0099] Exemplarily, when the face detection region does not include the facial feature of the preset target person, a contrast focusing method or a phase focusing method can be adopted to focus the shooting device for collecting the current video frame according to the head feature.

[0100] For example, when the contrast focusing method is adopted, a circumscribed rectangle frame of the head feature is obtained, then a contrast is calculated according to the pixels of the circumscribed rectangle frame of the head feature, and meanwhile, a moving direction of the lens of the shooting device is determined according to the calculated contrast, and when the position with the maximum contrast is found in the repeated moving process, the focusing of the lens of the shooting device is completed.

[0101] For example, when the phase focusing method is adopted, a circumscribed rectangle frame of the head feature is obtained, then a phase difference is calculated according to the PDAF pixel points of the circumscribed rectangle frame of the head feature, and then an offset value of focusing is determined according to the phase difference, and the lens of the shooting device is moved according to the offset value to realize focusing.

[0102] Step 304, when the face detection region includes the facial feature of the preset target person, it is detected whether the facial feature includes the eye feature; if not, the shooting device for collecting the current video frame is focused according to the facial feature; if yes, the shooting device for collecting the current video frame is focused according to the eye feature in the facial feature.

[0103] In some embodiments, when the face detection region includes the facial feature of the preset target person, it indicates that in the current video frame, the preset target person is facing the shooting device, in this case, the face detection region is further judged, and it is detected whether the facial feature in the collected first video frame includes the eye feature, if the facial feature includes the eye feature, the shooting device for collecting the current video frame is focused according to the eye feature, that is, the shooting device is focused on the eye of the preset target person. If the facial feature does not include the eye feature, the shooting device for collecting the current video frame is focused according to the facial feature, that is, the shooting device is focused on the face of the preset target person.

[0104] When the facial feature includes the eye feature, the shooting device is focused on the eye of the preset target person, so that the person in the video frame obtained by shooting is more spiritual, and the imaging effect is improved. For example, when the eye feature exists, the shooting device is focused on the mouth or nose of the person, which will make the video frame obtained by shooting not spiritual.

[0105] Exemplarily, when the facial feature includes the eye feature, the phase focusing method or the contrast focusing method can be adopted to focus the shooting device.

[0106] For example, when using the contrast focusing method, the circumscribed rectangular frame of the preset target person's eye can be considered as the focusing area, then the contrast is calculated according to the pixels of the focusing area, and then the moving direction of the lens of the shooting device is determined according to the calculated contrast, and when the position with the maximum contrast is found in the repeated movement, the focusing of the lens of the shooting device is completed.

[0107] For example, when using the phase focusing method, the circumscribed rectangular frame of the preset target person's eye can be considered as the focusing area, then the phase difference is calculated according to the PDAF pixel points corresponding to the focusing area, then the offset value of focusing is determined according to the phase difference, and then the lens of the shooting device is moved according to the offset value to realize focusing.

[0108] In some embodiments, the automatic focusing method of the shooting device further includes: when the preset target person is not at the preset position of the current video frame, adjusting the shooting device according to the preset target person to make the preset target person at the preset position of the current video frame.

[0109] Specifically, in the present embodiment, the preset position can refer to a pre-designated position. The preset position can be the center position of the current video frame, or other positions, which are not limited in the present application.

[0110] When the preset target person is not at the preset position of the current video frame, the shooting device is adjusted to rotate to make the preset target person at the preset position of the current video frame.

[0111] In some embodiments, the shooting device can be clamped by a gimbal to fix the shooting device. During the focusing process of the shooting device, the rotation of the gimbal can be adjusted to adjust the position of the preset target person in the video frame shot by the shooting device.

[0112] In some embodiments, the shooting device can be adjusted by the following steps: calculating the offset amount of the position of the preset target person in the current video frame from the preset position; sending a control instruction to the gimbal according to the offset amount to make the gimbal adjust according to the control instruction; wherein the control instruction is an instruction generated according to the offset amount and used to adjust the gimbal.

[0113] For example, when the offset amount is less than the offset threshold, it can be considered that the position of the preset target person in the current video frame and the preset position have a small phase difference, and the gimbal can not be adjusted, or can be fine-tuned according to the specific offset amount. That is, in this case, the shooting device can not generate a control instruction to make the gimbal remain stationary; or a control instruction can be generated and sent to the gimbal according to the specific offset amount, so that the gimbal is adjusted by a small amount according to the control instruction, so that the preset target person is at the preset position.

[0114] Exemplarily, when the offset is greater than or equal to the offset threshold, it can be considered that the position of the preset target person in the current video frame and the preset position are greatly different. In this case, the shooting device generates and sends a control instruction to the holder according to the specific offset, so that the holder adjusts according to the control instruction, so that the preset target person is in the preset position.

[0115] It should be noted that the adjustment processing of the shooting device is always ongoing to ensure that the preset target person is in the preset position of the video frame collected by the shooting device.

[0116] Please refer to Figure 4 In some embodiments, step 104 includes but is not limited to the following steps:

[0117] Step 402, when the tracking target is the head or head-shoulder part of the preset target person, in the current video frame, the head region or head-shoulder region corresponding to the region position and the first size is taken as the face detection region.

[0118] Step 404, when the tracking target is the body of the preset target person, in the current video frame, a target region is selected according to the region position, and the target region is taken as the face detection region; wherein the first size is greater than the second size of the target region.

[0119] Specifically, in the present embodiment, the head region can refer to the region corresponding to the head, such as the region covered by the circumscribed rectangular frame of the head. The head-shoulder region can refer to the region corresponding to the head and shoulder, such as the region covered by the circumscribed rectangular frame of the head-shoulder part.

[0120] When the tracking target is the head of the preset target person, then in the current video frame, the circumscribed rectangular frame corresponding to the head is determined according to the region position and the first size, and the circumscribed rectangular frame is taken as the head region corresponding to the head. The head region is taken as the face detection region.

[0121] When the tracking target is the head-shoulder part of the preset target person, then in the current video frame, the circumscribed rectangular frame corresponding to the head-shoulder part is determined according to the region position and the first size, and the circumscribed rectangular frame is taken as the head-shoulder region corresponding to the head-shoulder part. The head-shoulder region is taken as the face detection region.

[0122] When the tracking target is the body of the preset target person, then in the current video frame, a target region is selected according to the region position, and the circumscribed rectangular frame of the part is taken as the face detection region.

[0123] For example, a quarter of the human body above is selected as a target region, and the target region is taken as a face detection region, and the first size is four times the second size of the target region. Similarly, a third of the human body above can be selected as a target region, or a half of the human body above can be selected as a target region.

[0124] Please refer to Figure 5 Some embodiments of the present application provide an automatic focusing method of a shooting device, including but not limited to the following steps:

[0125] Step 502, in response to a target tracking instruction, detecting a second video frame to obtain a tracking target.

[0126] Step 504, detecting the coverage area of each tracking target of a person in a first video frame; marking the person corresponding to the coverage area with the largest first size as a preset target person.

[0127] Step 506, performing target tracking on the preset target person to obtain the area position and the first size of the tracking target of the preset target person in a current video frame.

[0128] Step 508, when the tracking target is the head or the shoulder part of the preset target person, taking the head region or the shoulder region corresponding to the area position and the first size as a face detection region in the current video frame; when the tracking target is the human body of the preset target person, selecting a target region according to the area position in the current video frame, and taking the target region as a face detection region, wherein the first size is greater than the second size of the target region.

[0129] Step 510, performing face detection on the face detection region to obtain a detection result.

[0130] In some embodiments, a pre-set face detection algorithm is used to perform face detection on the face detection region to obtain a detection result. If the face feature of the preset target person is detected, step 414 is executed, and if the detection result indicates that the face detection region does not contain the face feature of the preset target person, step 512 is executed.

[0131] Step 512, when the face detection region does not contain the face feature of the preset target person, focusing the shooting device collecting the current video frame according to the head feature.

[0132] Step 514, when the face detection region contains the face feature of the preset target person, detecting whether the face feature includes an eye feature; if not, focusing the shooting device collecting the current video frame according to the face feature; if yes, focusing the shooting device collecting the current video frame according to the eye feature in the face feature.

[0133] For example, when the facial feature includes a human eye feature, a phase focusing method or a contrast focusing method can be used to focus the photographing device.

[0134] For example, when the contrast focusing method is used, an outer rectangular frame of the human eye of the preset target person can be considered as a focusing area, and then a contrast is calculated according to pixels of the focusing area, and a moving direction of a lens of the photographing device is determined according to the calculated contrast. When a position with the maximum contrast is found in repeated movement, the focusing of the lens of the photographing device is completed.

[0135] For example, when the phase focusing method is used, an outer rectangular frame of the human eye of the preset target person can be considered as a focusing area, and then a phase difference is calculated according to PDAF pixel points corresponding to the focusing area, and a deflection value of focusing is determined according to the phase difference, and the lens of the photographing device is moved according to the deflection value to realize focusing.

[0136] The specific steps of steps 502-514 can refer to the embodiments of Figures 1 to 4 .

[0137] It should be understood that, although each step in the flowchart involved in each of the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.

[0138] Based on the same inventive concept, the present application also provides an automatic focusing device of a photographing device for implementing the automatic focusing method of the photographing device involved above.

[0139] In some embodiments, as shown in Figure 6 , an automatic focusing device of a photographing device is provided, comprising: a target tracking module 602, a face detection area determination module 604, a face detection module 606, and a focusing module 608, wherein:

[0140] The target tracking module 602 is configured to track a preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in a current video frame.

[0141] The face detection region determination module 604 is configured to determine a face detection region in the current video frame according to the region position and the first size.

[0142] The face detection module 606 is configured to perform face detection on the face detection region to obtain a detection result.

[0143] The focusing module 608 is configured to select a corresponding focusing parameter according to the detection result, and perform focusing on a shooting device that collects the current video frame based on the focusing parameter; when the detection result is different, the selected focusing parameter is different.

[0144] In some embodiments, when the current video frame contains multiple persons, the automatic focusing device of the shooting device further includes:

[0145] The center position determination module is configured to determine a center position of the first video frame.

[0146] The first coverage region detection module is configured to detect coverage regions of the tracking targets corresponding to the persons in the first video frame.

[0147] The center point calculation module is configured to calculate center points of the coverage regions.

[0148] The distance calculation module is configured to calculate distances between the center position and the center points corresponding to the tracking targets.

[0149] The first marking module is configured to mark a person corresponding to a minimum distance in the multiple distances as a preset target person.

[0150] In some embodiments, the automatic focusing device of the shooting device further includes but is not limited to:

[0151] The second marking module is configured to detect coverage regions of the tracking targets corresponding to the persons in the first video frame; and mark a person corresponding to a coverage region with a maximum first size among the coverage regions as a preset target person.

[0152] The third marking module is configured to mark a person specified by a target selection instruction in the first video frame as a preset target person in response to the input target selection instruction.

[0153] In some embodiments, the focusing parameter is a parameter for indicating focusing by using a head feature, a face feature, or a human eye feature, and the focusing module 608 includes:

[0154] The focusing unit is configured to select a focusing parameter for focusing by using a head feature, a face feature, or a human eye feature based on the detection result, and perform focusing on a shooting device that collects the current video frame according to the head feature, the face feature, or the human eye feature.

[0155] In some embodiments, the focusing unit includes:

[0156] The first focusing sub-unit is configured to focus on the shooting device for collecting the current video frame according to the head feature when the face feature of the preset target person is not included in the face detection region.

[0157] The second focusing sub-unit is configured to detect whether the eye feature is included in the face feature when the face feature of the preset target person is included in the face detection region; if not, focus on the shooting device for collecting the current video frame according to the face feature; if yes, focus on the shooting device for collecting the current video frame according to the eye feature in the face feature.

[0158] In some embodiments, the automatic focusing device of the shooting device further comprises:

[0159] The adjusting module is configured to adjust the shooting device according to the preset target person when the preset target person is not in the preset position of the current video frame, so as to make the preset target person in the preset position of the current video frame.

[0160] In some embodiments, the face detection region determining module 604 comprises:

[0161] The first determining unit is configured to, when the tracking target is the head or the head-shoulder part of the preset target person, select the head region or the head-shoulder region corresponding to the region position and the first size in the current video frame as the face detection region.

[0162] The second determining unit is configured to, when the tracking target is the human body of the preset target person, select a target region according to the region position in the current video frame, and take the target region as the face detection region; wherein the first size is greater than a second size of the target region.

[0163] The above-mentioned various modules in the automatic focusing device of the shooting device can be all or partially realized by software, hardware and combinations thereof. The above-mentioned various modules can be embedded in or independent of the processor in the electronic device in hardware form, or can be stored in the memory in the electronic device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned various modules.

[0164] In some embodiments, an electronic device is provided, which can be a terminal, and the internal structure diagram thereof can be as shown in Figure 7The electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The communication interface, the input device and the display screen of the electronic device are connected with the system bus through an I / O interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement an automatic focusing method of a shooting device. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.

[0165] Those skilled in the art can understand that, Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0166] In some embodiments, an electronic device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the following steps:

[0167] Target tracking is performed on the preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in the current video frame; a face detection region is determined in the current video frame according to the region position and the first size; face detection is performed on the face detection region to obtain a detection result; corresponding focusing parameters are selected according to the detection result, and a shooting device collecting the current video frame is focused based on the focusing parameters; wherein when the detection results are different, the selected focusing parameters are different.

[0168] In some embodiments, the processor executing the computer program further implements the following steps:

[0169] A center position of the first video frame is determined; a covered region of each tracking target in the first video frame is detected; a center point of the covered region is calculated; a distance between the center position and the center point corresponding to each tracking target is calculated; and a person corresponding to the minimum distance in the plurality of distances is marked as a preset target person.

[0170] In some embodiments, the processor, when executing the computer program, also implements the following steps:

[0171] Detecting an overlapping area of the tracking target corresponding to each character in the first video frame; marking the character corresponding to the overlapping area with the largest first size as the preset target character; or, in response to an input target selection instruction, marking the character designated by the target selection instruction in the first video frame as the preset target character.

[0172] In some embodiments, the processor, when executing the computer program, also implements the following steps:

[0173] Based on the detection result, selecting a focusing parameter for focusing using the head feature, the face feature, or the human eye feature, and focusing the shooting device collecting the current video frame according to the head feature, the face feature, or the human eye feature.

[0174] In some embodiments, the processor, when executing the computer program, also implements the following steps:

[0175] When the face detection area does not contain the face feature of the preset target character, focusing the shooting device collecting the current video frame according to the head feature; when the face detection area contains the face feature of the preset target character, detecting whether the face feature includes the human eye feature; if not, focusing the shooting device collecting the current video frame according to the face feature; if yes, focusing the shooting device collecting the current video frame according to the human eye feature in the face feature.

[0176] In some embodiments, the processor, when executing the computer program, also implements the following steps:

[0177] When the preset target character is not in the preset position of the current video frame, adjusting the shooting device according to the preset target character to make the preset target character in the preset position of the current video frame.

[0178] In some embodiments, the processor, when executing the computer program, also implements the following steps:

[0179] When the tracking target is the head or the shoulder part of the preset target character, in the current video frame, the head region or the shoulder region corresponding to the region position and the first size is taken as the face detection area; when the tracking target is the human body of the preset target character, in the current video frame, a target region is selected according to the region position, and the target region is taken as the face detection area; wherein the first size is larger than a second size of the target region.

[0180] In some embodiments, a computer readable storage medium is provided, which stores a computer program, and the computer program, when executed by a processor, implements the following steps:

[0181] Tracking the preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in the current video frame; determining a face detection region in the current video frame according to the region position and the first size; performing face detection on the face detection region to obtain a detection result; selecting a corresponding focusing parameter according to the detection result, and focusing a shooting device collecting the current video frame based on the focusing parameter; wherein the selected focusing parameter is different when the detection result is different.

[0182] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0183] Determining a center position of the first video frame; detecting a coverage region of each tracking target in the first video frame; calculating a center point of the coverage region; calculating a distance between the center position and the center point corresponding to each tracking target; and marking a person corresponding to a minimum distance in the plurality of distances as the preset target person.

[0184] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0185] Detecting a coverage region of each tracking target in the first video frame; marking a person corresponding to a coverage region with a maximum first size in the coverage regions as the preset target person; or, in response to an input target selection instruction, marking a person designated by the target selection instruction in the first video frame as the preset target person.

[0186] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0187] Selecting a focusing parameter for focusing using head features, face features, or human eye features based on the detection result, and focusing the shooting device collecting the current video frame according to the head features, the face features, or the human eye features.

[0188] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0189] When the face detection region does not contain face features of the preset target person, focusing the shooting device collecting the current video frame according to the head features; when the face detection region contains face features of the preset target person, detecting whether the face features include human eye features; if not, focusing the shooting device collecting the current video frame according to the face features; if yes, focusing the shooting device collecting the current video frame according to the human eye features in the face features.

[0190] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0191] When the preset target person is not at the preset position in the current video frame, the shooting device is adjusted according to the preset target person, so that the preset target person is at the preset position in the current video frame.

[0192] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0193] When the tracking target is the head or the head-shoulder part of the preset target person, in the current video frame, a head region or a head-shoulder region corresponding to the region position and the first size is selected as the face detection region; when the tracking target is the body of the preset target person, in the current video frame, a target region is selected according to the region position, and the target region is selected as the face detection region; wherein the first size is greater than a second size of the target region.

[0194] In one embodiment, a computer program product is provided, comprising a computer program which, when executed by the processor, implements the following steps:

[0195] The target tracking is performed on the preset target person to obtain a region position and a first size of a tracking target corresponding to the preset target person in the current video frame; a face detection region is determined in the current video frame according to the region position and the first size; face detection is performed on the face detection region to obtain a detection result; a corresponding focusing parameter is selected according to the detection result, and a shooting device for collecting the current video frame is focused based on the focusing parameter; wherein when the detection result is different, the selected focusing parameter is different.

[0196] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0197] The center position of the first video frame is determined; the coverage region of each tracking target corresponding to each person in the first video frame is detected; the center point of the coverage region is calculated; the distance between the center position and the center point corresponding to each tracking target is calculated; and the person corresponding to the minimum distance in the plurality of distances is marked as the preset target person.

[0198] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0199] The coverage region of each tracking target corresponding to each person in the first video frame is detected; the person corresponding to the coverage region with the largest first size in each coverage region is marked as the preset target person; or, in response to an input target selection instruction, the person designated by the target selection instruction in the first video frame is marked as the preset target person.

[0200] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0201] The focusing parameter for focusing by using the head feature, the face feature or the eye feature is selected based on the detection result, and the shooting device for collecting the current video frame is focused according to the head feature, the face feature or the eye feature.

[0202] In one embodiment, the computer program, when executed by the processor, further implements the following steps: when the face detection region does not contain the face feature of the preset target person, focusing the shooting device for collecting the current video frame according to the head feature; when the face detection region contains the face feature of the preset target person, detecting whether the face feature contains the eye feature; if not, focusing the shooting device for collecting the current video frame according to the face feature; if yes, focusing the shooting device for collecting the current video frame according to the eye feature in the face feature.

[0203] In one embodiment, the computer program, when executed by the processor, further implements the following steps: when the preset target person is not in the preset position of the current video frame, adjusting the shooting device according to the preset target person to make the preset target person in the preset position of the current video frame.

[0204] In one embodiment, the computer program, when executed by the processor, further implements the following steps: when the tracking target is the head or the shoulder part of the preset target person, in the current video frame, the head region or the shoulder region corresponding to the region position and the first size is taken as the face detection region; when the tracking target is the body of the preset target person, in the current video frame, a target region is selected according to the region position, and the target region is taken as the face detection region; wherein the first size is greater than a second size of the target region.

[0205] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. The non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. The volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0206] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0207] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be noted that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An autofocusing method of an image pickup apparatus, characterized by, The method includes: The target is tracked to obtain the region position and first size of the tracking target corresponding to the target in the current video frame. The tracking target includes any one of the body, head or head and shoulder parts of the target. Based on the region location and the first size, a face detection region is determined in the current video frame; Face detection is performed on the face detection area to obtain the detection results; Based on the detection results, the corresponding focus parameters are selected, and the shooting device that captures the current video frame is focused based on the focus parameters; wherein, when the detection results are different, the selected focus parameters are different.

2. The method of claim 1, wherein, Determining the face detection region in the current video frame based on the region location and the first size includes: If the tracking target corresponding to the preset target person is the head, then the area corresponding to the bounding rectangle of the head is taken as the face detection area. If the tracking target corresponding to the preset target person is the head and shoulders, then the area corresponding to the bounding rectangle of the head and shoulders is taken as the face detection area; If the target being tracked is a human body, then the area corresponding to one-third or one-quarter of the outer rectangle of the human body is taken as the face detection area.

3. The method of claim 1, wherein, The detection results include: a first detection result and a second detection result, wherein the first detection result indicates that facial features exist in the face detection area, and the second detection result indicates that facial features do not exist in the face detection area.

4. The method of claim 1, wherein, The current video frame contains multiple people; before tracking the preset target person, the method further includes: Determine the center position of the first video frame; Detect the coverage area of ​​the tracking target corresponding to each person in the first video frame; Calculate the center point of the coverage area; Calculate the distance between the center position and the center point corresponding to each of the tracked targets; The person corresponding to the smallest distance among the multiple distances is marked as the preset target person.

5. The method of claim 1, wherein, The current video frame contains multiple people; before tracking the preset target person, the method further includes: Detect the coverage area of ​​the tracking target corresponding to each of the aforementioned individuals in the first video frame; mark the individual corresponding to the largest coverage area in each of the aforementioned coverage areas as the preset target individual; or... In response to the input target selection command, the person specified by the target selection command in the first video frame is marked as the preset target person.

6. The method of claim 1, wherein, The focusing parameters are parameters used to represent focusing using head features, facial features, or human eye features; The step of selecting appropriate focus parameters based on the detection results and focusing the shooting device acquiring the current video frame based on the focus parameters includes: Based on the detection results, focus parameters that utilize head features, facial features, or human eye features are selected for focusing, and the shooting device that captures the current video frame is focused according to the head features, facial features, or human eye features.

7. The method of claim 6, wherein, The focusing on a shooting device collecting the current video frame according to the head feature, the face feature or the human eye feature comprises: when the face feature of the preset target person is not included in the face detection region, focusing on the shooting device collecting the current video frame according to the head feature; when the face feature of the preset target person is included in the face detection region, detecting whether the human eye feature is included in the face feature; if not, focusing on the shooting device collecting the current video frame according to the face feature; if yes, focusing on the shooting device collecting the current video frame according to the human eye feature in the face feature.

8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: when the preset target person is not in the preset position of the current video frame, adjusting the shooting device according to the preset target person to make the preset target person in the preset position of the current video frame.

9. The method according to any one of claims 1 to 7, characterized in that, The determining of the face detection region in the current video frame according to the region position and the first size comprises: when the tracking target is the head or the head-shoulder part of the preset target person, in the current video frame, the head region or the head-shoulder region corresponding to the region position and the first size is taken as the face detection region; when the tracking target is the human body of the preset target person, in the current video frame, a target region is selected according to the region position, and the target region is taken as the face detection region; wherein the first size is greater than a second size of the target region.

10. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to realize the steps of the method in any one of claims 1 to 9.