Shooting method and device, computer readable medium and electronic equipment

By recognizing preset gestures and generating shutter drive signals using the camera lens, the problem of increased costs associated with external remote control devices is solved, enabling self-portraits without the need for remote control devices, thus improving user experience and the accuracy of gesture recognition.

CN120957007APending Publication Date: 2025-11-14GODOX PHOTO EQUIPMENT CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511186633.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, an external remote control device is required when taking selfies or group photos, which increases the shooting cost, and users may easily forget to bring the remote control device, making it impossible to take photos.

Method used

The system captures the viewfinder image using a camera, identifies object information and detects preset gestures, defines a gesture recognition area, and generates a shutter drive signal based on gesture changes to perform the shooting operation, simplifying the process and improving the user experience.

Benefits of technology

Shooting can be performed without carrying a special remote control, simplifying the process, reducing the probability of misoperation, improving the accuracy and success rate of gesture recognition, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120957007A_ABST
    Figure CN120957007A_ABST
Patent Text Reader

Abstract

The invention discloses a shooting method and device, a computer readable medium and electronic equipment, and the method comprises the steps: collecting a camera view-finding picture through a camera of a camera, and carrying out the detection of the camera view-finding picture, so as to recognize the object information in the camera view-finding picture; if it is detected that the object information comprises a first preset gesture, determining a gesture recognition area in a camera view-finding picture according to the position of the first preset gesture; performing gesture detection on the gesture recognition area, and generating a shutter driving signal when detecting that a gesture in the gesture recognition area is changed from a first preset gesture to a second preset gesture; and executing shooting operation according to the shutter driving signal. According to the technical scheme, the recognition area is delimited through the first preset gesture, then shooting is triggered through the second preset gesture, the two-step operation logic of first positioning and then confirmation is formed, and the misoperation probability is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of camera technology, and specifically relates to a shooting method, apparatus, computer-readable medium, and electronic device. Background Technology

[0002] Photography allows for the creation of long-term records of fleeting events, making it an essential part of daily life and work. When taking selfies or group photos, the camera needs to be fixed in place, the user needs to be positioned, and then the shutter is controlled remotely. This method requires an external remote control device, increasing the cost. Furthermore, in practice, users often forget to bring the remote control, making it impossible to take a picture. Summary of the Invention

[0003] The purpose of this application is to provide a shooting method, apparatus, computer-readable medium, and electronic device to simplify the shooting process and improve the user experience.

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0005] According to one aspect of the embodiments of this application, a shooting method is provided, including: The camera captures the viewfinder image, and the viewfinder image is then detected to identify object information within it. If the object information is detected to include a first preset gesture, then the gesture recognition area in the camera viewfinder is determined based on the location of the first preset gesture. Gesture detection is performed on the gesture recognition area, and when the gesture in the gesture recognition area changes from the first preset gesture to the second preset gesture, a shutter drive signal is generated; The shutter is driven by the shutter drive signal to perform the shooting operation.

[0006] According to one aspect of the embodiments of this application, a shooting device is provided, comprising: The image acquisition module is used to capture the viewfinder image from the camera's webcam. The first detection module is used to detect the viewfinder image of the camera in order to identify object information in the viewfinder image of the camera. The gesture region determination module is used to determine the gesture recognition region in the camera viewfinder based on the location of the first preset gesture if the object information is detected to include a first preset gesture. The second detection module is used to perform gesture detection on the gesture recognition area, and generate a shutter drive signal when the gesture in the gesture recognition area changes from the first preset gesture to the second preset gesture. The shooting module is used to drive the camera shutter to perform the shooting operation according to the shutter drive signal.

[0007] In one embodiment of this application, the device further includes a gesture comparison module, which is specifically used for: Extract at least one object gesture from the object information; Each object gesture is compared with a standard gesture, and the object gesture that meets the matching requirement with the standard gesture is taken as the first preset gesture.

[0008] In one embodiment of this application, the gesture comparison module is specifically used for: Obtain the matching degree between each object's gesture and the standard gesture, and compare the matching degree with a preset matching threshold; When there are at least two object gestures that meet the matching degree higher than the preset matching threshold, the area parameter corresponding to the object gesture is calculated based on the area of ​​the object gesture in the camera viewfinder based on the matching degree higher than the preset matching threshold. The gesture corresponding to the largest area parameter is taken as the first preset gesture.

[0009] In one embodiment of this application, the gesture comparison module is specifically used for: Obtain the matching degree between each object gesture and the standard gesture, and generate a first score for each object gesture based on the matching degree between each object gesture and the standard gesture; A second score is generated for each object gesture based on the area of ​​each object gesture in the camera viewfinder; The first score and the second score are merged to generate a target score for each object's gesture; The gesture corresponding to the maximum target score is used as the first preset gesture.

[0010] In one embodiment of this application, the first detection module is specifically used for: Detect object information in the camera viewfinder; If the object information determines that the camera viewfinder includes at least one head facing the camera lens, then the first preset gesture is detected based on the object information corresponding to the at least one head facing the camera lens.

[0011] In one embodiment of this application, the gesture region determination module is specifically used for: Determine the center point position of the first preset gesture; Using the first center point as the center point of the gesture recognition area, an image area in the camera view that completely includes the first preset gesture is selected as the gesture recognition area.

[0012] In one embodiment of this application, the second detection module is specifically used for: The position of the gesture in the gesture recognition area is tracked; If a change in the gesture position is detected relative to the position of the first preset gesture, the gesture recognition area is updated according to the changed gesture position.

[0013] In one embodiment of this application, the second detection module is specifically used for: In a series of camera viewfinder frames, gestures in the gesture recognition area are detected, and it is determined whether the similarity of gestures in two adjacent camera viewfinder frames is within a preset range. If it is within the preset range, it is determined whether the gesture changes from the first preset gesture to the second preset gesture. If it changes to the second preset gesture, a shutter drive signal is generated.

[0014] In one embodiment of this application, the second detection module is further configured to: Timing begins from the moment the first preset gesture is detected. It is determined whether the duration of the timing is less than the first preset duration. If it is less than the first preset duration, the detection of gestures in the gesture recognition area continues. If it is not less than the first preset duration, the detection of gestures in the gesture recognition area stops.

[0015] In one embodiment of this application, the second detection module is further configured to: Determine the duration of the change from the first preset gesture to the second preset gesture; If the duration of the change is greater than the second preset duration, a shutter drive signal is generated.

[0016] In one embodiment of this application, the shooting module is specifically used for: The camera shutter is driven to perform a shooting operation according to the preset shooting conditions and the shutter drive signal; wherein, the preset shooting conditions include at least one of a delayed shooting time and a preset number of continuous shooting.

[0017] In one embodiment of this application, the first detection module is specifically used for: Based on the input image size requirements of the detection model, the size of the camera viewfinder is adjusted to obtain a standardized viewfinder image; The standardized image is detected using the detection model to identify object information in the camera viewfinder.

[0018] In one embodiment of this application, the second preset gesture includes a gesture sequence formed by multiple preset gestures.

[0019] According to one aspect of the embodiments of this application, a computer-readable medium is provided, on which a computer program is stored, which, when executed by a processor, implements the shooting method as described in the above technical solutions.

[0020] According to one aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor executes the executable instructions to cause the electronic device to perform the shooting method as described in the above technical solution.

[0021] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the imaging method as described in the above technical solutions.

[0022] In the technical solution provided in this application embodiment, the camera's viewfinder is first captured by the camera lens, and the viewfinder is detected to identify object information in the viewfinder. Then, if the detected object information includes a first preset gesture, the gesture recognition area in the viewfinder is determined based on the location of the first preset gesture. Finally, gesture detection is performed on the gesture recognition area, and when the gesture in the gesture recognition area changes from the first preset gesture to a second preset gesture, a shutter drive signal is generated. The shutter drive signal drives the camera's shutter to perform the shooting operation. Thus, users can perform shooting operations through gesture control, eliminating the need to carry a dedicated remote control device to control the camera, thereby simplifying the shooting process and improving the user experience. Furthermore, by first defining the recognition area with the first preset gesture and then triggering shooting with the second preset gesture, a two-step operation logic of positioning and confirmation is formed, reducing the probability of misoperation. Also, after determining the gesture recognition area with the first preset gesture, subsequent detection is only performed on that area, filtering out interference from other irrelevant objects in the image, significantly improving the success rate and accuracy of recognizing the second preset gesture.

[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0025] Figure 1 A flowchart illustrating an embodiment of the shooting method provided in this application is shown.

[0026] Figure 2 A flowchart illustrating an embodiment of the shooting method provided in this application is shown.

[0027] Figure 3 A schematic block diagram of the imaging device provided in an embodiment of this application is shown.

[0028] Figure 4 A schematic diagram of a computer system architecture suitable for implementing the embodiments of this application is shown. Detailed Implementation

[0029] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0030] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0031] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0033] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0034] The shooting method provided in this application will be described in detail below with reference to specific implementation methods.

[0035] Figure 1 A flowchart illustrating a photographing method provided in one embodiment of this application is shown schematically. This method can be implemented by a dedicated camera or an electronic device with camera functionality, such as an instant camera, mobile phone, or tablet computer. The execution process of the method is described below using a camera as the executing entity. Figure 1 As shown, the shooting method provided in this embodiment includes steps 110 to 140, as detailed below: Step 110: Capture the viewfinder image using the camera's webcam, and then detect the viewfinder image to identify the object information within it.

[0036] Specifically, camera viewfinder footage refers to the image captured by the camera when a device with shooting capabilities takes a picture. "Camera" can be any device capable of shooting, such as a dedicated camera, mobile phone, tablet, or wearable device with an integrated camera. The captured camera viewfinder footage can be continuous video frame data, with each frame representing a viewfinder image.

[0037] A camera viewfinder can include various objects, such as people, animals, plants, and buildings. Computer vision algorithms can be used to detect and identify all possible objects in the viewfinder, and output the object's category and location coordinates. The location coordinates can be represented by the coordinates of the detection box that selects the object. For example, if an animal is selected by a rectangle, then the animal's location coordinates are the coordinates of the rectangle, such as the top-left and bottom-right pixel coordinates of the rectangle, or the coordinates of the rectangle's center point.

[0038] In one embodiment of this application, step 110 specifically includes: adjusting the size of the camera viewfinder according to a preset input image size requirement to obtain a standardized viewfinder image; and detecting the standardized image to identify object information in the camera viewfinder.

[0039] The preset image input size can be the size requirement of the input data for the trained object detection model. The camera viewfinder is resized, such as by enlarging, reducing, or cropping, to obtain a standardized viewfinder image. Detection is then performed based on this standardized viewfinder image; for example, the standardized viewfinder image is input into the detection model, and the detection model outputs object information from the camera viewfinder. Step 120: If the detected object information includes a first preset gesture, the gesture recognition area in the camera viewfinder is determined based on the location of the first preset gesture.

[0040] Specifically, the first preset gesture is a gesture that indicates the user is ready to shoot, such as a V-sign, an open palm, or a pointing gesture. Gestures are usually human body movements, so the corresponding object information belongs to portrait information. The detection of portrait information includes whether the first preset gesture is included. If the first preset gesture is detected, it means that the user needs to perform a shooting operation, which is equivalent to starting the shooting operation.

[0041] Optionally, a neural network model can be used to detect the first preset gesture. Specifically, this can be done by extracting features such as the shape and skeletal key points of the gesture to determine whether the gesture made by the object is the first preset gesture.

[0042] The specific execution of the shooting operation is still determined by the object corresponding to the first preset gesture. Therefore, the gesture recognition area in the camera viewfinder is determined according to the location of the first preset gesture, so as to further determine when to perform the shooting operation based on the gesture recognition area.

[0043] In one embodiment of this application, the detection process of the first preset gesture includes: extracting at least one object gesture from the object information; comparing each object gesture with a standard gesture respectively, and taking the object gesture that meets the matching requirement with the standard gesture as the first preset gesture.

[0044] In other words, firstly, all object gestures are extracted from the object information, including all gestures detected in all portrait information. Then, these object gestures are compared with standard gestures, which are pre-set gestures in the camera used to trigger area delineation. If an object gesture meets the matching requirement with the standard gesture, then the object gesture that meets the requirement is the first preset gesture.

[0045] In one embodiment of this application, the similarity between each object gesture and a standard gesture can be calculated, and this similarity is the matching degree between the two. For example, the features of the object gesture and the features of the standard gesture are extracted separately, and then the cosine similarity between these two features is calculated to obtain the matching degree between the two.

[0046] In one embodiment of this application, it can be first detected whether the matching degree corresponding to each object gesture is greater than the matching threshold. If the matching degree corresponding to all object gestures is not greater than the matching threshold, it means that the first preset gesture is not included in the viewfinder. If there is an object gesture whose matching degree is greater than the matching threshold, then that object gesture can be used as the first preset gesture. If there are multiple object gestures whose matching degree is greater than the matching threshold, then the object gesture with the highest matching degree can be used as the first preset gesture.

[0047] In one embodiment of this application, the process of determining the first preset gesture may further include: obtaining the matching degree between each object gesture and the standard gesture, and comparing the matching degree with a preset matching threshold; when the number of object gestures with a matching degree higher than the preset matching threshold is at least two, calculating the area parameter corresponding to the object gesture based on the area of ​​the object gesture with a matching degree higher than the preset matching threshold in the camera viewfinder; and taking the object gesture corresponding to the largest area parameter as the first preset gesture.

[0048] In other words, when there are two or more object gestures with a matching degree higher than a preset matching threshold, the gesture with the largest area parameter in the camera viewfinder is selected as the first preset gesture. This can be achieved by directly using the area of ​​the object gesture in the camera viewfinder, or by using the proportion of the object gesture's area in the camera viewfinder. Using the object gesture with the largest area parameter as the first preset gesture increases the gesture recognition area defined by the first preset gesture, thereby improving the accuracy of subsequent second preset gesture recognition.

[0049] In one embodiment of this application, the process of determining the first preset gesture may further include: obtaining the matching degree between each object gesture and the standard gesture, and generating a first score for each object gesture based on the matching degree between each object gesture and the standard gesture; generating a second score for each object gesture based on the area of ​​each object gesture in the camera viewfinder; fusing the first score and the second score to generate a target score for each object gesture; and taking the object gesture corresponding to the maximum target score as the first preset gesture.

[0050] In other words, each object gesture is assigned a corresponding score based on its matching degree with the standard gesture and the gesture area, respectively. Thus, an object gesture corresponds to two types of scores: a first score based on the matching degree with the standard gesture and a second score based on the gesture area. The second score is generated based on the area, either directly based on the area size or based on the area percentage.

[0051] When setting scores, both the first and second scores are unified to the same range, for example, both scores can range from [0, 10]. The scores are then set based on the positive correlation between the metrics (matching degree, area) and the scores; that is, the higher the matching degree, the higher the first score, and the larger the gesture area, the higher the second score. Optionally, a mapping relationship between the metrics (matching degree, area) and scores can be predefined, and the scores can be set according to this mapping relationship. For example, to extract object gestures with a matching degree higher than 0.7 (maximum value is 1) with the standard gesture, assuming the weights for matching degree and area percentage are 1:5: If object gesture A has a matching degree of 0.8 and an area percentage of 0.1, its total score is 0.8 + 0.1 × 5 = 1.3. If object gesture B has a matching degree of 0.75 and an area percentage of 0.15, its total score is 0.75 + 0.15 × 5 = 1.5. Since object gesture B has a higher total score, it is used as the first preset gesture.

[0052] The two scores are then combined to obtain the target score for the object's gesture. For example, the two scores can be added directly. Alternatively, weights can be set for the matching degree and area respectively, and then the two scores can be weighted and summed to obtain the target score.

[0053] Finally, the gesture corresponding to the maximum target score can be taken as the first preset gesture.

[0054] In one embodiment of this application, when defining the gesture recognition area, the center point of the first preset gesture can be determined first. Then, using the first center point as the center point of the gesture recognition area, an image area in the camera viewfinder that completely includes the first preset gesture is selected as the gesture recognition area. The shape of the gesture recognition area can be set according to actual needs, such as a circle, rectangle, or palm shape. Expanding outwards from the center point of the first preset gesture, a region that completely includes the first preset gesture image is selected as the gesture recognition area. Therefore, the area of ​​the gesture recognition area will be greater than or equal to the total area of ​​the first preset gesture, and all coordinates of the first preset gesture will be within the gesture recognition area.

[0055] Preferably, the area of ​​the gesture recognition region can be set to be greater than the total area of ​​the first preset gesture, such as 1.5 times or 2 times the total area of ​​the first preset gesture, so that the second preset gesture can still be accurately detected when the user's gesture position changes to a certain extent.

[0056] In one embodiment of this application, step 110 specifically includes: detecting object information in the camera viewfinder; if it is determined from the object information that the camera viewfinder includes at least one head facing the camera lens, then detecting a first preset gesture based on the object information corresponding to the at least one head facing the camera lens.

[0057] Specifically, the system first detects object information within the camera's viewfinder and extracts human image information from this object information. Human image information includes all information about a person within the camera's viewfinder, such as the face, body, and limbs. From this human image information, the system identifies the head in the camera's viewfinder. If at least one head is facing the camera lens, it indicates that the user needs to take a photo, thus triggering the detection of the first preset gesture. This prevents misidentification of gestures from other objects as the first preset gesture, improving the accuracy of gesture-based shooting control.

[0058] Optionally, determining whether an avatar is facing the camera lens can be done by recognizing the completeness of the facial features in the avatar. When a user is not facing the camera, the camera viewfinder may only contain partial facial information. For example, if a user's head is turned to the left, the viewfinder may only show the right half of their facial features. Therefore, detecting the completeness of facial features can determine whether the user is facing the camera. Optionally, eye feature points can also be extracted using a model, and the user's face can be analyzed based on these feature points. For example, visible eyes are considered as facing the camera.

[0059] Step 130: Perform gesture detection on the gesture recognition area, and generate a shutter drive signal when the gesture in the gesture recognition area changes from the first preset gesture to the second preset gesture.

[0060] Step 140: Drive the camera shutter to perform the shooting operation according to the shutter drive signal.

[0061] Specifically, the second preset gesture is the shutter control gesture. Gesture detection within the gesture recognition area is performed to detect whether the user's gesture changes from the initial gesture to the shutter control gesture, that is, from the first preset gesture to the second preset gesture. When the gesture changes to the second preset gesture, a shutter drive signal is generated. This signal drives the camera's shutter to perform a shooting operation, which can be taking a photo or recording a video. Because shutter gesture detection is performed only within the gesture recognition area, compared to full-image detection, the computational load is reduced, real-time performance is improved, and detection efficiency is higher.

[0062] In one embodiment of this application, if the gesture does not change from the first preset gesture to the second preset gesture, such as changing from the first preset gesture to other gestures, it means that no shutter signal has been triggered. Here, we return to step 110 and continue to detect the viewfinder.

[0063] In one embodiment of this application, it is necessary to detect the continuous change process from the first preset gesture to the second preset gesture in order to determine that the gesture of the same object changes from the first preset gesture to the second preset gesture, thereby avoiding the accidental triggering of shooting operation caused by one person making the first preset gesture and another person making the second preset gesture.

[0064] In the technical solution provided in this application embodiment, the camera's viewfinder is first captured by the camera lens, and the viewfinder is detected to identify object information in the viewfinder. Then, if the detected object information includes a first preset gesture, the gesture recognition area in the viewfinder is determined based on the location of the first preset gesture. Finally, gesture detection is performed on the gesture recognition area, and when the gesture in the gesture recognition area changes from the first preset gesture to a second preset gesture, a shutter drive signal is generated. The shutter drive signal drives the camera's shutter to perform the shooting operation. Thus, users can perform shooting operations through gesture control, eliminating the need to carry a dedicated remote control device to control the camera, thereby simplifying the shooting process and improving the user experience. Furthermore, by first defining the recognition area with the first preset gesture and then triggering shooting with the second preset gesture, a two-step operation logic of positioning and confirmation is formed, reducing the probability of misoperation. Also, after determining the gesture recognition area with the first preset gesture, subsequent detection is only performed on that area, filtering out interference from other irrelevant objects in the image, significantly improving the success rate and accuracy of recognizing the second preset gesture.

[0065] In one embodiment of this application, considering that the user's hand position may change significantly when making a gesture, such as the gesture exceeding the gesture recognition area, the gesture position within the gesture recognition area is tracked during gesture detection to differentiate gestures. If a change in the gesture position relative to the location of a first preset gesture is detected, the gesture recognition area is updated based on the changed gesture position. The change in gesture position relative to the location of the first preset gesture can be determined based on the distance between the center point of the gesture position and the center point of the first preset gesture. If the distance between the two center points is greater than a threshold, the gesture position is considered to have changed; otherwise, it is considered not to have changed. When the gesture position changes, the gesture recognition area needs to be adjusted. Here, the center point of the changed gesture position can be used as the center point of the updated gesture recognition area, and then the area containing the changed gesture is selected as the updated gesture recognition area by expanding outward from this center point. Thus, by dynamically adjusting the gesture recognition area according to changes in gesture position, erroneous recognition or failure to recognize the second preset gesture due to movement of the camera or the user's position can be prevented, thereby improving the accuracy and success rate of the second preset gesture recognition, and ultimately improving the accuracy of camera gesture control.

[0066] In one embodiment of this application, the detection process of the second preset gesture may be as follows: In a continuous camera viewfinder frame, gestures in the gesture recognition area are detected. It is determined whether the similarity of gestures in two adjacent camera viewfinder frames is within a preset range. If it is within the preset range, it is determined whether the gesture changes from a first preset gesture to a second preset gesture. If it changes to the second preset gesture, a shutter drive signal is generated. Wherein, if the similarity of gestures in two adjacent camera viewfinder frames is within the preset range, the gesture is considered to be continuously changing. Therefore, in the case of continuously changing gestures, detecting a change from the first preset gesture to the second preset gesture indicates that the first preset gesture changed continuously to the second preset gesture, rather than suddenly changing to the second preset gesture. This avoids simultaneously identifying gestures made by different objects as both the first and second preset gestures, thus preventing accidental photography caused by gestures from different objects.

[0067] In one embodiment of this application, during the gesture detection process, timing begins from the detection of a first preset gesture. The system continuously checks whether the timeout is less than the first preset duration. If it is less, gesture detection in the gesture recognition area continues; if it is not less than the first preset duration, gesture detection in the gesture recognition area stops. That is, if gesture detection continues for a period of time without triggering a photo capture, gesture detection ceases to avoid prolonged gesture detection and wasting camera energy.

[0068] In one embodiment of this application, during the gesture detection process, after detecting a second preset gesture, the duration of the change from the first preset gesture to the second preset gesture is determined. If the duration of the change is greater than the second preset duration, a shutter drive signal is generated. That is, the continuous change duration of the gesture cannot be too short. If it is too short, it may be a gesture made by different objects. By limiting the change duration of the gesture to within the second preset duration, accidental photography caused by gestures made by different objects can be avoided.

[0069] In one embodiment of this application, the second preset duration is less than the first preset duration. During gesture detection, timing begins from the detection of the first preset gesture. When the timing duration is less than the first preset duration, the second preset gesture is detected. Upon detection of the second preset gesture, it is determined whether the duration of the change from the first preset gesture to the second preset gesture is greater than the second preset duration. If it is greater, the gesture change meets the condition, and a shutter drive signal is generated to drive the shutter for shooting. That is, the change from the first preset gesture to the second preset gesture needs to be completed within a preset duration range (defined by the second preset duration to the first preset duration) (too fast or too slow is not acceptable) to prevent other non-control gestures from being identified as control gestures, thereby improving the accuracy of control gesture recognition. If the second preset gesture is not detected when the timing duration is not less than the first preset duration, the gesture detection operation is no longer performed.

[0070] In one embodiment of this application, when performing the shooting operation, the technical solution further includes: performing the shooting operation according to preset shooting conditions; wherein, the preset shooting conditions include at least one of delaying the shooting for a set duration and taking a preset number of consecutive shots. For example, extending the shooting time by 2 seconds, taking 3 consecutive shots, etc.

[0071] Optionally, the delay duration can be provided via voice. The camera can detect ambient sound signals, and when it detects delay-related information in the ambient sound signals, it performs the corresponding delay shooting operation. The delay-related information can be a direct specification of the delay duration, such as "Take a picture in 3 seconds," or a countdown prompt, such as "Three, two, one, cheese."

[0072] In one embodiment of this application, the second preset gesture includes a gesture sequence formed by multiple preset gestures. For example, the second preset gesture may be a sequence of raising three fingers, then raising two fingers, and finally raising one finger, embodying a "3-2-1" change. Optionally, the first preset gesture and the second preset gesture may be combined, or they may be a gesture sequence of multiple preset gesture processes. For example, the first preset gesture may be raising three fingers, and the second preset gesture may include raising two fingers first, and then raising one finger.

[0073] Figure 2 A flowchart illustrating an embodiment of the shooting method provided in this application is shown schematically. This embodiment is a further refinement of the above embodiment. Figure 2 As shown, the shooting method provided in this embodiment includes the following steps: Step 1: The image sensor acquires image information in real time. Specifically, the user first activates the camera's Selfie mode. In Selfie mode, the camera's camera and image sensor acquire image information in real time, and the acquired image information is the camera viewfinder in this application.

[0074] Step 2: Preprocess the image information to adapt it to the intelligent recognition model, and then import the preprocessed image information into the intelligent recognition model. Specifically, adjust the size parameters of the current image to meet the requirements of the intelligent recognition model, and then import the adjusted image into the intelligent recognition model. The intelligent recognition model is the detection model in this application.

[0075] Step 3: The intelligent recognition model identifies human image information in the image and detects whether a preset preparatory gesture is generated in the human image information.

[0076] Specifically, the process involves: first, identifying all human image information from the image, including identifying all object information and then extracting human image information from the object information. Next, detecting whether any human image contains a preset pre-defined hand gesture. This preset pre-defined hand gesture is the gesture for entering the shooting preparation state, which is the first preset gesture mentioned earlier. For example, an open palm facing the camera. Then, detecting whether there are multiple suspected pre-defined hand gestures. If multiple suspected pre-defined hand gestures exist, the gesture with the highest reliability (i.e., matching degree) is selected as the pre-defined hand gesture through identification and comparison (i.e., comparing the current gesture with a pre-set standard gesture (e.g., calculating similarity), and the gesture closest to the standard gesture is the pre-defined hand gesture with the highest reliability). Alternatively, the gesture with the largest area occupied by the suspected hand gesture can be selected as the pre-defined hand gesture. The reliability and area of ​​suspected hand gestures can also be detected simultaneously, with preset weight values ​​for reliability and area respectively. Combining the weight values ​​with the scores of reliability and area, the gesture with the highest score is defined as the pre-defined hand gesture. If no pre-defined hand gesture is detected, the next frame of the image is identified until a pre-defined hand gesture is detected.

[0077] Step 4: Obtain the coordinates of the preparatory gesture in the image, and generate an effective recognition area based on its coordinates.

[0078] Specifically, a valid recognition area (i.e., gesture recognition area) is generated with the center coordinates of the preparatory gesture (i.e., the center point of the first preset gesture) as the center point. The area of ​​the valid recognition area is larger than the total area of ​​the preparatory gesture, and all coordinates of the preparatory gesture are within the valid recognition area.

[0079] Step 5: Within the effective recognition area, monitor the gesture information in real time.

[0080] Specifically, within the effective recognition area, it is determined whether a preset shutter control gesture (i.e., a second preset gesture) exists. This involves performing gesture detection on the gesture recognition area to determine if the second preset gesture exists. For example, the second preset gesture is a fist facing the camera. Preferably, it determines whether the preparatory gesture gradually transforms into a shutter control gesture, i.e., from an open palm facing the camera to a fist facing the camera. If it transforms into a shutter control gesture, a shutter control signal (i.e., a shutter drive signal) is generated. If the preparatory gesture gradually transforms into other gestures, the process returns to step 3.

[0081] By combining steps 4 and 5—first defining an effective recognition area, then detecting continuous changes in gestures within that area—errors caused by external interference are effectively prevented, thus significantly improving the accuracy of shutter control gesture recognition. The process for detecting continuous gesture changes involves determining whether the similarity between the gesture in the next frame and the gesture in the previous frame is within a preset range in the continuously acquired camera viewfinder images. If it is within the preset range, the gesture is considered a continuous change.

[0082] Step 6: After receiving the shutter control signal, drive the shutter according to the preset shutter control conditions; Specifically: The shutter control signal is the shutter drive signal generated when a shutter control gesture is detected. The preset shutter control conditions are the preset shooting conditions, which may include: a delayed shooting time, such as 2 seconds, or a continuous shooting number, such as 3 shots; then the camera will drive the shutter 2 seconds after receiving the shutter control signal and take 3 shots in a row.

[0083] Step 7: The camera generates the corresponding photo based on the shutter action.

[0084] The technical solution of this application allows users to control the camera shutter by making predetermined preparatory gestures and shutter control gestures in front of the camera, enabling convenient self-portrait operation. No external equipment needs to be installed on the camera, which effectively reduces shooting costs and allows users to take selfies at any time through simple operation, thus improving the user experience.

[0085] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0086] The following describes an embodiment of the apparatus of this application, which can be used to perform the shooting method in the above embodiments of this application. Figure 3 A schematic block diagram of the imaging device provided in an embodiment of this application is shown. Figure 3 As shown, the imaging device provided in this application embodiment includes: Image acquisition module 310 is used to acquire the viewfinder image from the camera via the camera's webcam; The first detection module 320 is used to detect the viewfinder of the camera in order to identify object information in the viewfinder of the camera. The gesture region determination module 330 is used to determine the gesture recognition region in the camera viewfinder based on the location of the first preset gesture if the object information is detected to include a first preset gesture. The second detection module 340 is used to perform gesture detection on the gesture recognition area, and generate a shutter drive signal when the gesture in the gesture recognition area changes from the first preset gesture to the second preset gesture. The shooting module 350 is used to drive the camera shutter to perform a shooting operation according to the shutter drive signal.

[0087] In one embodiment of this application, the device further includes a gesture comparison module, which is specifically used for: Extract at least one object gesture from the object information; Each object gesture is compared with a standard gesture, and the object gesture that meets the matching requirement with the standard gesture is taken as the first preset gesture.

[0088] In one embodiment of this application, the gesture comparison module is specifically used for: Obtain the matching degree between each object's gesture and the standard gesture, and compare the matching degree with a preset matching threshold; When there are at least two object gestures that meet the matching degree higher than the preset matching threshold, the area parameter corresponding to the object gesture is calculated based on the area of ​​the object gesture in the camera viewfinder based on the matching degree higher than the preset matching threshold. The gesture corresponding to the largest area parameter is taken as the first preset gesture.

[0089] In one embodiment of this application, the gesture comparison module is specifically used for: Obtain the matching degree between each object gesture and the standard gesture, and generate a first score for each object gesture based on the matching degree between each object gesture and the standard gesture; A second score is generated for each object gesture based on the area of ​​each object gesture in the camera viewfinder; The first score and the second score are merged to generate a target score for each object's gesture; The gesture corresponding to the maximum target score is used as the first preset gesture.

[0090] In one embodiment of this application, the first detection module 320 is specifically used for: Detect object information in the camera viewfinder; If the object information determines that the camera viewfinder includes at least one head facing the camera lens, then the first preset gesture is detected based on the object information corresponding to the at least one head facing the camera lens.

[0091] In one embodiment of this application, the gesture region determination module 330 is specifically used for: Determine the center point position of the first preset gesture; Using the first center point as the center point of the gesture recognition area, an image area in the camera view that completely includes the first preset gesture is selected as the gesture recognition area.

[0092] In one embodiment of this application, the second detection module 340 is specifically used for: The position of the gesture in the gesture recognition area is tracked; If a change in the gesture position is detected relative to the position of the first preset gesture, the gesture recognition area is updated according to the changed gesture position.

[0093] In one embodiment of this application, the second detection module 340 is specifically used for: In a series of camera viewfinder frames, gestures in the gesture recognition area are detected, and it is determined whether the similarity of gestures in two adjacent camera viewfinder frames is within a preset range. If it is within the preset range, it is determined whether the gesture changes from the first preset gesture to the second preset gesture. If it changes to the second preset gesture, a shutter drive signal is generated.

[0094] In one embodiment of this application, the second detection module 340 is further configured to: Timing begins from the moment the first preset gesture is detected. It is determined whether the duration of the timing is less than the first preset duration. If it is less than the first preset duration, the detection of gestures in the gesture recognition area continues. If it is not less than the first preset duration, the detection of gestures in the gesture recognition area stops.

[0095] In one embodiment of this application, the second detection module 340 is further configured to: Determine the duration of the change from the first preset gesture to the second preset gesture; If the duration of the change is greater than the second preset duration, a shutter drive signal is generated.

[0096] In one embodiment of this application, the shooting module 350 is specifically used for: The camera shutter is driven to perform a shooting operation according to the preset shooting conditions and the shutter drive signal; wherein, the preset shooting conditions include at least one of a delayed shooting time and a preset number of continuous shooting.

[0097] In one embodiment of this application, the first detection module 320 is specifically used for: Based on the preset input image size requirements, the size of the camera viewfinder is adjusted to obtain a standardized viewfinder image; The standardized image is detected to identify object information in the camera viewfinder.

[0098] In one embodiment of this application, the second preset gesture includes a gesture sequence formed by multiple preset gestures.

[0099] The specific details of the imaging device provided in the various embodiments of this application have been described in detail in the corresponding method embodiments, and will not be repeated here.

[0100] Figure 4 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.

[0101] It should be noted that, Figure 4 The computer system 400 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0102] like Figure 4As shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 402 or programs loaded from storage section 408 into random access memory (RAM). The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output interface 405 (I / O interface) is also connected to the bus 404.

[0103] The following components are connected to the input / output interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a local area network card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 410 as needed so that computer programs read from it can be installed into the storage section 408 as needed.

[0104] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit 401, it performs various functions defined in the system of this application.

[0105] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0107] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0108] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0109] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0110] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A shooting method, characterized in that, include: The camera captures the viewfinder image, and the viewfinder image is then detected to identify object information within it. If the object information is detected to include a first preset gesture, then the gesture recognition area in the camera viewfinder is determined based on the location of the first preset gesture. Gesture detection is performed on the gesture recognition area, and when the gesture in the gesture recognition area changes from the first preset gesture to the second preset gesture, a shutter drive signal is generated; The shutter is driven by the shutter drive signal to perform the shooting operation.

2. The shooting method according to claim 1, characterized in that, After identifying object information in the camera viewfinder, the method further includes: Extract at least one object gesture from the object information; Each object gesture is compared with a standard gesture, and the object gesture that meets the matching requirements with the standard gesture is taken as the first preset gesture.

3. The shooting method according to claim 2, characterized in that, The step of comparing each object gesture with a standard gesture and selecting the object gesture that meets the matching requirement with the standard gesture as the first preset gesture includes: Obtain the matching degree between each object's gesture and the standard gesture, and compare the matching degree with a preset matching threshold; When there are at least two object gestures that meet the matching degree higher than the preset matching threshold, the area parameter corresponding to the object gesture is calculated based on the area of ​​the object gesture in the camera viewfinder based on the matching degree higher than the preset matching threshold. The gesture corresponding to the largest area parameter is taken as the first preset gesture.

4. The shooting method according to claim 2, characterized in that, The step of comparing each object gesture with a standard gesture and selecting the object gesture that meets the matching requirement with the standard gesture as the first preset gesture includes: Obtain the matching degree between each object gesture and the standard gesture, and generate a first score for each object gesture based on the matching degree between each object gesture and the standard gesture; A second score is generated for each object gesture based on the area of ​​each object gesture in the camera viewfinder; The first score and the second score are merged to generate a target score for each object's gesture; The gesture corresponding to the maximum target score is used as the first preset gesture.

5. The shooting method according to claim 1, characterized in that, The camera viewfinder image is inspected, including: Detect object information in the camera viewfinder; If the object information determines that the camera viewfinder includes at least one head facing the camera lens, then the first preset gesture is detected based on the object information corresponding to the at least one head facing the camera lens.

6. The shooting method according to claim 1, characterized in that, Determining the gesture recognition area in the camera viewfinder based on the location of the first preset gesture includes: Determine the center point position of the first preset gesture; Using the first center point as the center point of the gesture recognition area, an image area in the camera view that completely includes the first preset gesture is selected as the gesture recognition area.

7. The shooting method according to claim 1, characterized in that, During the gesture detection process in the gesture recognition area, the method further includes: The position of the gesture in the gesture recognition area is tracked; If a change in the gesture position is detected relative to the position of the first preset gesture, the gesture recognition area is updated according to the changed gesture position.

8. The shooting method according to claim 1, characterized in that, When a change in gesture in the gesture recognition area from the first preset gesture to the second preset gesture is detected, a shutter drive signal is generated, including: In a series of camera viewfinder frames, gestures in the gesture recognition area are detected, and it is determined whether the similarity of gestures in two adjacent camera viewfinder frames is within a preset range. If it is within the preset range, it is determined whether the gesture changes from the first preset gesture to the second preset gesture. If it changes to the second preset gesture, a shutter drive signal is generated.

9. The shooting method according to claim 8, characterized in that, Detecting gestures in the gesture recognition area further includes: Timing begins from the moment the first preset gesture is detected. It is determined whether the duration of the timing is less than the first preset duration. If it is less than the first preset duration, the detection of gestures in the gesture recognition area continues. If it is not less than the first preset duration, the detection of gestures in the gesture recognition area stops.

10. The shooting method according to claim 8, characterized in that, Before generating the shutter drive signal, the method further includes: Determine the duration of the change from the first preset gesture to the second preset gesture; If the duration of the change is greater than the second preset duration, a shutter drive signal is generated.

11. The shooting method according to claim 1, characterized in that, The camera shutter is driven to perform a shooting operation according to the shutter drive signal, including: The camera shutter is driven to perform a shooting operation according to the preset shooting conditions and the shutter drive signal; wherein, the preset shooting conditions include at least one of a delayed shooting time and a preset number of continuous shooting.

12. The shooting method according to claim 1, characterized in that, Detecting the camera viewfinder to identify object information within it, including: Based on the preset input image size requirements, the size of the camera viewfinder is adjusted to obtain a standardized viewfinder image; The standardized image is detected to identify object information in the camera viewfinder.

13. The shooting method according to any one of claims 1-12, characterized in that, The second preset gesture includes a gesture sequence formed by multiple preset gestures.

14. A shooting device, characterized in that, include: The image acquisition module is used to capture the viewfinder image from the camera's webcam. The first detection module is used to detect the viewfinder image of the camera in order to identify object information in the viewfinder image of the camera. The gesture region determination module is used to determine the gesture recognition region in the camera viewfinder based on the location of the first preset gesture if the object information is detected to include a first preset gesture. The second detection module is used to perform gesture detection on the gesture recognition area, and generate a shutter drive signal when the gesture in the gesture recognition area changes from the first preset gesture to the second preset gesture. The shooting module is used to drive the camera shutter to perform the shooting operation according to the shutter drive signal.

15. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the shooting method according to any one of claims 1-13.

16. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the executable instructions to implement the shooting method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Control method of intelligent device and intelligent device

    CN113473198A

  • Virtual photographing method and device based on gesture recognition, electronic equipment and medium

    CN114339039A

  • Shooting method and electronic equipment

    CN115484391A

  • Gesture shooting method and device based on multi-person scene

    CN117319791A

  • Gesture recognition device, gesture recognition system, gesture recognition method, and program

    JP2025092381A