Photographing method and device
Patent Information
- Application Number
- PCT/CN2025/144030
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2025-12-19
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025144030_01102026_PF_FP_ABST
Abstract
Description
A shooting method and equipment
[0001] This application claims priority to Chinese Patent Application No. 202510388202.6, filed with the State Intellectual Property Office of China on March 28, 2025, entitled “A Shooting Method and Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal shooting technology, and in particular to a shooting method and device. Background Technology
[0003] With the continuous iteration of camera functions on electronic devices, photography performance has improved significantly, and using cameras on electronic devices for taking pictures has become a common habit. For example, the shooting functions of electronic devices have become increasingly sophisticated, offering features like quick-capture functionality. Quick-capture functionality allows electronic devices to take a picture as quickly as possible in response to a user's quick-capture command. This function enables users to capture fleeting moments.
[0004] Typically, to use the quick capture function, users need to unlock their electronic device and launch the camera app. They then use the quick capture mode within the camera app to take a picture. This requires a series of pre-operations before the quick capture function can be used, resulting in a delay compared to the desired instant capture and a poor user experience. Summary of the Invention
[0005] This application provides a shooting method and device that eliminates the need for users to perform a series of pre-operations; users can quickly take pictures simply by assuming a shooting position.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] In a first aspect, this application provides a shooting method applied to an electronic device, the method comprising:
[0008] Under preset conditions, the electronic device captures images using its camera. These preset conditions include: the electronic device is in a target posture; the facial expression of the subject in the captured preview image is a predetermined expression; and the target posture is either landscape or portrait mode. The electronic device can also save the captured images.
[0009] Thus, in this embodiment, the electronic device acquires and saves the captured image when preset conditions are met. No pre-processing is required from the user; simply assuming a shooting position (e.g., posing the electronic device in the desired posture and displaying a shooting expression) is sufficient for rapid image capture. This allows users to quickly capture desired moments, achieving true fast shooting and enhancing the user experience.
[0010] In one feasible approach, the preset conditions may include one or more of the following: the distance between the electronic device and the subject being photographed is within a preset distance range; the electronic device is held in a predetermined way; wherein the predetermined way of holding the electronic device is different depending on the target posture; and the degree of positional change of the electronic device is greater than a threshold.
[0011] Thus, this application embodiment can further determine the specific scenario when a user uses the mobile phone by considering multiple dimensions such as the device's movement state, the way the device is held, and the distance between the device and the user. This facilitates the subsequent precise triggering of the quick-capture process and improves the accuracy of the shooting.
[0012] In one feasible approach, the method further includes: determining whether the electronic device is in the target pose;
[0013] When the electronic device is in the target posture, a preview image is captured using a camera;
[0014] If the subject's expression in the preview image matches the predetermined expression, the preset conditions are confirmed to be met.
[0015] Thus, this embodiment of the application can directly trigger the quick-capture process after determining that the electronic device meets the above-mentioned preset conditions. Normal shooting is performed after confirming that the electronic device is in shooting mode and the user has made the facial expression required for shooting, thereby improving the success rate of shooting, greatly reducing the rate of unusable shots, and saving time and storage space.
[0016] In one feasible approach, the camera is configured in a low-power mode.
[0017] Thus, this embodiment of the application can acquire preview images in a low-power manner, facilitating subsequent expression detection on the preview images. While ensuring the detection quality of expressions, acquiring preview images in a low-power manner reduces the power consumption of the camera and related hardware. Simultaneously, the low-power hardware design and optimized software algorithms enable faster operating system response, improving the efficiency of the entire shooting process.
[0018] In one possible approach, the captured image is a frame from a preview stream captured by a camera in low-power mode.
[0019] In this way, the phone doesn't need to launch the target application (such as the camera app) to obtain the preview image, reducing the overall operating burden on the device and making the operating system more responsive. This can reduce the overall power consumption of the device and extend battery life.
[0020] In one feasible approach, the method also includes:
[0021] Electronic devices extract facial features of the subject from the preview image.
[0022] The electronic device matches the facial features of the photographed subject with a predetermined facial expression to obtain the facial expression similarity.
[0023] When the facial expression similarity is greater than or equal to the facial expression threshold, the electronic device determines that the facial expression of the subject in the preview image is the predetermined facial expression.
[0024] Thus, in this embodiment, the electronic device can utilize the facial features of the subject in the preview image to determine that the subject's expression in the preview image is a predetermined expression. This facilitates directly triggering the quick-capture process later. After confirming that the phone is in shooting mode and the user has made the desired facial expression, normal shooting is performed, improving the success rate of capturing images.
[0025] In one possible implementation, the electronic device further includes an expression detection service, and the method further includes: acquiring a preview image captured by a camera via the expression detection service; and identifying the expression of the subject in the preview image as a predetermined expression via the expression detection service.
[0026] Thus, this embodiment of the application can directly call the system-level expression detection service to acquire the current preview image through the camera and perform detection. The system service has high stability and reliability, ensuring the continuity and stability of preview image processing. Simultaneously, the system service can uniformly manage and schedule the camera, rationally allocating system resources and avoiding resource conflicts when multiple applications access the camera simultaneously. Furthermore, the system service can directly call relevant interfaces to quickly acquire the preview image captured by the camera, such as by calling the preview stream corresponding to the camera. This reduces the time delay experienced from acquiring the preview image to actually starting to save it.
[0027] In one possible implementation, the method further includes: the electronic device acquiring first data, the first data being used to characterize the attitude of the electronic device. The electronic device can also determine, based on the first data, that it is in a target attitude.
[0028] In one possible implementation, the method further includes: the electronic device can acquire second data, which characterizes the distance between the electronic device and the subject being photographed.
[0029] The electronic device can also determine that the distance between the electronic device and the subject being photographed is within a preset distance range when the second data is greater than or equal to the distance threshold.
[0030] Therefore, in addition to processing the first data, this embodiment of the application also utilizes sensor data collected by other different sensors for fusion and determination, further determining whether the phone is currently in a shooting scene based on both the phone's posture and position. This facilitates the determination of whether to subsequently acquire a preview image. It avoids the possibility of false captures caused by triggering subsequent preview image acquisition and quick-capture processes based on sensing information from a single sensor.
[0031] In one possible implementation, the method further includes: the electronic device acquiring touch data. Based on the touch data, the electronic device can determine a predetermined holding method. The predetermined holding method includes a single-handed holding method corresponding to the user holding the electronic device in portrait mode or a single-handed holding method corresponding to the user holding the electronic device in landscape mode.
[0032] Thus, this embodiment of the application can determine the holding method through touch data, more comprehensively and accurately determining whether the phone is currently in a shooting scene. A more precise determination of whether the phone is in a shooting scene triggers a quick-capture process, preventing accidental captures. This, in turn, improves the user experience.
[0033] In one feasible approach, during the process of determining the holding method of the electronic device as a predetermined holding method based on touch data, the confidence level of the predetermined holding method is determined based on the touch data. The confidence level of the predetermined holding method indicates the degree of confidence that the corresponding holding method of the electronic device is the predetermined holding method. When the confidence level of the predetermined holding method is greater than or equal to a target preset threshold, the holding method of the electronic device is determined to be the predetermined holding method.
[0034] In one possible implementation, the method further includes: the electronic device acquiring third data, the third data being used to characterize the movement of the electronic device, and the electronic device also determining, based on the third data, that the degree of position change of the electronic device is greater than a threshold.
[0035] Therefore, in addition to processing the first and second data, this embodiment of the application also utilizes sensor data collected by other different sensors for fusion and determination. By adding the dimension of device position change, it further determines whether the mobile phone is currently in a shooting scene. This facilitates the determination of whether to subsequently acquire a preview image. It avoids the occurrence of false shots caused by triggering subsequent preview image acquisition and quick capture processes based on sensing information from a single sensor.
[0036] In one feasible approach, before identifying the expression of the subject in the preview image as a predetermined expression, the method further includes: performing face detection on the preview image to obtain a face detection result, the face detection result including the completeness of the face corresponding to the subject in the preview image.
[0037] Secondly, embodiments of this application provide an electronic device, which includes a camera, a memory, one or more processors, and a display screen; the camera, display screen, and memory are coupled to the processor; wherein the camera is used to capture images, the display screen is used to display the captured images, and the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the shooting method as described in the first aspect above.
[0038] Thirdly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the shooting method described in the first aspect above.
[0039] Fourthly, embodiments of this application provide a computer program product, which includes a computer-readable storage medium storing a computer program, such that when at least one processor executes the computer program, the at least one processor performs the shooting method as described in the first aspect above. Attached Figure Description
[0040] Figure 1 is a structural schematic diagram of a mobile phone provided in an embodiment of this application;
[0041] Figure 2 is a software structure block diagram of a mobile phone provided in an embodiment of this application;
[0042] Figure 3 is a flowchart illustrating a shooting method provided in an embodiment of this application;
[0043] Figure 4 is a schematic diagram of the interface of a mode control corresponding to a target mode provided in an embodiment of this application;
[0044] Figure 5 is a schematic diagram of a process for obtaining a preview image according to an embodiment of this application;
[0045] Figure 6 is a schematic diagram of a process for acquiring and capturing images according to an embodiment of this application;
[0046] Figure 7 is a schematic diagram of an interface for displaying captured images provided in an embodiment of this application;
[0047] Figure 8 is a schematic diagram of a predetermined holding method provided in an embodiment of this application;
[0048] Figure 9 is a schematic diagram of the hardware structure of a mobile phone provided in an embodiment of this application. Detailed Implementation
[0049] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish the same or similar items with basically the same function and effect.
[0050] Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or order of execution, and that "first," "second," etc., are not necessarily different. Furthermore, in some embodiments of this application, words such as "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.
[0051] Furthermore, the device architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of device architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0052] With the widespread use of electronic devices, their functions have become increasingly powerful, especially smartphones. Many smartphones now come equipped with cameras, and the corresponding camera performance is constantly improving. For example, electronic devices can offer quick-capture functions, allowing users to capture fleeting moments.
[0053] Typically, to use the quick capture function, users need to unlock their electronic device and launch the camera app. They then use the quick capture mode within the camera app to take a picture. This requires a series of pre-operations before the quick capture function can be used, resulting in a delay compared to the desired instant capture and a poor user experience.
[0054] In related technologies, camera applications on electronic devices may include a snapshot mode, allowing users to trigger the shutter button to take a picture. Alternatively, the shutter button can be associated with the volume button on the electronic device, enabling users to trigger the volume button to quickly initiate the shooting process. However, using snapshot mode still requires a series of pre-operations (such as turning on the electronic device and launching the camera application) before the shot can be taken. In the scenario where the shutter button is associated with the volume button, most users are unlikely to immediately think of pressing the volume button, making it difficult to develop a habitual way of using the volume button for shooting. Based on these solutions, users cannot press the shutter button immediately, thus missing the opportunity to capture a snapshot and failing to achieve true quick shooting.
[0055] To address the aforementioned problems, this application provides a shooting method applied to an electronic device. The method includes: capturing an image using the electronic device's camera when the electronic device meets preset conditions. The preset conditions include: the electronic device is in a target posture; the facial expression of the subject in the captured preview image is a predetermined expression; and the target posture is either a landscape or portrait orientation. The electronic device can also save the captured image.
[0056] Thus, in this embodiment, the electronic device acquires and saves the captured image when preset conditions are met. No pre-processing is required from the user; simply assuming a shooting position (e.g., posing the electronic device in the desired posture and displaying a shooting expression) is sufficient for rapid image capture. This allows users to quickly capture desired moments, achieving true fast shooting and enhancing the user experience.
[0057] For example, taking a mobile phone as an electronic device, Figure 1 shows a schematic diagram of the structure of a mobile phone 100.
[0058] Mobile phone 100 may include processor 110, external memory interface 120, internal memory 121, Universal Serial Bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, buttons 190, motor 191, indicator 192, camera 193, display screen 194, and Subscriber Identification Module (SIM) card interface 195, etc.
[0059] The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and an image sensor 180N.
[0060] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0061] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0062] The controller can serve as the central nervous system and command center of the mobile phone 100. Based on the instruction operation code and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0063] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0064] The wireless communication function of mobile phone 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.
[0065] The mobile phone 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0066] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the mobile phone 100 may include one or N displays screens 194, where N is a positive integer greater than 1.
[0067] The mobile phone 100 can achieve shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0068] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0069] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0070] Here, the camera 193 can be located within the mobile phone 100. Alternatively, it can be an external component of the electronic device. In some implementations, the camera can also be located externally to the electronic device and connected via wired or wireless means. For example, the camera can connect to the electronic device via Bluetooth or a mobile hotspot. The electronic device can control the camera by sending or receiving commands.
[0071] The camera module can be located within the camera module 193. Alternatively, it can be placed in other locations within the phone 100. The camera module includes a lens, a focusing motor, a base, a circuit board, and an image sensor.
[0072] The base is fixedly connected to one side of the circuit board. The focusing motor is located on the side of the base away from the circuit board and is fixedly connected to the periphery of the base. The lens is mounted in the center of the focusing motor. The image sensor is fixed to the side of the circuit board facing the lens.
[0073] The lens is used to capture the light signal reflected from the subject. The focusing motor is used to drive the lens to move in a direction parallel to the optical axis. The optical axis refers to the line passing through the center of the lens. In some embodiments, the mobile phone 100 can control the focusing motor to move the lens to the focusing position, thereby completing the focusing process.
[0074] In some embodiments, the focusing motor may be a voice coil motor (VCM), a shape memory alloy (SMA) motor, a piezo motor (PM), or a stepper motor (STM), etc.
[0075] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0076] The mobile phone 100 can achieve audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0077] The image sensor 180N can be used to detect objects within the range captured by the camera, with each photosensitive unit corresponding to a pixel in the image sensor. The image sensor 180N may include a color (red, green, blue, RGB) image sensor, a monochrome image sensor, and an infrared image sensor, etc., but this embodiment does not limit the specific type. The image sensor 180N is used to acquire raw images, which may include RGB images, RYB images, monochrome images, and infrared images, etc.
[0078] In some embodiments, multiple image sensors can be arranged in the same camera. For example, a single-lens dual-sensor camera integrates both an RGB image sensor and a motion sensor within a single camera. Other examples include dual-lens dual-sensor cameras and single-lens triple-sensor cameras, where these sensors are used to image the same subject. When the number of lenses is less than the number of sensors, a beam splitter can be placed between the lenses and sensors to distribute the light entering through one lens across multiple sensors, ensuring that each sensor receives light. Furthermore, the number of processors in these cameras can be one or more. This application does not specifically limit the arrangement or number of components.
[0079] In other embodiments, only one image sensor may be provided in the same camera.
[0080] In some embodiments, the camera 193 can capture still images or moving images. The display screen 194 is used to display the still images or moving images.
[0081] In some examples, camera 193 can acquire a preview image of the subject when the electronic device meets preset conditions. Camera 193 can also acquire a captured image if the preview image contains a predetermined expression, and display screen 194 is used to display the captured image.
[0082] The electronic device provided in this application embodiment can run an operating system (OS). This operating system can be various operating systems used in industry, such as an operating system developed based on OpenHarmony, for example... Or other operating systems, such as The iOS mobile operating system. It can also be any open-source operating system or its derivatives, such as Linux. This includes other embedded operating systems, as well as future new operating systems, such as AI operating systems based on artificial intelligence. An operating system is a set of interconnected system software programs that manage and control the operation of electronic devices, utilize and run hardware and software resources, and provide public services to organize user interactions.
[0083] In electronic devices, the operating system connects to the physical devices at the hardware layer below and provides a runtime environment for application software above.
[0084] An operating system typically includes a kernel layer, a middleware layer, and an application layer. The application layer comprises applications, which can include system applications and third-party applications. The middleware layer includes a suite of software providing various services to application developers, or frameworks providing services such as databases, multimedia, and graphics, or capabilities such as distributed scheduling and system scaling.
[0085] The electronic devices we use in our daily lives come in various types and forms, and are applied in a wide range of scenarios. Therefore, based on the different forms and functions of electronic devices, different application scenarios, and different user needs, the operating systems used in these devices may also differ. The basic functions implemented by the electronic device provided in this application can be achieved through a general-purpose operating system or a dedicated operating system.
[0086] To more clearly illustrate the implementation of the embodiments of this application under a specific operating system, the following is shown. Based on the architecture, those skilled in the art can deduce the implementation of the embodiments of this application under other specific operating systems, such as... Implementation under operating systems, etc.
[0087] Figure 2 is a software structure block diagram of the mobile phone 100 according to an embodiment of this application.
[0088] The software architecture of Mobile Phone 100 can be divided into several layers. In some embodiments, from bottom to top, these layers are: kernel layer, system service layer, framework layer, and application layer. Layers communicate with each other through software interfaces. System functions can be tailored, added, or combined at the subsystem level depending on the deployment scenario of different device forms. Each subsystem can also be tailored, added, or combined at the functional level.
[0089] The kernel layer includes the kernel abstraction layer, the kernel subsystem, and the driver subsystem.
[0090] The system service layer comprises the core capabilities of the system, providing services to applications through the framework layer. This layer includes, but is not limited to, the following subsystems:
[0091] The system's basic capability subsystems provide the foundational capabilities for the operation, scheduling, and migration of distributed applications across multiple devices. These subsystems may include a distributed soft bus, distributed data management, distributed task scheduling, and the Ark multi-language runtime. They also include a multi-modal input subsystem, a graphics subsystem, a security subsystem, and an AI business subsystem.
[0092] The AI business subsystem may include an expression detection service, which is used to receive preview images captured by a camera and to identify whether a predetermined expression exists in the preview image.
[0093] Basic software service subsystems: These provide common and general software services. They may include an event notification subsystem, a telephone service subsystem, a multimedia subsystem, etc.
[0094] Enhanced Software Service Subsystem Set: Provides differentiated, enhanced software services tailored to different devices. This set may include proprietary business subsystems for smart screens, wearables, and IoT, among others.
[0095] Hardware service subsystem set: Provides hardware services. The hardware service subsystem set may include location service subsystem, user IAM (Identity and Access Management) subsystem, wearable proprietary hardware service subsystem, biometric identification subsystem, IoT proprietary hardware service subsystem, etc.
[0096] Distributed task scheduling enables distributed service management (discovery, synchronization, registration, and invocation), supporting remote startup, remote invocation, remote connection, and migration of applications across devices.
[0097] Distributed data management enables data synchronization, data storage, data sharing, and data access across all scenarios and devices.
[0098] The distributed soft bus provides communication-related capabilities for seamless interconnection between multiple devices, including: WLAN service capabilities, Bluetooth service capabilities, soft bus, inter-process communication RPC (Remote Procedure Call), and StarFlash communication capabilities.
[0099] Ark Multilingual Runtime is a unified compilation runtime platform designed to support the joint compilation and execution of multiple programming languages and multiple chip platforms.
[0100] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer includes: the ArkUI framework (which provides a complete infrastructure for UI development of system applications, including UI functions such as components, layouts, animations, and interactive events, as well as a real-time interface preview tool), the user application framework, and the Ability framework (an Ability is a lightweight application; the Ability framework schedules and manages the operation and lifecycle of Abilities). Different devices may have different operating systems, and the APIs they support may also differ.
[0101] Applications can include system apps and extended / third-party apps. System apps can include the desktop, control bar, settings, contacts, input method, gallery, etc., while extended / third-party apps can include social apps, travel apps, etc.
[0102] The following embodiments, with reference to the accompanying drawings, will illustrate the shooting method provided by this application, using a mobile phone having the structure shown in FIG1 as an example. Referring to FIG3, the method may include:
[0103] S300, the phone is in target mode.
[0104] For example, the target mode can be an intelligent capture mode. The target mode can also be an intelligent quick capture mode. This application does not limit the specific implementation of the target mode. The target mode is used to implement the quick capture or capture function in a mobile phone.
[0105] In some embodiments of this application, the electronic device is provided with a target mode, and the switch for the target mode is set in the system application of the mobile phone to support users in quickly capturing wonderful moments.
[0106] Specifically, the phone can display a target mode interface. Users can trigger operations on the target controls corresponding to the target mode in the target mode interface, and the phone can respond to the trigger operation on the target mode interface and activate the target mode.
[0107] In one implementation, the target mode interface may include a target control corresponding to the target mode. Users can use the target control to activate the target mode for subsequent quick shooting of the subject.
[0108] For example, the mobile phone can respond to an activation operation on the target control by activating the target mode corresponding to the target control. Subsequently, after activating the target mode, the mobile phone can trigger the quick capture process by sequentially performing posture detection and expression detection by collecting sensor data and preview images.
[0109] In some examples, the trigger entry point corresponding to the above target mode can be presented as an independent switch, such as through the interface of an application on a mobile phone.
[0110] For example, referring to Figure 4(A), the management interface of the settings application on the mobile phone may include a mode control 401 corresponding to the target mode. The user can operate on the mode control 401, such as by clicking it. In response to the user's click on the mode control 401, the mobile phone triggers entry into the target mode interface. The target mode interface includes a target control, such as the mode control 402 corresponding to Smart Quick Capture. The user can operate on the mode control 402, such as by clicking it. In response to the user's activation operation on the mode control 402, the mobile phone activates the target mode.
[0111] Of course, when the target mode is enabled, the user can interact with the mode control 402 again, such as by clicking it. The phone responds to the user's subsequent interaction with the mode control 402 and disables the target mode. It should be noted that the above examples only use the settings application and target mode interface as examples; this embodiment does not limit the specific location of the target control.
[0112] In another possible implementation, the trigger entry corresponding to the target mode can also be presented in the form of an integrated switch. For example, the target mode can be integrated with the mode switch corresponding to at least one of the modes: privacy mode, power saving mode, and do not disturb mode.
[0113] In other words, this application can reuse existing mode switches to control the target mode. When the user turns on an existing mode switch, the target mode is also turned on accordingly.
[0114] In other examples, the trigger entry point corresponding to this target mode can also be presented in the form of a physical button.
[0115] For example, referring to Figure 4(B), a physical button 403 corresponding to the target mode can be provided on the casing of the mobile phone. The user can operate the physical button 403, such as by pressing it. The mobile phone responds to the user's pressing operation on the physical button 403 and activates the target mode. Of course, when the target mode is activated, the user can operate the physical button 403 again, such as by pressing it, and the mobile phone responds to the user's pressing operation on the physical button 403 and deactivates the target mode. Of course, the user can also perform different operations on the physical button 403 to activate and deactivate the target mode.
[0116] It should be noted that the embodiments of this application do not specifically limit the implementation method of the physical buttons corresponding to the target mode.
[0117] In some embodiments of this application, after a user triggers the target control corresponding to the target mode, the mobile phone can respond to the user's activation of the target mode and activate it. This allows the mobile phone to quickly trigger shooting and obtain a captured image after activating the target mode. Consequently, users can capture desired moments more quickly, achieving true fast shooting and enhancing the user experience.
[0118] Thus, in this embodiment, the target mode switch can be set in the phone's system applications (such as the settings app). Users can quickly launch the quick-capture function by setting a shortcut at the system level, simplifying the user's operation steps. Simultaneously, setting a switch at the system level allows for on-demand loading of camera resources, enabling quick access when needed for shooting and timely release of resources after shooting, preventing the camera application from consuming excessive memory while running in the background for extended periods. This improves system performance.
[0119] In some embodiments of this application, the mobile phone can execute S300 before S301. Furthermore, after the user activates the target mode, the mobile phone will execute the subsequent quick-capture process as long as the target mode is not deactivated. That is, the mobile phone does not need to execute S300 before each execution of S301. Of course, the target mode can also be enabled by default, requiring no manual activation by the user; in other words, the mobile phone does not need to execute S300.
[0120] S301. Under preset conditions, the mobile phone uses its camera to capture images.
[0121] In real life, when users want to take a photo, they generally position their phones in a corresponding shooting posture and their faces display a corresponding expression, which is called a "shooting expression." In some examples, users can hold their phones horizontally or vertically while displaying a selfie expression. For instance, a user can hold their phone horizontally and use the front or rear camera to take a picture of themselves with a shooting expression. As another example, a user can hold their phone vertically and use the front or rear camera to take a picture of other users, and those users' faces will display a shooting expression.
[0122] The preset conditions include: the mobile phone is in the target posture, the expression of the subject in the captured preview image is the predetermined expression; the target posture is either landscape or portrait.
[0123] Therefore, in this embodiment of the application, when the mobile phone is in target mode, it is determined whether the mobile phone is in the target posture. When the mobile phone is in the target posture, a preview image is captured using the camera.
[0124] Specifically, the phone can perform posture detection based on sensor data to determine whether it is in the target posture. After determining that the phone's current posture meets the shooting requirements, it captures a preview image for expression detection. Then, if the aforementioned preset conditions are met, the phone uses its camera to capture and shoot the image.
[0125] For example, referring to Figure 5, the mobile phone can acquire first data, which characterizes the phone's posture. The phone can also determine whether it is in a target posture based on the first data. When the phone determines it is in the target posture based on the first data, it acquires a preview image. Otherwise, the process ends.
[0126] The first data includes sensor data collected by the first sensor.
[0127] The first sensor can be at least one of a gyroscope, accelerometer, orientation sensor, gravity sensor, or magnetic sensor. An accelerometer determines the phone's attitude by measuring the force applied along a specific axis. This force is represented by its direction (x-axis, y-axis, z-axis) and the magnitude of acceleration in that direction. A gyroscope measures the angle between the gyroscope's vertical axis and the phone within a three-dimensional coordinate system (including x-axis, y-axis, z-axis), calculates the angular velocity, and uses the angle and angular velocity to determine the phone's attitude in three-dimensional space. A magnetic sensor measures the angles between the phone and the four cardinal directions (north, south, east, west), and uses these angles to determine the phone's attitude. Orientation and gravity sensors detect the device's orientation, determining the phone's attitude by detecting changes in physical quantities such as magnetic fields, gravity, and light.
[0128] It should be noted that the first sensor provided in this application embodiment can be a single sensor or a type of sensor. The first sensing data can be collected by one sensor or by multiple sensors. For example, the first data can be collected by a gravity sensor, the first data can be collected by a gyroscope sensor, the first data can also be collected by a gyroscope sensor and a magnetic sensor, or the first data can be collected by a gyroscope sensor and an accelerometer sensor. This application embodiment does not specifically limit the type and number of the first sensor or the implementation method of the first data.
[0129] In some examples, the first set of data includes both acceleration and angular velocity data. Acceleration data can be from at least one of the three axes (x, y, z), or it can be acceleration data obtained by merging the acceleration data from the x, y, and z axes. Angular velocity data can be yaw, pitch, and roll angular velocities. The phone can use a complementary filtering algorithm to fuse the acceleration and angular velocity data to accurately determine whether the phone is in landscape or portrait orientation.
[0130] It should be noted that mobile phones can also use other algorithms to fuse acceleration and angular velocity data, and this application does not limit this.
[0131] In another possible implementation, the first data includes acceleration data collected by an accelerometer, magnetometer data collected by a magnetometer, and angular velocity data collected by a gyroscope. The phone can then fuse the magnetometer data with the acceleration and angular velocity data. For example, algorithms such as Kalman filtering can be used to comprehensively process the data from the three sensors to obtain a more accurate picture of the phone's attitude in space.
[0132] The acquisition of the first data described above is a periodically executed step. Typically, the first sensor can periodically collect sensor data. For example, the first sensor can collect the first data periodically at preset time intervals (preset time), such as 10ms, 100ms, 1s, 2s, 10s, 20s, 50s, 60s, 100s, etc. This allows for the continuous acquisition of updated first data, enabling the continuous detection of changes in the phone's posture through changes in the data within the first data.
[0133] Thus, the first data in this embodiment can be collected by the first sensor, which detects the phone's posture. This allows the phone to determine whether its current posture matches the shooting posture based on the first data. Consequently, after making an initial decision about the phone's posture, a fusion decision can be made with the user's facial expression. This avoids the problem of accidental snapshots caused by triggering a quick capture based solely on the phone being in a target posture.
[0134] In some embodiments of this application, the mobile phone obtains a preview image when it determines whether the phone is in landscape or portrait orientation based on the first data.
[0135] The phone includes a camera. The phone uses the camera to capture preview images. During the process of capturing preview images, the camera is configured in a low-power mode.
[0136] For example, a mobile phone can use the camera's low-power mode (such as front-facing swing capability) to acquire a preview image. The image quality of the preview image is lower than that of the image captured by the camera in normal shooting mode. Front-facing swing capability refers to the characteristic of the camera to minimize energy consumption while ensuring a certain image quality during the acquisition of preview images, through low-power hardware design and optimized software algorithms.
[0137] For example, a mobile phone can open its camera and use the configured image sensor parameters to capture and preview images. Developers can pre-configure the image sensor parameters corresponding to the target acquisition capability, such as configuring the camera after it leaves the factory. Image sensor parameters may include lower resolution, smaller pixel bit depth, and lower frame rate, etc. This application does not limit the specific values of the configured image sensor parameters.
[0138] Thus, this embodiment of the application can acquire preview images in a low-power manner, facilitating subsequent expression detection on the preview images. While ensuring the detection quality of expressions, acquiring preview images in a low-power manner reduces the power consumption of the camera and related hardware. Simultaneously, the low-power hardware design and optimized software algorithms enable faster operating system response, improving the efficiency of the entire shooting process.
[0139] In some examples, the preview image is a frame from a preview stream captured by a camera in low-power mode.
[0140] In other words, the preview image is a single frame from the preview stream captured by the camera, rather than an image captured by the target application. The target application can be an application with shooting capabilities, such as a camera app, image editing app, etc. This application does not limit the type of target application.
[0141] In this way, the phone doesn't need to launch the target application (such as the camera app) to obtain the preview image, reducing the overall operating burden on the device and making the operating system more responsive. This can reduce the overall power consumption of the device and extend battery life.
[0142] In some embodiments of this application, if the mobile phone is in the target posture and the expression of the subject in the preview image is a predetermined expression, then a preset condition is determined to be met. Thus, the mobile phone can capture an image when it recognizes the expression of the subject in the preview image as the predetermined expression.
[0143] The predetermined expressions include the facial expressions displayed by the subject during the shooting process. It is understood that predetermined expressions are pre-configured facial expressions. Predetermined expressions may include at least one of the following: happy expressions (such as a smile), sad expressions, angry expressions, surprised expressions, fearful expressions, and funny expressions. Of course, in addition to the above expressions, the subject may also display some complex expressions or subtle changes in expression, such as a surprised expression (a combination of surprise and happiness), a wry smile (a combination of sadness and helplessness), or a contemptuous expression (an expression of disdain and arrogance). This application does not limit the specific implementation of the predetermined expressions, or the implementation of determining the expression of the subject in the preview image as a predetermined expression.
[0144] In some embodiments of this application, after acquiring the preview image, the mobile phone can perform facial expression detection on the preview image.
[0145] Specifically, the phone also includes an expression detection service. This service can identify whether a predetermined expression exists in a preview image. The expression detection service is deployed in the framework layer of the operating system. The phone acquires the preview image captured by the camera through the expression detection service and identifies the expression of the subject in the preview image as the predetermined expression.
[0146] For example, referring to Figure 6, the expression detection service can acquire a preview image captured by a camera and determine whether a predetermined expression exists in the preview image. After determining that the predetermined expression exists in the preview image, the expression detection service can capture a photographic image using the camera. Otherwise, the process ends.
[0147] Therefore, in this embodiment, there is no need to launch the camera application; the system-level expression detection service can be directly invoked to acquire the current preview image through the camera and perform detection. The system service has high stability and reliability, ensuring the continuity and stability of preview image processing. Simultaneously, the system service can uniformly manage and schedule the camera, rationally allocating system resources and avoiding resource conflicts when multiple applications access the camera simultaneously. Furthermore, the system service can directly call relevant interfaces to quickly acquire the preview image captured by the camera, such as by calling the preview stream corresponding to the camera. This reduces the time delay experienced from acquiring the preview image to actually starting to save it.
[0148] In some embodiments of this application, during the process of recognizing the expression of the subject in the preview image as a predetermined expression, the mobile phone can determine whether the predetermined expression exists in the preview image based on the expression characteristics of the subject in the preview image.
[0149] Specifically, the expression detection service can extract the facial expression features of the subject in the preview image. The service can then match these features with a predetermined expression to obtain an expression similarity score. If the similarity score is greater than or equal to an expression threshold, the expression of the subject in the preview image is determined to be the predetermined expression.
[0150] In one implementation, the expression detection service uses a deep learning model to identify whether a predetermined expression exists in a preview image. The deep learning model can include a convolutional neural network (CNN) model. This deep learning model can also be an artificial intelligence model. The expression detection service can input the preview image into the trained CNN model, and the CNN model outputs a recognition result, which characterizes whether a predetermined expression exists in the preview image.
[0151] It should be noted that convolutional neural network models can be trained using a large amount of image data with predetermined facial expressions. The model parameters can be adjusted to enable accurate recognition of these expressions. The image data can encompass images of different ages, genders, ethnicities, and under various lighting conditions and shooting angles to ensure data diversity and representativeness. Furthermore, the expressions in these different images correspond to expression categories, such as happiness, sadness, anger, surprise, fear, and disgust.
[0152] During operation, a convolutional neural network model matches the facial features of a photographed subject with a predetermined expression to obtain an expression similarity score. When the expression similarity score is greater than or equal to an expression threshold, the expression of the photographed subject in the preview image is determined to be the predetermined expression. This expression similarity score can also be understood as the probability value of the expression in the preview image corresponding to the expression category. For example, if the expression in the preview image is a smiling expression, and the probability value of the smiling expression corresponding to the expression category (0.9) is greater than the expression threshold (0.75), then it is determined that a smiling expression exists in the preview image.
[0153] In some embodiments of this application, the mobile phone captures images when the aforementioned preset conditions are met. That is, when the mobile phone is in the target posture and the expression of the subject in the captured preview image is a predetermined expression, the quick capture process can be directly triggered.
[0154] Specifically, when the expression detection service identifies the expression of the subject in the preview image as a predetermined expression, it can send a shooting command to the camera. This shooting command instructs the camera to perform a shooting operation. In response to the shooting command, the camera performs the shooting operation and obtains the captured image.
[0155] Thus, this embodiment of the application can directly trigger the quick-capture process after the mobile phone meets the above-mentioned preset conditions. Normal shooting is performed after confirming that the mobile phone is in shooting mode and the user has made the facial expression required for shooting, thereby improving the success rate of shooting, greatly reducing the rate of unusable photos, and saving time and storage space.
[0156] In some embodiments of this application, before recognizing the expression of the subject in the preview image as a predetermined expression, the mobile phone can also perform face detection on the preview image to obtain a face detection result. The face detection result includes the completeness of the face corresponding to the subject in the preview image. In this way, the mobile phone can recognize whether the expression of the subject in the preview image is a predetermined expression after determining that the completeness of the face corresponding to the subject is greater than a preset threshold.
[0157] Thus, this embodiment of the application can identify whether a predetermined expression exists in the preview image when the completeness of the face corresponding to the subject being photographed is relatively high. On the one hand, a complete face provides sufficient facial feature information, making the recognition of the predetermined expression more accurate. On the other hand, it can more accurately determine whether the current user is in a shooting state, avoiding the problem of accidental shooting when the current user does not intend to shoot. This improves the reliability of the shooting process.
[0158] S302, Save captured images on your phone.
[0159] S303: The phone displays the captured image on its screen.
[0160] In some embodiments of this application, after the camera captures an image, it can save the image and transmit it to the display screen, activate the display screen, and display the captured image on the display screen.
[0161] For example, referring to Figure 7(A), the user can position the phone in a corresponding shooting state with facial expression control. There is no need to call the camera app to capture the image. Referring to Figure 7(B), after capturing the image, the phone can wake up the display screen to show the captured image. For example, the phone can use a photo album app to display the captured image. If the display screen can show the photo album app interface, the photo album app interface can include the captured image.
[0162] Thus, in this embodiment, users can quickly view the captured image after positioning their phone in the appropriate shooting state and making a selfie expression. The electronic device can take pictures without launching a target application (such as a camera app), quickly initiating the shooting function and rapidly recording the moment. Users can view the captured image on the display screen without performing a series of cumbersome operations. This enhances the user's shooting flexibility and user experience, achieving true fast shooting.
[0163] In some embodiments of this application, the preset conditions also include one or more of the following conditions:
[0164] The distance between the mobile phone and the subject is within a preset distance range, and the way the mobile phone is held is a predetermined holding method; wherein, the predetermined holding method is different when the mobile phone is in different target postures, and the degree of change of the mobile phone position is greater than a threshold.
[0165] Thus, this application embodiment can further determine the specific scenario when a user uses the mobile phone by considering multiple dimensions such as the device's movement state, the way the device is held, and the distance between the device and the user. This facilitates the subsequent precise triggering of the quick-capture process and improves the accuracy of the shooting.
[0166] In some scenarios where users want to take a photo, they need to place the phone at a certain distance from their face to capture a complete image. Therefore, optionally, the phone can determine its current distance from the user before acquiring the preview image. Only after determining that a certain distance exists between the phone and the user should the preview image be acquired.
[0167] In one feasible approach, the mobile phone can determine the distance between itself and the user before determining whether its posture is a target posture based on the first data. Once the mobile phone determines that the distance between itself and the subject is within a preset distance range, it performs the steps of acquiring the first data and determining whether the phone is in the target posture based on the first data.
[0168] Specifically, the mobile phone can acquire second data, which represents the distance between the mobile phone and the subject being photographed. If the second data is greater than or equal to a distance threshold, the mobile phone can determine that the distance between the mobile phone and the subject is within a preset distance range.
[0169] In another possible approach, the phone can determine whether its posture is the target posture based on the first data. For example, after determining that the phone's posture is the target posture, the distance between the phone and the user can be determined. If a certain distance is determined between the phone and the user, the step of acquiring a preview image can be performed.
[0170] For example, the second data may be data collected by a second sensor.
[0171] The second sensor may include an image sensor and / or a proximity sensor. The proximity sensor may include ultrasonic, laser, and infrared proximity sensors, etc. The image sensor and proximity sensor may be always-on or not always-on. If the second sensor is not always-on, the phone can trigger the activation of the second sensor and acquire the corresponding second data after determining the phone is in the target posture based on the first data. Alternatively, the phone can first activate the second sensor, acquire the corresponding second data, and then acquire the first data.
[0172] In some examples, the second sensor may include an image sensor. The image sensor may include a time-of-flight (TOF) sensor and an infrared image sensor, etc. The second data may include distance data and / or distance change data acquired by the image sensor; the distance data characterizes the distance between the phone and the user, and the distance change data characterizes how the distance between the phone and the user changes during movement.
[0173] For example, the distance threshold can be a single value or a range of values. When the second data includes distance data, the distance threshold can be a single value, such as 30 centimeters. When the second data includes distance variation data, the distance threshold can be a range of values, such as greater than or equal to 10 centimeters and less than or equal to 45 centimeters. This application does not specifically limit the distance threshold in its embodiments.
[0174] Taking a Time-of-Flight (TOF) sensor as an example, the second data includes distance data. A mobile phone can be equipped with a TOF sensor, which can be positioned close to the camera. The TOF sensor acquires a corresponding depth image. Then, the phone identifies the face region in the depth image and obtains the average depth value of the face region by mapping it to the corresponding area in the depth image. This average depth value of the face region is then used as the distance data.
[0175] Similarly, taking a TOF sensor, where the second data includes distance change data, as an example, the TOF sensor in a mobile phone can acquire images at different depths at two different times, and determine the first average depth value and the second average depth value of the face region respectively. Then, the difference between the first average depth value and the second average depth value is determined as the distance change data.
[0176] Subsequently, the phone determines the distance between itself and the user based on distance data and / or distance change data, compared to a distance threshold. Specifically, if the second data includes distance data, the phone can compare the distance data with the corresponding distance threshold to determine the distance between itself and the user. If the second data includes distance change data, the phone can compare the distance change data with the corresponding distance threshold to determine the distance between itself and the user. If the second data includes both distance data and distance change data, the phone can compare both with their respective distance thresholds to determine the distance between itself and the user.
[0177] As can be seen, in addition to processing the first data, this embodiment of the application also utilizes sensor data collected by other different sensors for fusion and determination, further determining whether the phone is currently in a shooting scene based on both the phone's posture and position. This facilitates the determination of whether to subsequently acquire a preview image. It avoids the possibility of false captures caused by triggering subsequent preview image acquisition and quick-capture processes based on sensing information from a single sensor.
[0178] In some embodiments of this application, before placing the phone at a certain distance from the face, the user will pick up the stationary phone and bring it in front of their face. The phone's movement typically changes from stationary to moving, and then back to stationary. Thus, the phone can determine whether a certain distance exists between the phone and the user after confirming the change in movement from stationary to moving and then back to stationary.
[0179] Specifically, the mobile phone can acquire third-party data, which characterizes the phone's movement. The phone can also determine the degree of positional change based on this third-party data. When the phone determines that the degree of positional change exceeds a threshold based on the third-party data, it proceeds to determine whether a certain distance exists between the phone and the user.
[0180] The third data can be acquired by a third sensor. The third sensor may include a gyroscope sensor.
[0181] Specifically, the mobile phone can acquire acceleration data at the first and second moments from the third data set. Based on these acceleration data, the phone can calculate the positional feature value corresponding to the acceleration change. If the positional feature value is greater than a threshold, the phone's position change is determined to be greater than the threshold, thus indicating a significant positional change.
[0182] The first moment is the time when the phone's motion state changes from a stationary state to a moving state, and the second moment is the time when the phone's motion state changes from a moving state to a stationary state. This application embodiment does not specifically limit the first and second moments. The process of the phone being picked up may include multiple first and second moments.
[0183] In the process of determining whether the position feature value is greater than the threshold, the mobile phone can analyze the movement angle of the phone in the three axes from the acceleration change data. If the movement angle in all three axes is greater than the threshold, then the position feature value is determined to be greater than the threshold.
[0184] For example, the mobile phone monitors and acquires acceleration data collected by the gyroscope sensor at the first and second moments in real time. This acceleration data includes acceleration data along three axes (X, Y, and Z). Since the acceleration data represents acceleration data collected along these three axes, a change in the acceleration data along one of these axes indicates that the phone has moved in that direction. Therefore, the movement of the phone can be determined by analyzing the changes in acceleration data along the three axes.
[0185] The embodiments of this application do not impose numerical limitations on the threshold. Those skilled in the art can determine the value of the threshold according to actual needs, such as 0.1, 0.2, 0.3, etc.
[0186] Meanwhile, since the acceleration data is collected in real time by the gyroscope sensor, multiple sets of acceleration data at different times can be acquired to determine whether the phone has moved, thus obtaining a more accurate judgment. For example, when acquiring multiple sets of acceleration data at different times, each set of acceleration data can be compared with a threshold following the above process. If the positional characteristic value of each set of acceleration data is greater than the threshold, it is determined that the degree of change in the phone's position is greater than the threshold.
[0187] Thus, this embodiment of the application can determine the specific scenario of a user using a mobile phone by fusing the first, second, and third data, considering three dimensions: device posture, device movement, and the distance between the device and the user. This facilitates subsequent expression detection and triggering of the quick-capture process based on the preview image, avoiding the occurrence of false captures due to determining the triggering process based on the sensing information of a single sensor.
[0188] In some embodiments of this application, in addition to determining the specific scenario of a user using a mobile phone based on the three dimensions of device posture, device movement, and distance between the device and the user, the dimension of how the user holds the phone can also be determined. Typically, in a selfie scenario, a user will use a predetermined holding method to hold the phone. The predetermined holding method differs depending on the target posture of the phone.
[0189] Referring to Figure 8(A), when the phone is in portrait mode, the predetermined holding method can be a single-handed holding method corresponding to the user maintaining the phone in portrait mode. Referring to Figure 8(B), when the phone is in landscape mode, the predetermined holding method can be a single-handed holding method corresponding to the user maintaining the phone in landscape mode. It should be noted that the predetermined holding method may also include a two-handed holding method; however, this application embodiment does not specifically limit the predetermined holding method.
[0190] Specifically, the phone can acquire touch data and, based on that data, determine the phone's grip as a predetermined method. By adding this process of confirming the current grip as the predetermined method, the phone can further determine whether it is currently in a shooting scenario.
[0191] The predetermined holding methods include a first holding method corresponding to the portrait mode or a second holding method corresponding to the landscape mode.
[0192] In some examples, when the phone is in portrait orientation, touch data is acquired, and the confidence level of a first holding method is determined based on the touch data. The confidence level of the first holding method indicates how credible it is that the phone's corresponding holding method is the first holding method. When the confidence level of the first holding method is greater than or equal to a first preset threshold, the phone's corresponding holding method is determined to be the first holding method.
[0193] Alternatively, when the phone is in landscape mode, touch data is acquired, and the confidence level of the second holding method is determined based on the touch data. The confidence level of the second holding method indicates how credible it is that the phone's corresponding holding method is the second holding method. When the confidence level of the second holding method is greater than or equal to a second preset threshold, the phone's corresponding holding method is determined to be the second holding method.
[0194] The aforementioned touch data may include screen capacitance data within a first preset time period before the phone is determined to be in the target posture, or screen capacitance data within a second preset time period after the phone is determined to be in the target posture. The first preset time period may be the same as or different from the second preset time period. Of course, the phone may also collect and cache the corresponding touch data in real time. This application embodiment does not limit the specific values corresponding to the first and second preset time periods.
[0195] For example, a mobile phone can collect touch data for fixed periods of time to determine the confidence level corresponding to different holding methods. Understandably, a higher confidence level indicates a higher degree of confidence in the user's use of that holding method, while a lower confidence level indicates a lower degree of confidence in the user's use of that holding method.
[0196] Specifically, the mobile phone can input screen capacitance data into the holding model to obtain the confidence level of a first holding method or a second holding method. This embodiment of the application also trains the holding model. During training, a large amount of screen capacitance data can be used as training samples, and two types of holding features are set, such as holding features for the first holding method and holding features for the second holding method. This enables the holding model to learn the ability to recognize the confidence levels of the first and second holding methods.
[0197] In some embodiments, the confidence level described above can be calculated using Convolutional Neural Networks (CNN), Region Proposal Networks (RPN), Regions with CNN features (RCNN), Faster RCNN, MobileV2, and residual networks, or a combination of multiple networks. This application does not limit the specific form of the holding method model.
[0198] Thus, this embodiment of the application can determine the holding method through touch data, more comprehensively and accurately determining whether the phone is currently in a shooting scene. A more precise determination of whether the phone is in a shooting scene triggers a quick-capture process, preventing accidental captures. This, in turn, improves the user experience.
[0199] It should be noted that the embodiments of this application do not limit the execution timing of the processes of determining that the distance between the mobile phone and the subject is within a preset distance range, that the mobile phone is held in a predetermined way, and that the degree of change in the position of the mobile phone is greater than a threshold.
[0200] In some solutions, multiple embodiments of this application can be combined, and the combined solution can be implemented. Optionally, some operations in the process of each method embodiment may be combined, and / or the order of some operations may be changed. Furthermore, the execution order between the steps of each process is merely exemplary and does not constitute a limitation on the execution order between steps; other execution orders are also possible. It is not intended to indicate that the execution order is the only possible order in which these operations can be performed.
[0201] Those skilled in the art will conceive of various ways to reorder the operations described in the embodiments of this application. Furthermore, it should be noted that process details involved in one embodiment of this application are similarly applicable to other embodiments, or different embodiments can be combined.
[0202] Furthermore, some steps in the method embodiments can be equivalently replaced with other possible steps. Alternatively, some steps in the method embodiments may be optional and can be deleted in certain use cases. Or, other possible steps may be added to the method embodiments.
[0203] Furthermore, the various method embodiments can be implemented individually or in combination.
[0204] This application also provides an electronic device, such as the mobile phone described above, as shown in FIG9. The mobile phone may include one or more processors 910, memory 920 and communication interface 930.
[0205] The memory 920, communication interface 930, and processor 910 are coupled together. For example, the memory 920, communication interface 930, and processor 910 can be coupled together via bus 940.
[0206] The communication interface 930 is used for data transmission with other devices. The memory 920 stores computer program code. The computer program code includes computer instructions, which, when executed by the processor 910, cause the electronic device to perform the relevant method steps described in this embodiment.
[0207] Processor 910 may be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. The processor may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0208] Bus 940 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The aforementioned bus 940 can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in Figure 9, but this does not indicate that there is only one bus or one type of bus.
[0209] This application also provides an electronic device, which includes a memory and one or more processors. The memory is coupled to the processors. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device performs the relevant method steps described in the above method embodiments.
[0210] This application also provides an imaging device, which includes a memory and one or more processors. The memory is coupled to the processors. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the imaging device performs the relevant method steps described in the above method embodiments.
[0211] This application also provides a computer-readable storage medium storing computer program code. When the processor executes the computer program code, the electronic device executes the relevant method steps in the above method embodiments.
[0212] This application also provides a computer program product containing instructions that, when executed on a computer or processor, cause the computer or processor to perform the relevant method steps as described in the above method embodiments.
[0213] This application also provides a chip system, including: a processor coupled to a memory, the memory being used to store programs or instructions, and when the program or instructions are executed by the processor, the chip system enables the methods in any of the above method embodiments.
[0214] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.
[0215] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application embodiment does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application embodiment does not specifically limit the type of memory or the arrangement of the memory and processor.
[0216] For example, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0217] The electronic devices, computer storage media, or computer program products provided in this application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0218] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0219] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0220] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units, located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0221] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0222] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the contributing parts, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0223] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A shooting method, characterized in that, Applied to electronic devices, the method includes: Under preset conditions, the camera of the electronic device is used to capture images; The preset conditions include: the electronic device is in a target posture, and the expression of the subject in the captured preview image is a predetermined expression; the target posture is a landscape posture or a portrait posture. Save the captured image.
2. The method according to claim 1, characterized in that, The preset conditions also include one or more of the following conditions: The distance between the electronic device and the subject being photographed is within a preset distance range; The electronic device is held in a predetermined manner; wherein, the predetermined holding method differs depending on the target posture of the electronic device. The positional change of the electronic device is greater than a threshold.
3. The method according to claim 1, characterized in that, The method further includes: Determine whether the electronic device is in the target posture; With the electronic device in the target posture, the preview image is captured using the camera; If the expression of the subject in the preview image is the predetermined expression, then the preset condition is determined to be met.
4. The method according to claim 3, characterized in that, The camera is configured in low-power mode.
5. The method according to any one of claims 1-4, characterized in that, The preview image is a frame from the preview stream captured by the camera in low-power mode.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Extract the facial expression features of the subject in the preview image; The facial expression features of the photographed subject are matched with the predetermined facial expression to obtain the facial expression similarity. If the expression similarity is greater than or equal to the expression threshold, the expression of the subject in the preview image is determined to be the predetermined expression.
7. The method according to any one of claims 1-6, characterized in that, The electronic device also includes an expression detection service, and the method further includes: The preview image captured by the camera is obtained through the facial expression detection service; The facial expression detection service identifies the facial expression of the subject in the preview image as a predetermined expression.
8. An electronic device, characterized in that, The electronic device includes a camera, a memory, one or more processors, and a display screen; the camera, the display screen, the memory, and the processor are coupled; wherein the camera is used to capture images, the display screen is used to display the captured images, the memory stores computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the shooting method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to perform the shooting method as described in any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product includes a computer-readable storage medium storing a computer program that, when executed by at least one processor, causes the at least one processor to perform the shooting method as described in any one of claims 1-7.