Capturing method, electronic device and storage medium
By predicting the position of the subject in subsequent image frames and adjusting the camera focus, the clarity problem of electronic devices when shooting moving objects is solved, achieving higher shooting effects and user experience.
Patent Information
- Application Number
- PCT/CN2024/142905
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-17
AI Technical Summary
During the shooting process, especially when moving objects, it is difficult for existing electronic devices to accurately track the subject, resulting in blurred images and affecting the user experience.
By predicting the position of the photographing subject in subsequent image frames, adjusting the focus of the camera, adopting the active focus pursuit method, combining depth data and phase data, accurately adjusting the focus position, and displaying the detection frame to indicate the position of the photographing subject.
It improves the shooting clarity in sports scenes, reduces the deviation between the detection frame and the shooting subject, and improves the user's shooting experience.
Smart Images

Figure CN2024142905_17072025_PF_FP_ABST
Abstract
Description
Shooting method, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 10, 2024, with application number 202410042523.6 and invention name “A shooting method, electronic device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of terminal technology, and in particular to a shooting method, electronic device and storage medium. Background Art
[0003] With the development of electronic devices, the functions they provide are also increasing. The shooting function is one of the most frequently used functions in electronic devices. Users can use the shooting function of electronic devices to shoot images or videos. As people's demand for the shooting effects of electronic devices continues to increase, the shooting functions of electronic devices are becoming more and more diverse. The shooting function of electronic devices has a variety of shooting modes. For example, portrait mode, night scene mode, and other shooting modes.
[0004] However, the shooting functions provided by current electronic devices are not yet perfect. When shooting a subject, the subject may be blurred in the captured image, affecting the user's shooting experience. Summary of the Invention
[0005] The present application provides a shooting method, an electronic device, and a storage medium, which can track and focus on a shooting subject during the shooting process, thereby improving the clarity of the shooting subject in the captured image and enhancing the user's shooting experience.
[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, a shooting method is provided, the method comprising: receiving a first operation from a user, the first operation being used to instruct to turn on a shooting function of a target application. In response to the first operation, a first shooting interface of the target application is displayed, the first shooting interface including a first preview window, the first preview window being used to display a first image frame captured by a camera. In response to a second operation from the user, a motion focus mode is entered, and after detecting a target object, a second shooting interface of the target application is displayed, the second shooting interface including a second preview window and a first detection frame, the second preview window being used to display a second image frame captured by the camera, the first detection frame being predicted based on a first image position of the target object in the first image frame, and the first detection frame being used to indicate a second image position of the target object in the second image frame.
[0008] In this way, the electronic device predicts the second image position of the subsequent image frame based on the detected first image position, and then displays the detection frame in the shooting interface based on the predicted second image position. This allows the detection frame to accurately indicate the image position of the subject in the image frame, reducing the deviation between the detection frame and the subject. Even if the subject is moving, the electronic device's detection frame can accurately follow the subject and focus, improving the shooting experience of moving scenes.
[0009] In a possible implementation of the first aspect, the focus position of the camera is obtained based on the second image position of the target object in the second image frame, and the focus of the camera is adjusted according to the focus position of the camera. In this way, the electronic device can predetermine the focus position that the camera will use when capturing image frames later based on the predicted second image position, and the electronic device can adjust the focus of the camera in advance so that the camera focuses on the focus position. In this way, the impact of delays in image frame detection, focus position calculation, etc. on the focus tracking effect is reduced. Even when the subject is a moving object, the camera of the electronic device can accurately follow the subject to focus, thereby improving the shooting experience of moving scenes.
[0010] In another possible implementation manner of the first aspect, the second image frame is the i-th image frame acquired after the first image frame, where i is a positive integer.
[0011] In another possible implementation of the first aspect, phase data and / or depth data of the second image frame is obtained, and the acquisition time of the depth data of the second image frame is consistent with the acquisition time of the second image frame. The focus position is obtained based on the phase data and / or depth data of the second image frame and the second image position of the target object in the second image frame. In this way, the electronic device can predict the focus position of the camera based on the depth data and / or phase data corresponding to the second image position, thereby further improving the accuracy of the camera focus and enabling the camera to focus on the target object.
[0012] In another possible implementation of the first aspect, the electronic device includes a laser device configured to provide depth data corresponding to a preset number of image regions (e.g., 40×30). In this manner, the electronic device can obtain accurate depth data of the target object through the laser device.
[0013] In another possible implementation of the first aspect, the electronic device collects depth data of image frames at a first preset frame rate using a laser device. The first preset frame rate is 60 frames per second.
[0014] In another possible implementation of the first aspect, when the target object is a moving object, the first detection frame includes multiple sub-frames, each of which covers a display area corresponding to the moving object. The electronic device uses the first detection frame formed by the multiple sub-frames to indicate the moving object in a shooting interface, thereby improving the recognition of the moving object, thereby better notifying the user of the moving object in the shooting interface and enhancing the user's shooting experience.
[0015] In another possible implementation of the first aspect, if the moving object includes a human face, the multiple subframes of the first detection frame cover the display area other than the moving object's face. In this implementation, the first detection frame displayed in the capture interface of the electronic device avoids the human face, thereby providing the user with a clear, unobstructed view of the human face and improving the user experience.
[0016] In another possible implementation of the first aspect, the electronic device determines, based on the depth data of the moving object in the second image frame, multiple image areas corresponding to the moving object in a preset number of image areas corresponding to the second image frame, and the laser device of the electronic device is used to provide depth data of a preset number of image areas, such as providing depth data of 40×30 image areas. The sizes of these image areas can be the same and evenly arranged. Furthermore, the electronic device generates multiple sub-frames included in the first detection frame based on the multiple image areas corresponding to the moving object. The depth data of the moving object is the depth data with the smallest distance in the depth data corresponding to the second image area. In this way, the first detection frame provided by the electronic device can more accurately indicate the image position of the moving object.
[0017] In another possible implementation of the first aspect, the number of sub-frames in the first detection frame is less than or equal to a preset number. For example, the preset number is 27, 20, or other values. If the number of sub-frames in the first detection frame is too large, it will affect the aesthetics of the shooting interface and the user's perception.
[0018] In another possible implementation of the first aspect, the preset number is equal to 1 / 16 of the total number of preset coordinate grids at the full-frame ratio. At the full-frame ratio, the electronic device displays the preview window in full screen, and the preset coordinate network is used to evenly divide the display area of the preview window.
[0019] In another possible implementation of the first aspect, the camera is a wide-angle camera or a telephoto camera. This implementation is applicable to zoom scenarios of the wide-angle camera or the telephoto camera, such as zoom tracking focus with a zoom ratio of 1-5 times for the wide-angle camera or the telephoto camera.
[0020] In another possible implementation of the first aspect, the electronic device displays a function setting window in the shooting interface, which includes a function key for the motion focus mode. When the function key for the motion focus mode is selected, the target application is in the motion focus mode.
[0021] In another possible implementation of the first aspect, the function settings window further includes a function key for subject focus mode. The method further includes: in response to a second operation, switching the target application from motion focus mode to subject focus mode; in subject focus mode, the electronic device prioritizes focusing on faces.
[0022] In another possible implementation of the first aspect, the target application also includes a snapshot mode. If the target application is in snapshot mode, the target application uses the motion focus mode for focusing. If the target application is in non-snapshot mode, the target application uses the subject focus mode for focusing. Of course, the user can also set the combination of the snapshot mode and the focus mode according to their own preferences or shooting needs. For example, when the snapshot mode is turned on, the subject focus mode is selected. When the snapshot mode is not turned on, the motion focus mode is selected.
[0023] In another possible implementation of the first aspect, in motion focus mode, the priority order of focus subjects is: object indicated by user operation (which may be referred to as a manually triggered object) > moving object > registered human face > unregistered human face > human body > cat or dog. For example, if, in motion focus mode, the electronic device does not detect a moving object but detects a human face, and that face is a registered face, the electronic device will perform focus tracking on the registered face. A first detection frame, such as a double-lined yellow frame, is displayed in the display area for focus tracking on the registered face.
[0024] In another possible implementation of the first aspect, in the subject focus mode, the priority order of the focus subject is: object indicated by user operation (which may be called manual trigger object) > registered face > unregistered face > human body > cat = dog.
[0025] In a second aspect, the present application provides an electronic device comprising: a camera, a memory, and a processor; the memory and a display are respectively coupled to the processor. The camera is configured to capture image frames. The memory stores computer program code, which includes computer instructions. When executed by the processor, the electronic device performs the method described in the first aspect and any possible implementation thereof.
[0026] In a third aspect, the present application provides a computer-readable storage medium comprising computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method described in the first aspect and any possible implementation thereof.
[0027] In a fourth aspect, the present application provides a computer program product comprising program instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any possible implementation thereof. For example, the computer may be the electronic device described above.
[0028] In a fifth aspect, the present application provides a chip system, which is applied to an electronic device. The chip system includes an interface circuit and a processor. The interface circuit and the processor are interconnected via a circuit. The interface circuit is configured to receive signals from a memory and send signals to the processor, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device executes the method described in the first aspect and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a schematic diagram of an electronic device capturing an image of a moving object provided by an embodiment of the present application;
[0030] FIG2 is a hardware structure block diagram of a mobile phone 100, an example of an electronic device provided in an embodiment of the present application;
[0031] FIG3 is a software structure block diagram of a mobile phone 100, an example of an electronic device provided in an embodiment of the present application;
[0032] FIG4 is a timing diagram of a focusing process provided by an embodiment of the present application;
[0033] FIG5 is a flow chart of a photographing method provided in an embodiment of the present application;
[0034] FIG6 is a schematic diagram of an interface for setting a focus mode according to an embodiment of the present application;
[0035] FIG7 is a schematic diagram of a first detection frame provided in an embodiment of the present application;
[0036] FIG8 is a schematic diagram of a preview window at different aspect ratios provided by an embodiment of the present application;
[0037] FIG9 is a schematic diagram of another first detection frame provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] To meet users' photography needs, electronic devices are equipped with photography functions. With the miniaturization and portability of electronic devices, users can use electronic devices with photography functions to take photos anytime and anywhere. Consequently, users' requirements for photography functions are becoming increasingly sophisticated. In response to users' increasing demands for photography functions, electronic devices provide a variety of photography modes, such as portrait mode, night scene mode, and high-dynamic range (HDR) mode, to meet users' diverse photography needs.
[0039] In the actual shooting process, many factors affect the shooting effect. For example, focus speed, shutter delay, image processing, etc. all affect the shooting effect. To this end, electronic devices can improve the shooting effect to a certain extent through some image processing methods. For example, electronic devices use HDR mode to synthesize multiple captured image frames. During the synthesis process, the noise in the multiple image frames can be filtered out and the details of the multiple image frames are combined to obtain the final image presented to the user.
[0040] To further enhance photography, electronic devices also offer a tracking focus mode. In this mode, the electronic device tracks and focuses on the subject. The subject is the primary object, or target, of the electronic device's photography. To achieve a clear image of the subject, the electronic device focuses on it. The following describes the tracking focus process of an electronic device using an example.
[0041] After capturing an image frame, the electronic device determines the image region within the frame where the subject is located and calculates the focus position corresponding to the subject within the frame based on the image region where the subject is located. The image region where the subject is located can be referred to as a region of interest (ROI). If the deviation between the focus position corresponding to the subject and the focus position currently used by the electronic device (i.e., the focus of the camera) is greater than a preset threshold, this indicates that the current focus position used by the electronic device is significantly different from the focus position corresponding to the subject. If the electronic device remains at the current focus position, it will be difficult for the electronic device to achieve a clear image of the subject. In this case, the electronic device adjusts the focus of the camera to move the camera's focus to the focus position corresponding to the subject to track the focus of the subject. If the deviation between the focus position corresponding to the subject and the current focus position used by the electronic device is less than or equal to a preset threshold, this indicates that the current focus position used by the electronic device is slightly different from the focus position corresponding to the subject. If the electronic device remains at the current focus position, the electronic device can achieve a clear image of the subject. In this case, the electronic device does not adjust the camera's focus and continues to maintain the camera's focus at the current focus position, using the current focus position to track the focus of the subject. This type of focus tracking method can be called passive focus tracking method.
[0042] However, during the shooting process of an electronic device, due to delays in image frame detection and focusing, the ROI determined by the electronic device may deviate from the actual position of the subject in the image frame. The detection frame displayed by the electronic device in the shooting interface is difficult to accurately indicate the image position of the subject, and the camera of the electronic device is difficult to follow the subject and focus. The subject in the image presented to the user by the electronic device may be blurred, especially when the subject is a moving object, the moving object in the image is difficult to be clearly imaged. For example, if the subject of the electronic device is a moving car, the electronic device follows the car and focuses. The image shot by the electronic device is shown in Figure 1. It can be seen that the image of the car in the image shot by the electronic device is blurred and not clear enough.
[0043] The embodiment of the present application also provides another focusing solution for improving the accuracy of focus tracking and enhancing the shooting effect. Specifically, the electronic device displays the image frames captured by the camera in the preview window of the shooting interface. When the electronic device captures the first image frame, it detects the first image frame and determines the image area where the shooting subject is located in the first image frame. The image area where the shooting subject is located is the ROI. The ROI detected by the electronic device in the first image frame and the ROI detected before the first image frame can provide the movement trend of the shooting subject. Furthermore, the electronic device predicts the ROI corresponding to the second image frame based on the ROI detected in the first image frame. The second image frame is an image frame captured after the first image frame. The electronic device further displays a detection frame for indicating the shooting subject in the shooting interface based on the ROI corresponding to the second image frame.
[0044] In this way, the electronic device uses the detected ROI to predict the ROI of the next image frame, and then displays the detection frame in the shooting interface based on the predicted ROI. This detection frame accurately indicates the image position of the captured subject in the image frame, reducing the deviation between the detection frame and the subject. Even if the subject is moving, the electronic device's detection frame can accurately follow the subject and focus, improving the shooting experience of moving scenes.
[0045] For example, the electronic devices described in the embodiments of the present application may be mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) and virtual reality (VR) devices, media players, wearable devices, and the like. The embodiments of the present application do not impose any particular restrictions on the specific form of the electronic devices.
[0046] In the embodiment of the present application, the electronic device is a mobile phone 100 as an example, and the hardware structure of the electronic device is described through the mobile phone 100. As shown in Figure 2, the mobile phone 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195.
[0047] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), a driver processor, etc. Different processing units may be independent devices or integrated into one or more processors. The processor 110 may be the nerve center and command center of the mobile phone 100. The processor 110 may generate an operation control signal based on the instruction opcode and timing signal to complete the control of instruction fetching and execution.
[0048] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0049] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage. For example, the electronic device can save captured images to the external memory card.
[0050] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. For example, in an embodiment of the present application, the processor 110 can execute instructions stored in the internal memory 121, and the internal memory 121 can include a program storage area and a data storage area. The electronic device can also save the captured images in the internal memory.
[0051] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. In some embodiments, the power management module 141 and the charging management module 140 can also be provided in the same device.
[0052] The sensor module 180 may include sensors such as a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a Hall sensor, a touch sensor, an ambient light sensor, and a bone conduction sensor. The mobile phone 100 may collect various data through the sensor module 180. For example, the mobile phone 100 may collect touch data through the touch sensor in the sensor module 180, thereby allowing the mobile phone 100 to identify user operations based on the touch data.
[0053] Mobile phone 100 implements display functionality through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0054] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, or a quantum dot light-emitting diode (QLED). For example, mobile phone 100 displays a camera interface, captured images, and the like through display screen 194.
[0055] In some implementations, the touch sensor can be disposed within the display screen 194 , with the touch sensor and display panel forming a touch screen, also known as a "touch screen." The touch sensor, also known as a "touch panel," is configured to detect touch operations, such as clicks and slides, applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to an application processor to determine the type of touch event. The mobile phone 100 can provide visual output related to the touch operations via the display screen 194 .
[0056] The mobile phone 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display 194 and application processor.
[0057] The ISP processes data fed back by camera 193. For example, when the mobile phone 100 is shooting, the camera shutter opens, and light is transmitted through the lens to the camera's photosensitive element. The camera's photosensitive element converts the light signal into an electrical signal, which is then passed to the ISP for processing and converted into an image visible to the human eye. The ISP can also perform algorithmic optimization on image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0058] The camera 193 is used to capture static or dynamic images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the mobile phone 100 may include one or more cameras 193, such as the mobile phone 100 includes a front camera and a rear camera.
[0059] The mobile phone 100 may also include a time of flight (TOF) sensor. The TOF sensor is used to measure the time it takes for a light pulse to travel between the ToF sensor and the object being measured (such as referred to as TOF data). The TOF data is used to generate distance depth data of the object being measured. In an embodiment of the present application, the TOF sensor may use a laser device to measure the distance between the subject and the lens so that the camera 193 can achieve autofocus. The time of flight sensor can be set inside the camera 193 or outside the camera 193.
[0060] It should be understood that the interface connection relationship between the modules illustrated in this embodiment is merely illustrative and does not constitute a structural limitation on the electronic device. In other embodiments, the electronic device may include more or fewer modules than those provided in the above embodiment, and the modules may also adopt different interface connection methods from those in the above embodiment, or a combination of multiple interface connection methods. The methods in the following embodiments can all be implemented in an electronic device having the above hardware structure.
[0061] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. In the embodiment of the present application, the Android system of the layered architecture and the electronic device being the mobile phone 100 are taken as an example to illustrate the software structure of the electronic device.
[0062] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system may include an application layer, an application framework layer, an Android runtime and system libraries, a hardware abstraction layer (HAL), and a kernel layer.
[0063] The application layer may include a series of application packages. For example, the application package may include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video and other applications, which are not limited in this embodiment of the present application.
[0064] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, content provider, view system, telephony manager, resource manager, notification manager, and sensor manager, etc., but the embodiments of the present application do not impose any restrictions on this.
[0065] The Android runtime consists of core libraries and a virtual machine (VM). The Android runtime is responsible for scheduling and management of the Android system. The core libraries consist of two parts: one containing the Java language's callable functions and the other the Android core library. The application layer and the application framework layer run in the VM. The VM executes the Java files in the application layer and application framework layer as binary files. The VM is responsible for performing functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0066] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0067] The Hardware Abstraction Layer (HAL) encapsulates Linux kernel drivers, provides interfaces to the kernel, and shields the underlying hardware from implementation details. For example, the HAL might include a camera HAL, a Wi-Fi HAL, or a Bluetooth HAL.
[0068] The kernel layer is the layer between hardware and software. It includes drivers and system services. Drivers include at least display drivers, camera drivers, audio drivers, and sensor drivers.
[0069] FIG3 is a structural block diagram of the mobile phone 100 according to an embodiment of the present application. The following uses the application processor of the mobile phone 100 having the layered architecture shown in FIG3 to exemplify the software and hardware workflow of the mobile phone 100.
[0070] In this example, a user operates mobile phone 100 to open the camera application. In response to the user's operation to open the camera application, the camera application issues a camera start instruction via the camera access interface in the application framework layer. The camera start instruction issued by the camera application is then transmitted to camera 193 via the camera access interface, the camera hardware abstraction layer, and the camera driver. In response to the camera start instruction, camera 193 of mobile phone 100 begins capturing image frames and transmits the image frames to the image signal processor. The image signal processor performs image processing on the image frames transmitted by camera 193, such as image front-end processing (IFE) such as color adjustment and noise reduction, and generates preview image frames through the image processing engine (IPE). The image signal processor further transmits the processed image frames to the camera hardware abstraction layer via the camera driver. The camera hardware abstraction layer includes camera-related algorithm modules, such as a detection module and a focus module. The detection module is used to detect image frames and determine the image frame's ROI. The focus module is used to determine the camera's focus position. Furthermore, the camera hardware abstraction layer reports the ROI position to the camera application via the camera access interface. The camera application can perform subsequent shooting processes based on the position of the ROI reported by the camera hardware abstraction layer, such as displaying a detection frame in the shooting interface.
[0071] In the following embodiments, the electronic device is taken as a mobile phone 100 as an example to introduce the method provided by the embodiments of the present application.
[0072] As described above, during the capture process of an electronic device, due to delays in image frame detection, focusing, and other aspects, the ROI determined by the electronic device may deviate from the actual position of the subject in the image frame, making it difficult for the electronic device's camera to focus on the subject. To facilitate understanding of the solutions provided in the embodiments of this application, the following first describes, with reference to the accompanying drawings, the reasons why the ROI determined by the electronic device deviates from the actual position of the subject in the image frame.
[0073] Taking a mobile phone 100 as an example, as shown in FIG4 , the mobile phone 100 periodically captures image frames using a camera, for example, at a frame rate of 30 frames per second (FPS). The horizontal axis of FIG4 represents time. The camera captures image frames in a row-by-row exposure manner. That is, the camera begins exposing the first row of the image frame, and after the first row is exposed, the camera reads the first row of data. After a row exposure cycle, the second row begins to be exposed, and after the second row is exposed, the camera reads the second row of data, and so on, until the entire image frame is fully read. Therefore, the image frames captured by the camera appear as parallelograms in FIG4 . The position of the ROI in the image frame changes over time. The camera transmits the read image frames to the image signal processor, which processes the image signal processor in the mobile phone 100. From the start of camera exposure to the reading of an image frame, a certain amount of time is consumed, for example, 0-16 milliseconds. The image signal processor also consumes a certain amount of time to process the image frames, for example, 19 milliseconds. It takes a certain amount of time, such as 1 millisecond, for the image signal processor to transmit the processed image frame to the camera hardware abstraction layer. The camera hardware abstraction layer uses the detection algorithm provided by the detection module to detect the image frame and determine the ROI of the image frame. It takes a certain amount of time, such as 11 milliseconds, for the detection algorithm in the camera hardware abstraction layer to determine the ROI of the image frame. After obtaining the ROI of the image frame, the detection module transmits the ROI of the image frame to the focus module. It takes a certain amount of time, such as 25 milliseconds, for the detection module to transmit the ROI of the image frame to the focus module. It can be seen that the time delay from the camera capturing an image frame (such as image frame 1) to the focus module obtaining the ROI of the image frame is approximately 54 milliseconds to 70 milliseconds. At this time, the camera of the mobile phone 100 has started to capture or is about to capture the third image frame (i.e., image frame 3).
[0074] If the focus module of the mobile phone 100 uses the ROI detected by the detection module to calculate the focus position at this time, that is, focusing is performed based on the ROI detected in the first image frame, since the position of the ROI changes over time, the position of the ROI in image frame 3 has deviated from the position of the ROI in image frame 1. Then, when the camera captures image frame 3, the focus position corresponding to the ROI of image frame 1 is used, and it is difficult for the camera to capture a clear subject.
[0075] In view of this, in this embodiment, the mobile phone 100 uses the ROI of image frame 1 and the ROI detected in the image frame before image frame 1 to predict the ROI of the image frame collected after image frame 1, such as predicting the ROI of n (such as 4) image frames collected after image frame 1. Wherein, n is a positive integer. Furthermore, the mobile phone 100 determines the focus position of the camera based on the predicted ROI, so that the camera tracks the focus of the subject. This focusing method can be called active focusing. Active focusing can reduce the impact of the above-mentioned time delay on the focus position of the camera, and has a more accurate focusing effect even if the subject is a fast-moving moving object.
[0076] In view of the fact that the above-mentioned active focus tracking method has a relatively accurate focus tracking effect even when the shooting subject is a moving object, in some implementations, the mobile phone 100 can set a focus mode for scenes where the shooting subject is a moving object. For example, the shooting function of the mobile phone 100 can have a motion focus mode and a subject focus mode. The motion focus mode is a mode that prioritizes focusing on moving objects. In the motion focus mode, the mobile phone 100 uses an active focus tracking method to track the shooting subject. The subject focus mode is a mode that prioritizes focusing on faces. In the subject focus mode, the position of the shooting subject changes little, and the mobile phone 100 uses the above-mentioned passive focus tracking method to track the shooting subject. Of course, regardless of whether the shooting subject is a moving object, the mobile phone 100 can use the above-mentioned active focus tracking method to track the focus. The embodiments of the present application are not limited to this.
[0077] The following uses the example of the mobile phone 100 using the active tracking focus mode in the motion focus mode to illustrate the method provided by the embodiment of the present application. As shown in Figure 5, the method provided by the embodiment of the present application includes the following steps:
[0078] S501 , the mobile phone 100 activates a shooting function of a target application in response to a first operation, captures a first image frame through a camera, and displays the first image frame in a shooting interface of the target application.
[0079] Mobile phone 100 has installed a target application that provides a shooting function. For example, the target application is an application with a shooting function, such as a camera application or a chat application. In response to the first operation, mobile phone 100 activates the shooting function of the target application, begins capturing image frames through the camera, and displays the shooting interface of the target application.
[0080] It is understandable that the mobile phone 100 periodically captures image frames through the camera, such as capturing image frames at 30 FPS. The first image frame is an image frame captured by the mobile phone 100 through the camera at the first moment.
[0081] The first operation is used to instruct the target application to be activated. For example, the first operation is a user operation in which a user clicks on the target application icon. Of course, the first operation can also be other user operations instructing the activation of the shooting function of the target application, such as a user operation of pressing a physical button on the mobile phone 100 that activates the shooting function. The embodiments of the present application do not limit the specific implementation of the first operation.
[0082] The target application's shooting interface includes a preview window. The preview window is used to display image frames captured by the mobile phone 100 through the camera. Through the preview window, the user can see the scene currently being captured by the mobile phone 100. The user can adjust the shooting angle of the mobile phone 100 based on the image displayed in the preview window to accurately capture the subject. The shooting interface of the electronic device includes a first shooting interface, which includes a first preview window. The first preview window is used to display the first image frame captured by the camera.
[0083] In addition to the preview window, the shooting interface also includes various function keys or controls. For example, function keys for various shooting modes, such as aperture, night scene, portrait, photo, video, movie, and professional, are located at the bottom and top of the shooting interface. Images captured by the mobile phone 100 in different shooting modes may differ in color, clarity, and imaging. Users can select any shooting mode to capture an image.
[0084] S502 : In a motion focus mode, the mobile phone 100 determines a target object in a first image frame and a first image position of the target object in the first image frame.
[0085] After the mobile phone 100 activates the shooting function, it enters the motion focus mode in response to the user's second operation. The shooting interface of the target application is provided with a function key for the motion focus mode and / or a function key for the subject focus mode. The user can select a focus mode based on actual shooting needs and the shooting scene. For example, the mobile phone 100 displays a function bar in the shooting interface of the target application, and the function bar includes a function key for the motion focus mode. The mobile phone 100 receives the second operation of the user clicking the function key for the motion focus mode in the function bar and enters the motion focus mode.
[0086] Exemplarily, as shown in FIG6 , a pull-up function key is provided at the bottom of the shooting interface (or referred to as the shallow level). If the mobile phone 100 receives a user operation of clicking or sliding up the pull-up function key, the mobile phone 100 pops up a function bar (or referred to as a function setting window) in the shooting interface. A focus mode function key is provided in the function bar. If the mobile phone 100 receives a click operation of the user on the focus mode function key, the mobile phone 100 displays a function key for the motion focus mode (such as “motion focus priority”) and a function key for the subject focus mode (such as “subject focus priority”) in the function bar. If the mobile phone 100 receives a user operation of the user selecting the motion focus mode (an example of the second operation), the mobile phone 100 sets the focus mode of the shooting function to the motion focus mode. Of course, if the mobile phone 100 receives a user operation of the user selecting the subject focus mode, the mobile phone 100 sets the focus mode of the shooting function to the subject focus mode.
[0087] In the embodiment of the present application, after each image frame is captured by the camera of the mobile phone 100, the captured image frame is transmitted to the image signal processor. Taking the first image frame as an example, the image signal processor performs image processing on the first image frame and transmits the processed first image frame to the detection module. The detection module of the mobile phone 100 detects the first image frame using a detection algorithm, such as a target detection algorithm, a single target tracking algorithm, an optical flow method, a frame difference method and other detection algorithms, analyzes the changes in pixels in adjacent image frames, and determines the target object (i.e., the shooting subject) in each first image frame and the first image position of the target object in the first image frame (i.e., the ROI of the first image frame).
[0088] S503 : The mobile phone 100 predicts a second image position of the target object in the second image frame according to the first image position of the target object in the first image frame.
[0089] Exemplarily, the detection module of the mobile phone 100 determines the moving trajectory of the target object based on the first image position of the target object in the first image frame and the third image position of the target object in one or more third image frames (such as 8 third image frames) collected before the first moment. The moving trajectory is used to indicate the moving trend (or movement trend) of the moving object. For example, the moving trajectory includes a moving direction and a moving speed. Furthermore, the detection module of the mobile phone 100 predicts the second image position of the moving object in the second image frame collected after the first image frame based on the moving trajectory of the target object, that is, predicts the image position of the target object after the first moment. The second image frame is the i-th image frame collected after the first image frame, where i is a positive integer. For example, the second image frame is the 1st image frame, the 2nd image frame, or the 3rd image frame collected after the first image frame.
[0090] In another example, the detection module of the mobile phone 100 can use a preset machine learning model to predict the second image position of the target object in the second image frame. Specifically, the detection module of the mobile phone 100 can input the first image frame marked with the image position of the target object and multiple third image frames into the trained preset machine learning model to obtain the prediction result output by the preset machine learning model. The first image frame and the multiple third image frames input into the preset machine learning model are arranged in chronological order. The prediction result output by the preset machine learning model represents the image position of the target object after the first moment, such as outputting the second image position of the target object in the second image frame collected after the first moment, or outputting the image position of the target object in multiple image frames (such as 4 image frames) collected after the first moment. The image positions in these multiple image frames include the second image position in the second image frame.
[0091] The preset machine learning model can be obtained by training a machine learning model with a specific model structure using the collected historical image frames and the positions of the reference objects in the historical image frames by the mobile phone 100. Alternatively, the preset machine learning model is obtained by training other electronic devices. The mobile phone 100 can directly install the trained preset machine learning model. Through the preset machine learning model, the mobile phone 100 can more accurately predict the image position of the target object at a future moment. The model structure of the preset machine learning model is not limited in the embodiments of the present application. For example, the preset machine learning model can adopt a model structure such as a support vector machine, a convolutional neural network, etc.
[0092] S504 , the mobile phone 100 determines the focus position of the camera based on the second image position of the target object in the second image frame, and displays a first detection frame on the shooting interface of the target application.
[0093] In an embodiment of the present application, after predicting the second image position of the target object in the second image frame, the detection module of the mobile phone 100 transmits the second image position of the target object in the second image frame to the focus module. The focus module of the mobile phone 100 determines the focus position of the camera based on the second image position of the target object in the second image frame. For example, the focus module of the mobile phone 100 can use an autofocus (AF) algorithm to determine the focus position of the camera, and further adjust the focus of the camera according to the focus position of the camera, such as the camera motor pushing the lens focus of the camera to the focus position, so that the position where the target object is located can be clearly imaged.
[0094] In addition, the mobile phone 100 also generates a first detection frame based on the second image position of the target object in the second image frame, and displays the first detection frame in the shooting interface of the target application. For example, the mobile phone 100 can generate a first detection frame based on the depth data corresponding to the second image position and the second image position, and display the first detection frame in the shooting interface of the target application. The first detection frame is used to indicate the second image position of the target object in the second image frame. The first detection frame is predicted based on the first image position of the target object in the first image frame, so that even due to delays in image frame detection and other aspects, the first detection frame can accurately indicate the image position of the target object in the shooting interface. At this time, the shooting interface of the electronic device can be a second shooting interface, and the second shooting interface includes a second preview window and a first detection frame, and the second preview window is used to display the second image frame captured by the camera. The first detection frame can be a focus frame.
[0095] In order to further improve the accuracy of the mobile phone 100 focusing on the target object, in some implementations, the mobile phone 100 can also obtain the phase data and / or depth data of the second image frame, and further determine the focus position of the camera when capturing the next frame of the second image frame based on the phase data and / or depth data of the second image frame and the second image position of the target object in the second image frame.
[0096] In one example of this implementation, the mobile phone 100 is configured with a laser device. The laser device is used to measure the time it takes for a laser pulse to travel back and forth between the laser device and the object being measured, thereby obtaining TOF data. The TOF data can be converted into depth data of the distance between the object being measured and the laser device. The focus module of the mobile phone 100 can obtain the depth data corresponding to the second image position of the second image frame, that is, obtain the depth data of the target object at the acquisition moment of the second image frame (such as the second moment). Further, the focus module determines the focus position of the camera when acquiring the next frame of the second image frame (such as the third moment) based on the depth data corresponding to the second image position of the second image frame. For example, the focus module of the mobile phone 100 can predict the depth data of the target object at the third moment based on the depth data of the target object at the second moment and the depth data of the target object before the second acquisition moment. Further, the focus module of the mobile phone 100 determines the focus position of the camera when acquiring the next frame of the second image frame based on the depth data of the target object at the third moment.
[0097] In another example of this implementation, the focus module of the mobile phone 100 can obtain the phase (phase detection, PD) data (or phase difference data) of the second image frame. The mobile phone 100 further obtains the phase data corresponding to the second image position based on the second image position of the second image frame. Furthermore, the focus module determines the focus position of the camera when capturing the next frame of the second image frame based on the phase data corresponding to the second image position of the second image frame. For example, the focus module of the mobile phone 100 can determine the focus position corresponding to the second image frame based on the phase data corresponding to the second image position of the second image frame. Further, the focus module of the mobile phone 100 determines the focus position change trend based on the focus position corresponding to the second image frame and the focus position corresponding to the first image frame. The mobile phone 100 then predicts the focus position of the camera when capturing the next frame of the second image frame based on the focus position change trend.
[0098] In another example of this implementation, the focus module of the mobile phone 100 can obtain the phase data and depth data corresponding to the second image position of the second image frame based on the second image position of the second image frame. The focus module of the mobile phone 100 further predicts the focus position of the camera when capturing the next frame of the second image frame based on the phase data and depth data corresponding to the second image position of the second image frame. For example, the focus module of the mobile phone 100 predicts the focus position corresponding to the next frame of the second image frame (such as called the first focus position) based on the phase data corresponding to the second image position of the second image frame. The focus module of the mobile phone 100 predicts the focus position corresponding to the next frame of the second image frame (such as called the second focus position) based on the depth data corresponding to the second image position of the second image frame. Further, the focus module of the mobile phone 100 determines the focus position that the camera ultimately uses when capturing the next frame of the second image frame based on the first focus position and the second focus position, such as taking a weighted average of the first focus position and the second focus position.
[0099] In the embodiment of the present application, the mobile phone 100 can obtain the depth data and / or phase data corresponding to the second image position by predicting the second image position, and then predict the focus position to be adopted by the camera when capturing the next frame of the second image frame based on the depth data and / or phase data corresponding to the second image position. In this way, the accuracy of camera focus can be further improved, allowing the camera to focus on the target object.
[0100] It is understandable that since there is a certain time delay between the focusing module acquiring the second image position transmitted by the detection module, the camera of the mobile phone 100 has already captured the image frame after the first image frame when the focusing module acquires the second image position. At this time, if the focusing module of the mobile phone 100 uses the first image position detected by the first image frame to focus, for a target object with a faster movement speed, especially for a target object that moves quickly in the direction of movement of the camera lens, it is difficult for the camera of the mobile phone 100 to focus on the target object, and the target object in the captured second image is blurred. Therefore, in order to cope with the rapid movement of the target object, the mobile phone 100 uses the predicted second image position to adjust the focus in advance to achieve accurate focus on the target object.
[0101] As shown in Figure 4, the first image frame is image frame 1, and the second image frame is image frame 3. When the focus module obtains the second image position (predicted) of image frame 3 transmitted by the detection module, the mobile phone 100 has already captured image frame 2 but has not yet captured image frame 3. After the mobile phone 100 exposes and reads the second image position of image frame 3, because the detection module takes a certain amount of time to detect image frame 3, the mobile phone 100 does not immediately obtain the detection result of image frame 3 (i.e., the detected image position of the target object in image frame 3) after reading image frame 3. The focus module begins to predict the camera's focus position based on the second image position (predicted) of image frame 3, depth data, and / or phase data. The depth data corresponding to image frame 3 is collected at the same time as the acquisition time of image frame 3. That is, the time when the first row of image data of image frame 3 is read after exposure coincides with the time when the depth data of image frame 3 is read after exposure. The depth data collected by the laser device is collected using a global exposure method, that is, all depth data corresponding to image frame 3 is collected in one go. It takes a certain amount of time, such as 2-3 milliseconds, for the TOF data measured by the laser device to be converted into depth data. It also takes a certain amount of time, such as 2-4 milliseconds, for the focus module to obtain the depth data corresponding to image frame 3. The focus module predicts the focus position of the camera based on the depth data and / or phase data corresponding to the second image position of image frame 3, which takes a certain amount of time, such as 5 milliseconds. It takes a certain amount of time, such as 0.3 milliseconds, for the focus module to transmit the predicted focus position to the camera. The camera motor pushes the camera lens according to the focus position, which takes a certain amount of time, such as 10 milliseconds. At this time, the mobile phone 100 has not yet started to capture image frame 4, but the focus of the camera has moved to the focus position corresponding to image frame 4. In this way, the mobile phone 100 can pre-adjust the focus of the camera so that the camera lens is pushed to the focus position corresponding to image frame 4 in advance. When the mobile phone 100 captures the image frame 4, the lens of the camera has been pushed to the focus position corresponding to the image frame 4 in advance, thereby achieving focus on the target object.
[0102] It is understood that, considering the rapid movement of the target object, in order to improve focusing accuracy, in motion focus mode, the mobile phone 100 can use the above-mentioned active focus method to predict the focus position corresponding to each image frame. The camera of the mobile phone 100 controls the motor to move the lens according to the focus position corresponding to each image frame. This focusing method can also be called frame-by-frame focusing.
[0103] In some implementations, in order to provide valid depth data, the frame rate of the depth data collected by the mobile phone 100 (which may be referred to as a first preset frame rate) is an integer multiple of the frame rate of the image frame (which may be referred to as a second preset frame rate). In one example, the mobile phone 100 collects image frames at 30FPS and collects depth data at 60FPS. In addition, the acquisition time of the image frames collected by the mobile phone 100 is synchronized with the acquisition time of the depth data. For example, as shown in Figure 4, the exposure time of the first row of each image frame is consistent with the exposure time of the depth data of the image frame, that is, each image frame is aligned in time with the depth data of the image frame. In this way, each image frame has valid depth data, which can provide data support for active focus.
[0104] In some implementations, to provide more accurate depth data, the laser device configured in the mobile phone 100 can provide depth data corresponding to 40×30 image regions. The laser device can be a laser matrix that can provide depth data for each region. For example, if an image frame includes 40×30 image regions, the laser device can provide depth data for one or more of the 40×30 image regions. In this way, the focus module of the mobile phone 100 can obtain accurate depth data of the target object.
[0105] In an embodiment of the present application, the mobile phone 100 can use a first detection frame to inform the user of the subject currently being tracked and focused. To highlight that the subject being tracked and focused is a moving object, in some implementations, the mobile phone 100 displays a first detection frame in the shooting interface that follows the moving object. This first detection frame includes multiple sub-frames, and the multiple sub-frames of the first detection frame cover the display area corresponding to the moving object.
[0106] For example, let's take the case where the moving object in the scene captured by the mobile phone 100 is a ball. As shown in (1) of FIG7 , in the shooting interface displayed by the mobile phone 100, the display area where the ball is located is covered with a first detection frame formed by multiple sub-frames. The first detection frame moves with the ball. After the ball moves to the position shown in (2) of FIG7 , the first detection frame moves with the ball and also moves to a position consistent with the ball.
[0107] In an embodiment of the present application, the mobile phone 100 prompts the user of moving objects in the shooting interface through a first detection frame formed by multiple sub-frames, which can improve the recognition of moving objects, thereby better prompting the user of moving objects in the shooting interface and improving the user's shooting experience.
[0108] In some implementations, the first detection frame is irregularly shaped, and may also be referred to as a special-shaped frame. The shape of the first detection frame corresponds to the shape of the moving object, and moving objects of different shapes correspond to different first detection frame shapes. Mobile phone 100 can adaptively adjust the shape of the first detection frame based on the shape of the moving object, allowing the first detection frame to present a variety of irregular shapes, providing a better visual experience for the user.
[0109] In some implementations, the number of sub-frames in the first detection frame is less than a preset number. For example, the preset number is 27, 20, or other values. If the number of sub-frames in the first detection frame is too large, it will affect the aesthetics of the shooting interface and the user's perception.
[0110] In one example, the preset number is equal to 1 / 16 of the total number of preset coordinate grids in the full-frame ratio. For example, the total number of preset coordinate grids in the full-frame ratio is 24×18, where 24 is the vertical number of the preset coordinate grids and 18 is the horizontal number of the preset coordinate grids. The frame ratio is used to indicate the aspect ratio of the preview window. The aspect ratio of the preview window is different under different frame ratios. For example, as shown in Figure 8, the frame ratios of the preview window in the shooting interface are 4:3, 1:1, and full screen. Among them, the full-frame ratio is consistent with the aspect ratio of the screen, and can also be 4:3.
[0111] The preset coordinate grid is used to evenly divide the display area of the mobile phone 100 screen. The preset coordinate grid includes multiple grids. The size of each grid can be 8×8 device independent pixels (dp). When the aspect ratio is 4:3, the preview window includes 24×18 grids. When the aspect ratio is 1:1, the preview window includes 18×18 grids. The mobile phone 100 can determine the grid corresponding to the moving object based on the display area where the moving object is located, and the grid corresponding to the moving object can be used as multiple sub-frames in the first detection frame.
[0112] In some implementations, the mobile phone 100 can generate multiple sub-frames of the first detection frame based on the depth data of a preset number (such as 40×30) of image areas provided by the laser device. For example, the mobile phone 100 can determine the depth data of the moving object based on the depth data corresponding to the second image frame. The depth data of the moving object is the depth data with the minimum distance in the depth data corresponding to the second image area. Further, the mobile phone 100 determines the multiple image areas where the moving object is located in the above-mentioned preset number of image areas based on the depth data of the moving object. Further, the mobile phone 100 determines the grid corresponding to the moving object based on the correspondence between the preset number of image areas provided by the laser device and the above-mentioned preset coordinate grid, and the grid corresponding to the moving object can be used as multiple sub-frames in the first detection frame.
[0113] In this way, the first detection frame provided by the mobile phone 100 can more accurately indicate the image position of the moving object.
[0114] In some implementations, if the display area of the moving object is large and the number of grids corresponding to the moving object exceeds a preset number, in this case, the mobile phone 100 selects the grids in the central area of the moving object as multiple sub-frames in the first detection frame, so that the number of multiple sub-frames is less than the preset number.
[0115] When the first detection frame covers the moving object, the first detection frame may visually block part of the moving object, causing inconvenience to the user. To reduce the visual impact of the first detection frame on the user, in some implementations, if the moving object includes a face, multiple sub-frames of the first detection frame cover the display area other than the face of the moving object. In this way, the first detection frame displayed in the shooting interface of the mobile phone 100 will avoid the face and display the first detection frame in the display area other than the face.
[0116] It is understandable that, in general, users are very likely to want to see the face clearly through the shooting interface. Therefore, in this implementation, the first detection frame displayed by the mobile phone 100 in the shooting interface will avoid the face, thereby providing the user with a clear and unobstructed face and improving the user experience.
[0117] For example, as shown in FIG9 , in motion focus mode, if the moving object detected by the mobile phone 100 is a person, a first detection frame is displayed on the shooting interface. The first detection frame covers the display area other than the person's face, avoiding the display area where the face is located.
[0118] In some implementations, the display position of the first detection frame at the current moment is predicted based on the image position of the moving object before the current moment. For example, the display position of the first detection frame in the second image frame (corresponding to the image frame captured at the current moment) is predicted based on the image position of the moving object in the first image frame (corresponding to the image frame captured before the current moment). Even when the moving object moves faster than the moving object, the first detection frame provided by the electronic device can accurately indicate the moving object, reduce the situation where the first detection frame is separated from the moving object, and improve the user experience.
[0119] To further improve the applicability of the frame-by-frame focus method, in some implementations, the camera is a wide-angle camera or a telephoto camera. The frame-by-frame focus method provided in the embodiments of the present application is applicable to zoom scenarios of wide-angle cameras or telephoto cameras, such as zoom tracking focus with a zoom ratio of 1-5x for wide-angle cameras or telephoto cameras.
[0120] In some implementations, if the mobile phone 100 receives a continuous shooting operation from the user, such as a continuous shooting operation in which the user long presses the shooting button, the mobile phone 100 responds to the continuous shooting operation and tracks the focus of the shooting subject in a frame-by-frame tracking manner.
[0121] In some implementations, if a subject (e.g., a face, body, or object) reenters the camera's field of view within a preset time (e.g., 2 seconds, 3 seconds, etc.) after leaving the camera's field of view, the phone 100 continues to track and focus on the subject. The subject may be automatically determined by the phone 100 or determined by the phone 100 in response to a user's click.
[0122] In other implementations, if the mobile phone 100 does not detect a moving object in motion focus mode, it can focus on a non-moving object. In motion focus mode, the priority order of the photographed subject is: object indicated by user operation (which may be called a manual trigger object) > moving object > registered human face > non-registered human face > human body > cat = dog.
[0123] When the mobile phone 100 focuses on a non-moving object, the mobile phone 100 displays a second detection frame in the display area of the non-moving object. The second detection frame is a rectangular frame, for example, a single-line frame, a double-line frame, etc.
[0124] The registered face is a face pre-registered in the mobile phone 100. For example, the registered face may be a face pre-registered in the mobile phone 100 with power-on permission. Another example is a face pre-registered by the target application. In other implementations, the registered face is determined by the mobile phone 100 based on the frequency of appearance of faces in stored images. For example, the mobile phone 100 may count the frequency of appearance of each face in images in the gallery application and select the j faces with the highest frequency of appearance as the registered faces. j is a positive integer.
[0125] In some implementations, in order to make the shooting interface concise and unified and reduce different elements in the shooting interface, there may be only one detection frame in the shooting interface at the same time, such as only displaying the first detection frame including multiple sub-frames.
[0126] As mentioned above, the camera function of the mobile phone 100 includes a subject focus mode in addition to the motion focus mode. In subject focus mode, the mobile phone 100 identifies a face in a captured image frame and determines the face's location. Furthermore, the face is focused based on its location. For example, the mobile phone 100 may use an autofocus algorithm to adjust the camera's focus so that the face's location in the image frame is clearly imaged.
[0127] It is understandable that, considering that the moving speed of the subject being focused is usually slow, in order to save power consumption of the mobile phone 100, the mobile phone 100 can focus on the subject every several image frames, for example, once every four image frames.
[0128] In subject focus mode, the priority order of the photographed subjects is: objects indicated by user operations (which can be called manual trigger objects) > registered faces > non-registered faces > human bodies > cats = dogs.
[0129] In the embodiment of the present application, the mobile phone 100 provides the user with multiple focusing options. Without receiving any user operation, the mobile phone 100 can automatically track the focus on a moving object or a face. The user can also lock the focus subject by manually clicking the screen, meeting the user's diverse shooting needs.
[0130] In some implementations, the shooting mode of the mobile phone 100 also includes a snapshot mode. In snapshot mode, the mobile phone 100 selects the image frame with the best image quality from multiple consecutively captured image frames as the final photo provided to the user. Snapshot mode can also be combined with focus mode. For example, when snapshot mode is enabled, the mobile phone 100 defaults to motion focus mode. When snapshot mode is not enabled, the mobile phone 100 defaults to subject focus mode.
[0131] Of course, the user can also set the combination of snapshot mode and focus mode according to his or her own preferences or shooting needs. For example, when the snapshot mode is turned on, select the subject focus mode. When the snapshot mode is not turned on, select the motion focus mode. The mobile phone 100 will automatically record the user's selection. After the mobile phone 100 turns off the shooting function under the control of the user, when the shooting function is turned on again, the focus mode and shooting mode of the camera function are the same as the settings before turning off the shooting function. For example, when the mobile phone 100 turns on the shooting function for the first time, the subject focus mode and the snapshot mode are started. When the shooting function is turned on for the second time, the mobile phone 100 starts the subject focus mode and the snapshot mode by default. In this way, the mobile phone 100 can maintain the last user's setting for the shooting function, and the user can shoot images in the mode he or she likes or is accustomed to without manual adjustment.
[0132] It should be noted that the personal information (such as facial images, etc.) used in the technical solution of this application is limited to information for which the individual's separate consent has been obtained, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) before using the function, and signing the agreement (authorization) including authorization of relevant user information.
[0133] In other embodiments of the present application, an electronic device is provided, comprising: a memory, a camera, and one or more processors. The memory and the camera are respectively coupled to the processor. The camera is used to capture image frames. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device can perform the various functions or steps in the above-mentioned method embodiments. Of course, the electronic device can also include other hardware structures. For example, the electronic device also includes hardware structures such as sensors and communication modules. The structure of the electronic device can refer to the structure of the mobile phone 100 shown in Figure 2.
[0134] An embodiment of the present application also provides a chip system, which is applied to an electronic device. The chip system includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected by lines. For example, the interface circuit can be used to receive signals from other devices (such as memories). For another example, the interface circuit can be used to send signals to other devices (such as processors). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which are not specifically limited in the embodiments of the present application.
[0135] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned electronic device, the electronic device executes each function or step in the above-mentioned method embodiment.
[0136] The present application also provides a computer program product, which, when executed on a computer, enables the computer to perform the functions or steps of the above method embodiment. For example, the computer may be the above electronic device.
[0137] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0138] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0139] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0142] The above content is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in the present application should be included in the protection scope of the present application.
Claims
1. A shooting method, characterized in that, Applied to an electronic device, the electronic device including a camera; the method includes: Receiving a first operation of a user, the first operation being used to indicate turning on the shooting function of a target application; In response to the first operation, displaying a first shooting interface of the target application, the first shooting interface including a first preview window, the first preview window being used to display a first image frame collected by the camera; In response to a second operation of the user, entering a motion focus mode; After detecting a target object, displaying a second shooting interface of the target application, the second shooting interface including a second preview window and a first detection frame, the second preview window being used to display a second image frame collected by the camera, the first detection frame being predicted based on a first image position of the target object in the first image frame, the first detection frame being used to indicate a second image position of the target object in the second image frame.
2. The method according to claim 1, characterized in that, The method further includes: Obtaining a focusing position of the camera based on the second image position of the target object in the second image frame; Adjusting the focus of the camera according to the focusing position of the camera.
3. The method according to claim 2, wherein The second image frame is the i-th frame image collected after the first image frame, where i is a positive integer.
4. The method according to any one of claims 1 to 3, characterized in that The obtaining the focusing position of the camera based on the second image position of the target object in the second image frame includes: Obtaining phase data and / or depth data of the second image frame; wherein, the acquisition time of the depth data of the second image frame is the same as the acquisition time of the second image frame; Obtaining the focusing position according to the phase data and / or depth data of the second image frame, and the second image position of the target object in the second image frame.
5. The method according to claim 4, wherein The electronic device includes a laser device, the laser device being used to provide depth data corresponding to a preset number of image regions, the preset number being 40×30.
6. The method according to claim 5, wherein The electronic device acquires depth data of image frames at a first preset frame rate through the laser device.
7. The method according to claim 6, characterized in that, The first preset frame rate is 60 frames per second.
8. The method according to any one of claims 1-7, characterized in that, When the target object is a moving object, the first detection frame includes a plurality of sub-frames, the plurality of sub-frames covering a display area corresponding to the moving object.
9. The method according to claim 8, characterized in that The method further includes: Determining, according to depth data of the moving object in the second image frame, a plurality of image regions corresponding to the moving object in a preset number of image regions corresponding to the second image frame, wherein the laser device of the electronic device is used to provide depth data of the preset number of image regions; Generating the plurality of sub-frames included in the first detection frame based on the plurality of image regions corresponding to the moving object.
10. The method according to claim 9, characterized in that, If the moving object includes a human face, the plurality of sub-frames of the first detection frame cover other display areas except the human face of the moving object.
11. The method according to claim 9 or 10, characterized in that, The number of sub-frames in the first detection frame is less than or equal to a preset number.
12. The method according to claim 10, wherein The preset number is equal to 1 / 16 of the total number of preset coordinate grids in the full-frame ratio. In the full-frame ratio, the electronic device displays the preview window full screen, and the preset coordinate network is used to evenly divide the display area of the preview window.
13. The method according to any one of claims 1-11, characterized in that, The camera is a wide-angle camera or a telephoto camera.
14. An electronic device, characterized in that, Comprising: a camera, a memory, and one or more processors; the camera and the memory are respectively coupled to the processor; the camera is used for acquiring image frames; wherein, computer program code is stored in the memory, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1-13.
15. A computer-readable storage medium, characterized in that, including computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-13.
Citation Information
Patent Citations
Focusing control method and electronic equipment
CN106961552A
Focusing processing method and equipment
CN108496350A
Focusing method, mobile terminal, and computer readable storage medium
CN109167910A
Target focus tracking method and device, electronic equipment and computer readable storage medium
CN110650291A
Focusing method and device applied to terminal equipment and terminal equipment
CN111050060A