Shooting method, electronic equipment and storage medium
By predicting the position of the subject in subsequent image frames and adjusting the camera focus, the problem of electronic devices detecting frame deviation when shooting moving objects is solved, achieving a clearer image display and a better user experience.
Patent Information
- Application Number
- CN202410042523.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-18
Smart Images

Figure CN120343395A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to a shooting method, an electronic device, and a storage medium. Background Art
[0002] With the development of electronic devices, the functions provided by electronic devices are becoming more and more diverse. The shooting function is one of the frequently used functions in electronic devices. Users can use the shooting function of the electronic device to take pictures and obtain images or videos. As people's requirements for the shooting effect of electronic devices continue to increase, the shooting function of electronic devices is becoming more and more abundant. The shooting function of electronic devices has a variety of shooting modes. For example, shooting modes such as portrait mode and night scene mode.
[0003] However, the shooting functions provided by current electronic devices are not yet perfect. In the case of shooting a shooting subject, the shooting subject in the captured image may be blurred, affecting the user's shooting experience. Summary of the Invention
[0004] This application provides a shooting method, an electronic device, and a storage medium, which can track and focus on the shooting subject during the shooting process, improve the clarity of the shooting subject in the captured image, and enhance the user's shooting experience.
[0005] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, a shooting method is provided. The method includes: receiving a first operation of a user, where the first operation is used to indicate turning on the shooting function of a target application. In response to the first operation, a first shooting interface of the target application is displayed. The first shooting interface includes a first preview window, and the first preview window is used to display a first image frame captured by a camera. In response to a second operation of the user, enter the motion focus mode. After detecting a target object, a second shooting interface of the target application is displayed. The second shooting interface includes a second preview window and a first detection frame. The second preview window is used to display a second image frame captured by the camera. The first detection frame is predicted based on a first image position of the target object in the first image frame, and the first detection frame is used to indicate a second image position of the target object in the second image frame.
[0007] In this way, the electronic device predicts the second image position of the subsequent image frame based on the detected first image position, and thus displays a detection frame in the shooting interface according to the predicted second image position. In this way, the detection frame can accurately indicate the image position of the shooting subject in the image frame, reducing the deviation between the detection frame and the shooting subject. Even when the shooting subject is a moving object, the detection frame of the electronic device can accurately follow the shooting subject for focusing, improving the shooting experience in a motion scene.
[0008] In a possible implementation of the first aspect, the focusing position of the camera is obtained based on the second image position of the target object in the second image frame, and the focus of the camera is adjusted according to the focusing position of the camera. In this way, the electronic device can pre-determine the focusing position to be used by the camera when collecting image frames based on the predicted second image position, and the electronic device can adjust the focus of the camera in advance so that the camera focuses at this focusing position. Thus, the influence of delays in aspects such as image frame detection and focusing position calculation on the tracking focus effect is reduced. Even when the shooting subject is a moving object, the camera of the electronic device can accurately follow the shooting subject to focus, improving the shooting experience in a moving scene.
[0009] In another possible implementation of the first aspect, the second image frame is the i-th frame image collected after the first image frame, where i is a positive integer.
[0010] In another possible implementation of the first aspect, the phase data and / or depth data of the second image frame are obtained. The acquisition time of the depth data of the second image frame is consistent with the acquisition time of the second image frame. The focusing position is obtained according to the phase data and / or depth data of the second image frame, and the second image position of the target object in the second image frame. In this way, the electronic device can combine the depth data and / or phase data corresponding to the second image position to predict the focusing position of the camera, thereby further improving the accuracy of the camera focusing, so that the camera can follow the target object to focus.
[0011] In another possible implementation of the first aspect, the electronic device includes a laser device, and the laser device is used to provide depth data corresponding to a preset number of (such as 40×30) image regions. In this way, the electronic device can obtain accurate depth data of the target object through the laser device.
[0012] In another possible implementation of the first aspect, the electronic device collects the depth data of the image frame at a first preset frame rate through the laser device. The first preset frame rate is 60 frames per second.
[0013] In another possible implementation of the first aspect, when the target object is a moving object, the first detection frame includes a plurality of sub-frames, and the plurality of sub-frames cover the display area corresponding to the moving object. The electronic device prompts the moving object in the shooting interface through the first detection frame formed by the plurality of sub-frames, which can improve the recognition rate of the moving object, thereby better prompting the user of the moving object in the shooting interface and improving the user's shooting experience.
[0014] In another possible implementation of the first aspect, if the moving object includes a human face, multiple sub-frames of the first detection frame cover other display areas of the moving object except the human face. In this implementation, the first detection frame displayed by the electronic device in the shooting interface will avoid the human face, thereby providing the user with a clear and unobstructed human face and improving the user experience.
[0015] In another possible implementation of the first aspect, the electronic device determines multiple image areas corresponding to the moving object in a preset number of image areas corresponding to the second image frame according to the depth data of the moving object in the second image frame. The laser device of the electronic device is used to provide the depth data of the preset number of image areas, such as providing the depth data of 40×30 image areas. The sizes of these image areas can be the same and are arranged evenly. Further, the electronic device generates multiple sub-frames included in the first detection frame based on the multiple image areas corresponding to the moving object. The depth data of the moving object is the depth data with the minimum distance in the depth data corresponding to the second image area. In this way, the first detection frame provided by the electronic device can more accurately indicate the image position of the moving object.
[0016] In another possible implementation of the first aspect, the number of multiple sub-frames in the first detection frame is less than or equal to a preset number. For example, the preset number is a value such as 27, 20, etc. If the number of multiple sub-frames in the first detection frame is too large, it will affect the aesthetics of the shooting interface and the user's perception.
[0017] In another possible implementation of the first aspect, the preset number is equal to 1 / 16 of the total number of preset coordinate grids in the full-frame ratio. In the full-frame ratio, the electronic device displays a preview window full-screen, and the preset coordinate network is used to evenly divide the display area of the preview window.
[0018] In another possible implementation of the first aspect, the camera is a wide-angle camera or a telephoto camera. This implementation is applicable to the zoom scenarios of wide-angle cameras or telephoto cameras, such as applicable to the zoom tracking focus with a zoom ratio of 1-5 times for wide-angle cameras or telephoto cameras.
[0019] In another possible implementation of the first aspect, the electronic device displays a function setting window in the shooting interface. The function setting window includes a function key for the motion focus mode. When the function key for the motion focus mode is in the selected state, the target application is in the motion focus mode.
[0020] In another possible implementation of the first aspect, the function setting window further includes a function key for the subject focus mode. The method further includes: in response to a second operation, switching the target application from the motion focus mode to the subject focus mode; in the subject focus mode, the electronic device preferentially focuses on the human face.
[0021] In another possible implementation of the first aspect, the target application further includes a snapshot mode. When the target application is in the snapshot mode, the target application uses the motion focus mode for focusing. When the target application is in a non-snapshot mode, the target application uses the subject focus mode for focusing. Of course, the user can also set the combination of the snapshot mode and the focus mode according to their own preferences or shooting needs. For example, when the snapshot mode is turned on, the subject focus mode is selected. When the snapshot mode is not turned on, the motion focus mode is selected.
[0022] In another possible implementation of the first aspect, in the motion focus mode, the priority order of the focus subjects is: the object indicated by the user operation (which can be called the manual trigger object) > the moving object > the registered face > the unregistered face > the human body > the cat = the dog. For example, if in the motion focus mode, the electronic device does not detect a moving object but detects a face, and the face is a registered face, the electronic device performs tracking focus on the registered face. A first detection frame is displayed in the display area of the tracking focus on the registered face, such as a yellow frame with double lines.
[0023] In another possible implementation of the first aspect, in the subject focus mode, the priority order of the focus subjects is: the object indicated by the user operation (which can be called the manual trigger object) > the registered face > the unregistered face > the human body > the cat = the dog.
[0024] In a second aspect, the present application provides an electronic device, including: a camera, a memory, and a processor; the memory and the display screen are respectively coupled to the processor. The camera is used to collect image frames. The memory stores computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device is caused to execute the method described in the first aspect and any of its possible implementations above.
[0025] In a third aspect, the present application provides a computer-readable storage medium, including computer instructions. When the computer instructions are run on an electronic device, the electronic device is caused to execute the method described in the first aspect and any of its possible implementations above.
[0026] In a fourth aspect, the present application provides a computer program product containing program instructions. When the computer program product runs on a computer, the computer can be caused to execute the method described in the first aspect and any of its possible implementations above. For example, the computer can be the above-mentioned electronic device.
[0027] In a fifth aspect, the present application provides a chip system, which is applied to an electronic device. The chip system includes an interface circuit and a processor. The interface circuit and the processor are interconnected through a line. The interface circuit is configured to receive a signal from a memory and send the signal to the processor, and the signal includes computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device executes the method described in the first aspect and any possible implementation manner thereof above. Description of the Drawings
[0028] Figure 1 FIG. is a schematic diagram of an image of a moving object captured by an electronic device provided in an embodiment of the present application;
[0029] Figure 2 FIG. is a hardware structure block diagram of an example mobile phone 100 of an electronic device provided in an embodiment of the present application;
[0030] Figure 3 FIG. is a software structure block diagram of an example mobile phone 100 of an electronic device provided in an embodiment of the present application;
[0031] Figure 4 FIG. is a timing schematic diagram of a focusing process provided in an embodiment of the present application;
[0032] Figure 5 FIG. is a flowchart of a shooting method provided in an embodiment of the present application;
[0033] Figure 6 FIG. is a schematic diagram of an interface for setting a focusing mode provided in an embodiment of the present application;
[0034] Figure 7 FIG. is a schematic diagram of a first detection frame provided in an embodiment of the present application;
[0035] Figure 8 FIG. is a schematic diagram of a preview window at different frame ratios provided in an embodiment of the present application;
[0036] Figure 9 FIG. is another schematic diagram of a first detection frame provided in an embodiment of the present application. Detailed Embodiments
[0037] To meet the shooting needs of users, an electronic device is configured with a shooting function. With the miniaturization and portability of electronic devices, users can take pictures at any time and place through an electronic device with a shooting function. As a result, users' requirements for the shooting function are also getting higher and higher. In the face of the increasing requirements of users for the shooting function, the electronic device provides various shooting modes such as portrait mode, night scene mode, and high-dynamic range (HDR) mode to meet the diverse shooting needs of users.
[0038] In the actual shooting process, there are many factors that affect the shooting effect. For example, focusing speed, shutter lag, image processing, etc. all affect the shooting effect. In this regard, the electronic device can improve the shooting effect of the finished film to a certain extent through some image processing methods. For example, the electronic device synthesizes multiple captured image frames through the HDR mode. During the synthesis process, noise in the multiple image frames can be filtered, and the details of the multiple image frames can be combined to obtain the final image presented to the user.
[0039] To further improve the shooting effect, the shooting function of the electronic device also provides a tracking focus mode. In the tracking focus mode, the electronic device tracks the shooting subject for focusing. The shooting subject is the main object captured by the electronic device, or is also called the target object. In order to make the shooting subject clearly imaged, the electronic device focuses on the shooting subject. The following introduces the tracking focus process of the electronic device through an embodiment.
[0040] After the electronic device captures an image frame, it determines the image area where the shooting subject is located in the image frame, and calculates the corresponding focus position of the shooting subject in the image frame according to the image area where the shooting subject is located. The image area where the shooting subject is located can be called the Region Of Interest (ROI). If the deviation between the focus position corresponding to the shooting subject and the currently used focus position of the electronic device (i.e., the focus of the camera) is greater than the preset threshold, it indicates that the difference between the currently used focus position of the electronic device and the focus position corresponding to the shooting subject is relatively large. If the electronic device continues at the current focus position, it is difficult for the electronic device to make the shooting subject clearly imaged. In this case, the electronic device adjusts the focus of the camera to move the focus of the camera to the focus position corresponding to the shooting subject to track the shooting subject for focusing. If the deviation between the focus position corresponding to the shooting subject and the currently used focus position of the electronic device is less than or equal to the preset threshold, it indicates that the difference between the currently used focus position of the electronic device and the focus position corresponding to the shooting subject is relatively small. If the electronic device continues at the current focus position, the electronic device can make the shooting subject clearly imaged. In this case, the electronic device does not adjust the focus of the camera and continues to keep the focus of the camera at the current focus position and use the current focus position to track the shooting subject for focusing. This tracking focus method can be called a passive tracking focus method.
[0041] However, during the shooting process of the electronic device, due to delays in aspects such as image frame detection and focusing, the ROI determined by the electronic device may deviate from the actual position of the shooting subject in the image frame. It is difficult for the detection frame displayed on the shooting interface of the electronic device to accurately indicate the image position of the shooting subject, and it is difficult for the camera of the electronic device to follow the shooting subject for focusing. The shooting subject in the image presented to the user may be blurred. Especially when the shooting subject is a moving object, it is difficult for the moving object in the image to be clearly imaged. For example, taking the shooting subject of the electronic device as a moving car, the electronic device follows the car for focusing. The image captured by the electronic device is as Figure 1 shown. It can be seen that the car in the image captured by the electronic device is blurred and not clear enough.
[0042] The embodiment of the present application also provides another focusing scheme to improve the accuracy of tracking focus and enhance the shooting effect. Specifically, the electronic device displays the image frame collected by the camera in the preview window of the shooting interface. When the electronic device collects the first image frame, it detects the first image frame to determine the image area where the shooting subject is located in the first image frame. The image area where the shooting subject is located is the ROI. The ROI detected in the first image frame and the ROI detected before the first image frame can provide the movement trend of the shooting subject. Further, the electronic device predicts the ROI corresponding to the second image frame according to the ROI detected in the first image frame. The second image frame is the image frame collected after the first image frame. The electronic device further displays a detection frame for indicating the shooting subject in the shooting interface according to the ROI corresponding to the second image frame.
[0043] In this way, the electronic device predicts the ROI of the subsequent image frames through the detected ROI, and thus displays the detection frame in the shooting interface according to the predicted ROI. In this way, the detection frame can accurately indicate the image position of the shooting subject in the image frame, reduce the deviation between the detection frame and the shooting subject. Even when the shooting subject is a moving object, the detection frame of the electronic device can accurately follow the shooting subject for focusing, improving the shooting experience in a moving scene.
[0044] Exemplarily, the electronic device described in the embodiments of the present application may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, a media player, a wearable device, and other devices. The embodiments of the present application do not impose special restrictions on the specific form of the electronic device.
[0045] In the embodiments of the present application, taking the electronic device as the mobile phone 100 as an example, the hardware structure of the electronic device is introduced through the mobile phone 100. As Figure 2 shown, the mobile phone 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0046] Among them, the processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), a driving processor, etc. Among them, different processing units may be independent devices or integrated in one or more processors. The processor 110 may be the nerve center and command center of the mobile phone 100. The processor 110 may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.
[0047] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may store instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0048] The external memory interface 120 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, the electronic device can save the captured images in the external memory card.
[0049] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. For example, in the embodiments of the present application, the processor 110 can execute the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. The electronic device can also save the captured images in the internal memory.
[0050] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. In some embodiments, the power management module 141 and the charging management module 140 may also be provided in the same device.
[0051] The sensor module 180 may include sensors such as a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a Hall sensor, a touch sensor, an ambient light sensor, and a bone conduction sensor. The mobile phone 100 can collect various data through the sensor module 180. For example, the mobile phone 100 collects touch data through the touch sensor in the sensor module 180, and thus the mobile phone 100 identifies user operations through the touch data.
[0052] The mobile phone 100 realizes the display function through a GPU, a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0053] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. For example, the mobile phone 100 displays the camera interface, the captured images, etc. through the display screen 194.
[0054] In some implementation manners, the above touch sensor may be disposed in the display screen 194, and the touch sensor and the display panel form a touch screen, also called a "touch control screen". The touch sensor is also called a "touch control panel", which is used to detect touch operations acting on or near it, such as click operations, swipe operations, etc. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. The mobile phone 100 can provide visual output related to the touch operation through the display screen 194.
[0055] The mobile phone 100 can realize the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.
[0056] The ISP is used to process the data fed back by the camera 193. For example, when the mobile phone 100 takes a photo, the shutter of the camera is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The camera photosensitive element converts the optical signal into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the human eye. The ISP can also perform algorithm optimization on the noise, brightness, and color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP may be disposed in the camera 193.
[0057] The camera 193 is used to capture static or dynamic images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the mobile phone 100 may include one or more cameras 193. For example, the mobile phone 100 includes a front camera and a rear camera.
[0058] The mobile phone 100 may further include a Time of Flight (TOF) sensor. The TOF sensor is used to measure the time taken for a light pulse to travel to and from the TOF sensor and the object being measured (such as the TOF data). The TOF data is used to generate depth data on the distance of the object being measured. In the embodiments of the present application, the TOF sensor may use a laser device to measure the distance between the shooting subject and the lens, so that the camera 193 can achieve autofocus. The time of flight sensor may be disposed inside the camera 193 or outside the camera 193.
[0059] It can be understood that the interface connection relationship between the modules illustrated in this embodiment is only for illustrative purposes and does not constitute a structural limitation on the electronic device. In some other embodiments, the electronic device may also include more or fewer modules than those provided in the above embodiments, and different interface connection methods as described in the above embodiments, or a combination of multiple interface connection methods, may be adopted between the various modules. The methods in the following embodiments can all be implemented in an electronic device having the above hardware structure.
[0060] The software system of the electronic device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, taking the Android system with a layered architecture and the electronic device being the mobile phone 100 as an example, the software structure of the electronic device is illustrated by way of example.
[0061] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system may include an application layer, an application framework layer, an Android runtime, system libraries, a hardware abstraction layer (HAL), and a kernel layer.
[0062] The application layer may include a series of application packages. For example, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, etc., and the embodiments of the present application do not make any restrictions on this.
[0063] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, and a sensor manager, etc., and the embodiments of the present application do not make any restrictions on this.
[0064] The Android runtime includes a core library and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core library contains two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0065] The system library may include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc.
[0066] The hardware abstraction layer is an encapsulation of the Linux kernel driver, provides an interface upward, and shields the implementation details of the underlying hardware. For example, the hardware abstraction layer may include a camera hardware abstraction layer, a Wi-Fi hardware abstraction layer, a Bluetooth hardware abstraction layer, etc.
[0067] The kernel layer is the layer between the hardware and the software. The kernel layer includes driver programs and system service programs. The driver programs at least include a display driver, a camera driver, an audio driver, a sensor driver, etc.
[0068] Figure 3 is the structural block diagram of the mobile phone 100 in the embodiments of the present application. Below, taking the application processor of the mobile phone 100 having Figure 3 the hierarchical architecture shown, the working processes of the software and hardware of the mobile phone 100 are exemplarily described.
[0069] In the example, the user operates the mobile phone 100 to open the camera application. In response to the user operation of opening the camera application, the camera application issues an instruction to turn on the camera through the camera access interface in the application framework layer. The instruction to turn on the camera issued by the camera application is sequentially transmitted to the camera 193 through the camera access interface, the camera hardware abstraction layer, and the camera driver. In response to the instruction to turn on the camera, the camera 193 of the mobile phone 100 starts to collect image frames and transmits the image frames to the image signal processor. The image signal processor performs image processing on the image frames transmitted by the camera 193, such as performing front-end image processing (Image Front End, IFE) such as color adjustment and noise reduction, and generating preview image frames through the Image Processing Engine (IPE). The image signal processor further transmits the processed image frames to the camera hardware abstraction layer through the camera driver. The camera hardware abstraction layer includes camera-related algorithm modules, such as a detection module, a focusing module, etc. Among them, the detection module is used to detect the image frames and determine the ROI of the image frames. The focusing module is used to determine the focusing position of the camera. Further, the camera hardware abstraction layer reports the position of the ROI to the camera application through the camera access interface. The camera application can execute subsequent shooting processes based on the position of the ROI reported by the camera hardware abstraction layer, such as displaying a detection frame in the shooting interface, etc.
[0070] In the following embodiments, the electronic device is taken as an example of the mobile phone 100 to introduce the method provided by the embodiments of the present application.
[0071] As described above, during the shooting process of the electronic device, due to delays in aspects such as image frame detection and focusing, the ROI determined by the electronic device may deviate from the actual position of the shooting subject in the image frame, and it is difficult for the camera of the electronic device to follow the shooting subject to focus. To facilitate understanding of the solution provided by the embodiments of the present application, the reasons for the deviation between the ROI determined by the electronic device and the actual position of the shooting subject in the image frame will be first introduced below with reference to the accompanying drawings.
[0072] Taking the electronic device as an example of the mobile phone 100, as Figure 4 shown, the mobile phone 100 periodically collects image frames through the camera, for example, collects image frames at a frame rate of 30 frames per second (Frames Per Second, FPS). Figure 4 The horizontal axis of represents time. The camera collects image frames in a progressive exposure manner, that is, the camera starts to expose from the first row of the image frame. After the exposure of the first row is completed, the camera reads the data of the first row. After a line exposure period, the second row starts to be exposed. After the exposure of the second row is completed, the camera reads the data of the second row, and so on, until the entire image frame is read. Therefore, the image frames collected by the camera are in Figure 4A parallelogram is presented. The position of the ROI in the image frame changes continuously over time. The camera transmits the read image frame to the image signal processor, and the mobile phone 100 processes it through the image signal processor. It takes a certain amount of time from the start of camera exposure to reading an image frame, such as 0 - 16 milliseconds. It also takes a certain amount of time for the image signal processor to process the image frame, such as 19 milliseconds. It takes a certain amount of time for the image signal processor to transmit the processed image frame to the camera hardware abstraction layer, such as 1 millisecond. The camera hardware abstraction layer uses the detection algorithm provided by the detection module to detect the image frame and determine the ROI of the image frame. It takes a certain amount of time for the detection algorithm in the camera hardware abstraction layer to determine the ROI of the image frame, such as 11 milliseconds. After the detection module obtains the ROI of the image frame, it then transmits the ROI of the image frame to the focusing module. It takes a certain amount of time for the detection module to transmit the ROI of the image frame to the focusing module, such as 25 milliseconds. It can be seen that from the start of the camera collecting an image frame (such as Image Frame 1) to the time when the focusing module obtains the ROI of this image frame, the time delay is approximately 54 milliseconds to 70 milliseconds. At this time, the camera of the mobile phone 100 has already started collecting or is about to collect the 3rd image frame (i.e., Image Frame 3).
[0073] If at this time the focusing module of the mobile phone 100 calculates the focusing position using the ROI detected by the detection module, that is, focuses according to the ROI detected in the 1st image frame, since the position of the ROI changes continuously over time and the position of the ROI in Image Frame 3 has deviated from the position of the ROI in Image Frame 1, then when the camera collects Image Frame 3 and uses the focusing position corresponding to the ROI of Image Frame 1, it is difficult for the camera to capture a clear subject.
[0074] In view of this, in this embodiment, the mobile phone 100 predicts the ROI of the image frames collected after Image Frame 1 based on the ROI of Image Frame 1 and the ROIs detected in the image frames before Image Frame 1, such as predicting the ROIs of n (such as 4) image frames collected after Image Frame 1. Herein, n is a positive integer. Further, the mobile phone 100 determines the focusing position of the camera according to the predicted ROI so that the camera tracks the subject for focusing. This focusing method can be called active focusing. Active focusing can reduce the influence of the above time delay on the focusing position of the camera, and even if the subject is a fast - moving object, it has a relatively accurate focusing effect.
[0075] In view of the fact that the above-mentioned active autofocus method also has a relatively accurate autofocus effect when the shooting subject is a moving object, in some implementation manners, the mobile phone 100 can set a focusing mode for a scene where the shooting subject is a moving object. For example, the shooting function of the mobile phone 100 can have a motion focusing mode and a subject focusing mode. The motion focusing mode is a mode that preferentially performs autofocus on a moving object. In the motion focusing mode, the mobile phone 100 uses the active autofocus method to perform autofocus on the shooting subject. The subject focusing mode is a mode that preferentially performs autofocus on a human face. In the subject focusing mode, the position change of the shooting subject is relatively small, and the mobile phone 100 uses the above-mentioned passive autofocus method to perform autofocus on the shooting subject. Of course, regardless of whether the shooting subject is a moving object, the mobile phone 100 can use the above-mentioned active autofocus method for autofocus. The embodiments of the present application do not limit this.
[0076] Taking the mobile phone 100 using the active autofocus method in the motion focusing mode as an example, the method provided by the embodiments of the present application will be described by way of example. As Figure 5 shown, the method provided by the embodiments of the present application includes the following steps:
[0077] S501, in response to a first operation, the mobile phone 100 starts the shooting function of the target application, acquires a first image frame through the camera, and displays the first image frame in the shooting interface of the target application.
[0078] The mobile phone 100 is installed with a target application for providing a shooting function. For example, the target application is an application program with a shooting function such as a camera application or a chat application. In response to the first operation, the mobile phone 100 starts the shooting function of the target application, starts to acquire image frames through the camera, and displays the shooting interface of the target application.
[0079] It can be understood that the mobile phone 100 periodically acquires image frames through the camera, such as acquiring image frames at 30 FPS through the camera. The first image frame is an image frame acquired by the mobile phone 100 through the camera at a first moment.
[0080] The first operation is used to indicate to start the target application. For example, the first operation is a user operation of clicking on the icon of the target application. Of course, the first operation can also be other user operations that indicate to start the shooting function of the target application, such as a user operation of pressing a physical button on the mobile phone 100 that has a function of starting the shooting function. The embodiments of the present application do not limit the specific implementation manner of the first operation.
[0081] The shooting interface of the target application includes a preview window. The preview window is used to display the image frames captured by the camera of the mobile phone 100. Through the preview window, the user can see the scene picture currently captured by the mobile phone 100, and the user can change the shooting angle of the mobile phone 100 through the picture displayed in the preview window so that the mobile phone 100 captures the shooting subject. Among them, the shooting interface of the electronic device includes a first shooting interface, and the first shooting interface includes a first preview window, and the first preview window is used to display the first image frame captured by the camera.
[0082] In addition to including a preview window, the shooting interface also includes various function keys or controls. For example, at the bottom and top of the shooting interface, function keys for various shooting modes such as aperture, night scene, portrait, photo taking, video recording, movie, and professional are also set. In different shooting modes, the images captured by the mobile phone 100 are different in terms of color, clarity, or imaging method, etc. The user can select any shooting mode to capture images.
[0083] S502, in the motion focus mode, the mobile phone 100 determines the target object in the first image frame and the first image position of the target object in the first image frame.
[0084] After the mobile phone 100 starts the shooting function, in response to the second operation of the user, it enters the motion focus mode. The shooting interface of the target application is provided with a function key for the motion focus mode and / or a function key for the subject focus mode. The user can select the focus mode according to the actual shooting requirements and shooting scene. For example, the mobile phone 100 displays a function bar in the shooting interface of the target application, and the function bar includes a function key for the motion focus mode. The mobile phone 100 receives the second operation of the user clicking the function key for the motion focus mode in the function bar and enters the motion focus mode.
[0085] Exemplarily, as Figure 6 shown, a pull-up function key is set at the bottom (or called the shallow layer) of the shooting interface. If the mobile phone 100 receives the user operation of the user clicking or swiping up the pull-up function key, the mobile phone 100 pops up a function bar (or called a function setting window) in the shooting interface. The function bar is provided with function keys for the focus mode. If the mobile phone 100 receives the click operation of the function key for the focus mode by the user, the mobile phone 100 displays a function key for the motion focus mode (such as "motion focus priority") and a function key for the subject focus mode (such as "subject focus priority") in the function bar. If the mobile phone 100 receives the user operation of selecting the motion focus mode (an example of the second operation), the mobile phone 100 sets the focus mode of the shooting function to the motion focus mode. Of course, if the mobile phone 100 receives the user operation of selecting the subject focus mode, the mobile phone 100 sets the focus mode of the shooting function to the subject focus mode.
[0086] In an embodiment of the present application, after the camera of the mobile phone 100 captures each image frame, the captured image frame is transmitted to the image signal processor. Taking the first image frame as an example, the image signal processor performs image processing on the first image frame and transmits the processed first image frame to the detection module. The detection module of the mobile phone 100 uses a detection algorithm to detect the first image frame, such as using detection algorithms such as object detection algorithms, single-object tracking algorithms, optical flow methods, frame difference methods, etc., analyzes the changes in pixel points in adjacent image frames, and determines the target object (i.e., the shooting subject) in each first image frame and the first image position of the target object in the first image frame (i.e., the ROI of the first image frame).
[0087] S503, the mobile phone 100 predicts the second image position of the target object in the second image frame according to the first image position of the target object in the first image frame.
[0088] Exemplarily, the detection module of the mobile phone 100 determines the movement trajectory of the target object according to the first image position of the target object in the first image frame and the third image positions in one or more third image frames (such as 8 third image frames) captured before the first moment. This movement trajectory is used to indicate the movement trend (or motion trend) of the moving object. For example, this movement trajectory includes the movement direction and the movement speed. Further, the detection module of the mobile phone 100 predicts the second image position of the moving object in the second image frame captured after the first image frame according to the movement trajectory of the target object, that is, predicts the image position of the target object after the first moment. The second image frame is the i-th frame image captured after the first image frame, and i is a positive integer. For example, the second image frame is the 1st frame image, the 2nd frame image, or the 3rd frame image captured after the first image frame.
[0089] In another example, the detection module of the mobile phone 100 can use a preset machine learning model to predict the second image position of the target object in the second image frame. Specifically, the detection module of the mobile phone 100 can input the first image frame marked with the image position of the target object and multiple third image frames into the trained preset machine learning model to obtain the prediction result output by the preset machine learning model. The first image frame and the multiple third image frames input into the preset machine learning model are arranged in chronological order. The prediction result output by the preset machine learning model represents the image position of the target object after the first moment, such as outputting the second image position of the target object in the second image frame captured after the first moment, or outputting the image positions of the target object in multiple image frames (such as 4 image frames) captured after the first moment. The image positions in these multiple image frames include the second image position in the second image frame.
[0090] The preset machine learning model can be obtained by the mobile phone 100 training a machine learning model with a specific model structure using the collected historical image frames and the positions of the reference objects in the historical image frames. Alternatively, the preset machine learning model is trained by other electronic devices. The mobile phone 100 can directly install the preset machine learning model that has been trained. Through the preset machine learning model, the mobile phone 100 can relatively accurately predict the image position of the target object at a future moment. In the embodiments of the present application, the model structure of the preset machine learning model is not limited. For example, the preset machine learning model can adopt model structures such as support vector machines and convolutional neural networks.
[0091] S504, the mobile phone 100 determines the focusing position of the camera based on the second image position of the target object in the second image frame, and displays a first detection frame on the shooting interface of the target application.
[0092] In the embodiments of the present application, after the detection module of the mobile phone 100 predicts the second image position of the target object in the second image frame, it transmits the second image position of the target object in the second image frame to the focusing module. The focusing module of the mobile phone 100 determines the focusing position of the camera according to the second image position of the target object in the second image frame. For example, the focusing module of the mobile phone 100 can use an autofocus (AF) algorithm to determine the focusing position of the camera, and further adjust the focus of the camera according to the focusing position of the camera. For example, the motor of the camera pushes the lens focus of the camera to the focusing position, so that the position where the target object is located can be clearly imaged.
[0093] In addition, the mobile phone 100 also generates a first detection frame based on the second image position of the target object in the second image frame, and displays the first detection frame on the shooting interface of the target application. For example, the mobile phone 100 can generate a first detection frame according to the depth data corresponding to the second image position and the second image position, and display the first detection frame on the shooting interface of the target application. The first detection frame is used to indicate the second image position of the target object in the second image frame. The first detection frame is predicted based on the first image position of the target object in the first image frame. Thus, even due to delays in aspects such as image frame detection, the first detection frame can accurately indicate the image position of the target object on the shooting interface. At this time, the shooting interface of the electronic device can be a second shooting interface, and the second shooting interface includes a second preview window and a first detection frame. The second preview window is used to display the second image frame collected by the camera. The first detection frame can be a focusing frame.
[0094] In order to further improve the accuracy of the mobile phone 100 in focusing on the target object, in some implementation manners, the mobile phone 100 may further obtain the phase data and / or depth data of the second image frame, and further determine the focusing position of the camera when collecting the next image frame of the second image frame according to the phase data and / or depth data of the second image frame, and the second image position of the target object in the second image frame.
[0095] In an example of this implementation manner, the mobile phone 100 is configured with a laser device. The laser device is used to measure the time taken for a laser pulse to travel back and forth between the laser device and the object to be measured, and obtain TOF data. The TOF data can be converted into depth data of the distance between the object to be measured and the laser device. The focusing module of the mobile phone 100 can obtain the depth data corresponding to the second image position of the second image frame, that is, obtain the depth data of the target object at the acquisition moment of the second image frame (such as called the second moment). Further, the focusing module determines the focusing position of the camera when collecting the next image frame of the second image frame (such as called the third moment) according to the depth data corresponding to the second image position of the second image frame. For example, the focusing module of the mobile phone 100 can predict the depth data of the target object at the third moment according to the depth data of the target object at the second moment and the depth data of the target object before the second acquisition moment. Further, the focusing module of the mobile phone 100 determines the focusing position of the camera when collecting the next image frame of the second image frame according to the depth data of the target object at the third moment.
[0096] In another example of this implementation manner, the focusing module of the mobile phone 100 can obtain the phase (phase detection, PD) data (or called phase difference data) of the second image frame. The mobile phone 100 further obtains the phase data corresponding to the second image position according to the second image position of the second image frame. Further, the focusing module determines the focusing position of the camera when collecting the next image frame of the second image frame according to the phase data corresponding to the second image position of the second image frame. For example, the focusing module of the mobile phone 100 can determine the focusing position corresponding to the second image frame according to the phase data corresponding to the second image position of the second image frame. Further, the focusing module of the mobile phone 100 determines the change trend of the focusing position according to the focusing position corresponding to the second image frame and the focusing position corresponding to the first image frame. The mobile phone 100 then predicts the focusing position of the camera when collecting the next image frame of the second image frame according to the change trend of the focusing position.
[0097] In another example of this implementation manner, the focusing module of the mobile phone 100 can obtain the phase data and depth data corresponding to the second image position of the second image frame according to the second image position of the second image frame. The focusing module of the mobile phone 100 further predicts the focusing position when the camera captures the next image frame of the second image frame according to the phase data and depth data corresponding to the second image position of the second image frame. For example, the focusing module of the mobile phone 100 predicts the focusing position corresponding to the next image frame of the second image frame (such as referred to as the first focusing position) according to the phase data corresponding to the second image position of the second image frame. The focusing module of the mobile phone 100 predicts the focusing position corresponding to the next image frame of the second image frame (such as referred to as the second focusing position) according to the depth data corresponding to the second image position of the second image frame. Further, the focusing module of the mobile phone 100 determines the final focusing position used when the camera captures the next image frame of the second image frame according to the first focusing position and the second focusing position, such as performing weighted averaging on the first focusing position and the second focusing position.
[0098] In the embodiments of the present application, the mobile phone 100 can obtain the depth data and / or phase data corresponding to the second image position through the predicted second image position, and then predict the focusing position adopted when the camera captures the next image frame of the second image frame according to the depth data and / or phase data corresponding to the second image position. In this way, the accuracy of the camera focusing can be further improved, so that the camera can focus on the target object.
[0099] It can be understood that since there is a certain time delay in the focusing module obtaining the second image position transmitted by the detection module, the camera of the mobile phone 100 has captured the image frame after the first image frame when the focusing module obtains the second image position. At this time, if the focusing module of the mobile phone 100 uses the first image position detected in the first image frame for focusing, for a target object with a relatively fast moving speed, especially for a target object moving rapidly in the moving direction of the camera lens, it is difficult for the camera of the mobile phone 100 to focus on the target object, and the target object in the captured second image is blurred. Therefore, in order to cope with the rapid movement of the target object, the mobile phone 100 uses the predicted second image position to adjust the focus in advance to achieve accurate focusing on the target object.
[0100] Such as Figure 4As shown, take the first image frame as image frame 1 and the second image frame as image frame 3 as an example. When the focusing module obtains the second image position of image frame 3 (predicted) transmitted by the detection module, the mobile phone 100 has already captured image frame 2 and has not yet captured image frame 3. After the mobile phone 100 exposes and reads the second image position of image frame 3, since it takes a certain amount of time for the detection module to detect image frame 3, the mobile phone 100 will not immediately obtain the detection result of image frame 3 (i.e., the image position of the target object detected in image frame 3) after reading image frame 3. The focusing module starts to predict the focusing position of the camera according to the second image position of image frame 3 (predicted), depth data, and / or phase data. The acquisition time of the depth data corresponding to image frame 3 is the same as the acquisition time of image frame 3, that is, the time when the first row of image data of image frame 3 is read after exposure is the same as the time when the depth data of image frame 3 is read after exposure. The depth data collected by the laser device is collected in a global exposure mode, that is, all the depth data corresponding to image frame 3 can be collected at one time. It takes a certain amount of time to convert the TOF data measured by the laser device into depth data, such as 2 - 3 milliseconds. It also takes a certain amount of time for the focusing module to obtain the depth data corresponding to image frame 3, such as 2 - 4 milliseconds. It takes a certain amount of time for the focusing module to predict the focusing position of the camera according to the depth data and / or phase data corresponding to the second image position of image frame 3, such as 5 milliseconds. It takes a certain amount of time for the focusing module to transmit the predicted focusing position to the camera, such as 0.3 milliseconds. The motor of the camera needs to consume a certain amount of time to push the lens of the camera according to the focusing position, such as 10 milliseconds. At this time, the mobile phone 100 has not started to capture image frame 4, but the focus of the camera has already moved to the focusing position corresponding to image frame 4. In this way, the mobile phone 100 can pre-adjust the focus of the camera to push the lens of the camera to the focusing position corresponding to image frame 4 in advance. When the mobile phone 100 captures image frame 4, the lens of the camera has already been pushed to the focusing position corresponding to image frame 4 in advance, so as to achieve focusing on the target object.
[0101] It can be understood that considering the relatively fast moving speed of the target object, in order to improve the focusing accuracy, in the motion focusing mode, the mobile phone 100 can use the above active focusing method to predict the focusing position corresponding to each image frame, and the camera of the mobile phone 100 controls the motor to push the lens according to the focusing position corresponding to each image frame. This focusing method can also be called frame-by-frame focusing.
[0102] In some implementations, to provide effective depth data, the frame rate of the depth data collected by the mobile phone 100 (which can be referred to as the first preset frame rate) is an integer multiple of the frame rate of the image frames (which can be referred to as the second preset frame rate). In one example, the mobile phone 100 captures image frames at 30 FPS and depth data at 60 FPS. In addition, the acquisition time of the image frames captured by the mobile phone 100 is synchronized with the acquisition time of the depth data. For example, as Figure 4 shown, the exposure time of the first row of each image frame is the same as the exposure time of the depth data of that image frame, that is, each image frame is temporally aligned with the depth data of that image frame. In this way, each image frame has effective depth data, which can provide data support for active focusing.
[0103] In some implementations, to provide more accurate depth data, the laser device configured in the mobile phone 100 can provide depth data corresponding to 40×30 image regions. The laser device can be a laser matrix, which can provide depth data for sub-regions. For example, for an image frame including 40×30 image regions, the laser device can provide depth data for one or more of the 40×30 image regions. In this way, the focusing module of the mobile phone 100 can obtain accurate depth data of the target object.
[0104] In the embodiments of the present application, the mobile phone 100 can prompt the user of the current tracking and focusing shooting subject through a first detection frame. To highlight that the tracking and focusing shooting subject is a moving object, in some implementations, the mobile phone 100 displays a first detection frame that follows the moving object in the shooting interface. The first detection frame includes a plurality of sub-frames, and the plurality of sub-frames of the first detection frame cover the display area corresponding to the moving object.
[0105] Exemplarily, taking the moving object in the shooting scene of the mobile phone 100 being a ball as an example. As Figure 7 shown in (1) of, in the shooting interface displayed by the mobile phone 100, the display area where the ball is located is covered with a first detection frame formed by a plurality of sub-frames. The first detection frame follows the ball. After the ball moves to the position shown in (2) of Figure 7 , the first detection frame moves with the movement of the ball and also moves to the same position as the ball.
[0106] In the embodiments of the present application, the mobile phone 100 prompts the moving object in the shooting interface through the first detection frame formed by a plurality of sub-frames, which can improve the recognition rate of the moving object, so as to better prompt the moving object in the shooting interface and improve the user's shooting experience.
[0107] In some implementations, the above-mentioned first detection frame has an irregular shape and can also be referred to as a special-shaped frame. The shape of the first detection frame corresponds to the shape of the moving object, and the shapes of the first detection frames corresponding to moving objects of different shapes are different. The mobile phone 100 can adaptively adjust the shape of the first detection frame according to the shape of the moving object, so that the first detection frame can present various irregular shapes, providing a better visual experience for users.
[0108] In some implementations, the number of multiple sub-frames in the first detection frame is less than a preset number. For example, the preset number is a value such as 27 or 20. If the number of multiple sub-frames in the first detection frame is too large, it will affect the beauty of the shooting interface and the user's perception.
[0109] In one example, the preset number is equal to 1 / 16 of the total number of preset coordinate grids in the full-frame ratio. For example, the total number of preset coordinate grids in the full-frame ratio is 24×18, where 24 is the longitudinal number of the preset coordinate grids and 18 is the transverse grid number of the preset coordinate grids. The frame ratio is used to indicate the aspect ratio of the preview window. The aspect ratios of the preview windows under different frame ratios are different. For example, as Figure 8 shown, the frame ratios of the preview windows in the shooting interface are 4:3, 1:1, and full screen respectively. Among them, the full-frame ratio is the same as the aspect ratio of the screen and can also be 4:3.
[0110] The preset coordinate grids are used to evenly divide the display area of the screen of the mobile phone 100. The preset coordinate grids include multiple grids. The size of each grid can be 8×8 device independent pixels (dp). When the frame ratio is 4:3, the preview window includes 24×18 grids. When the frame ratio is 1:1, the preview window includes 18×18 grids. The mobile phone 100 can determine the grid corresponding to the moving object according to the display area where the moving object is located, and the grid corresponding to the moving object can be used as multiple sub-frames in the first detection frame.
[0111] In some implementations, the mobile phone 100 can generate multiple sub-frames of the first detection frame according to the depth data of a preset number (such as 40×30) of image areas provided by the laser device. For example, the mobile phone 100 can determine the depth data of the moving object according to the depth data corresponding to the second image frame. The depth data of the moving object is the depth data with the smallest distance in the depth data corresponding to the second image area. Further, the mobile phone 100 determines the multiple image areas where the moving object is located in the above-mentioned preset number of image areas according to the depth data of the moving object. Further, the mobile phone 100 determines the grid corresponding to the moving object according to the correspondence between the preset number of image areas provided by the laser device and the above-mentioned preset coordinate grids, and the grid corresponding to the moving object can be used as multiple sub-frames in the first detection frame.
[0112] In this way, the first detection frame provided by the mobile phone 100 can more accurately indicate the image position of the moving object.
[0113] In some implementation manners, if the display area of the moving object is large and the number of grids corresponding to the moving object exceeds the preset number, in this case, the mobile phone 100 selects the grids in the central area of the moving object as multiple sub-frames in the first detection frame, so that the number of multiple sub-frames is less than the preset number.
[0114] When the first detection frame covers the moving object, the first detection frame may visually block a part of the area of the moving object, causing inconvenience to the user. To reduce the visual impact of the first detection frame on the user, in some implementation manners, if the moving object includes a human face, the multiple sub-frames of the first detection frame cover other display areas outside the human face of the moving object. In this way, the first detection frame displayed by the mobile phone 100 in the shooting interface will avoid the human face and display the first detection frame in other display areas outside the human face.
[0115] It can be understood that generally, the user is very likely to want to see the human face clearly through the shooting interface. Therefore, in this implementation manner, the first detection frame displayed by the mobile phone 100 in the shooting interface will avoid the human face, so as to provide the user with a clear and unobstructed human face and improve the user experience.
[0116] Exemplarily, as Figure 9 shown, in the motion focus mode, if the moving object detected by the mobile phone 100 is a person, the first detection frame is displayed in the shooting interface, and the first detection frame covers other display areas outside the human face of the person, avoiding the display area where the human face is located.
[0117] In some implementation manners, the display position of the first detection frame at the current moment is predicted based on the image position of the moving object before the current moment. For example, the display position of the first detection frame in the second image frame (corresponding to the image frame collected at the current moment) is predicted based on the image position of the moving object in the first image frame (corresponding to the image frame collected before the current moment). Even when the moving object is moving rapidly, the first detection frame provided by the electronic device can accurately indicate the moving object, reduce the situation where the first detection frame is separated from the moving object, and improve the user experience.
[0118] To further improve the applicability of the frame-by-frame focusing method, in some implementation manners, the above camera is a wide-angle camera or a telephoto camera. The frame-by-frame focusing method provided by the embodiments of the present application is applicable to the zoom scenarios of wide-angle cameras or telephoto cameras, such as zoom tracking focusing with a zoom ratio of 1-5 times for wide-angle cameras or telephoto cameras.
[0119] In some implementations, if the mobile phone 100 receives a continuous shooting operation from the user, such as a continuous shooting operation where the user long-presses the shooting button, the mobile phone 100 responds to this continuous shooting operation and tracks the focus of the shooting subject using the frame-by-frame tracking focus method.
[0120] In some implementations, if the shooting subject (such as a face, a human body, an object, etc.) enters the field of view of the camera again within a preset duration (such as 2 seconds, 3 seconds, etc.) after going out of the field of view of the camera, the mobile phone 100 continues to perform tracking focus on this shooting subject. This shooting subject can be a shooting subject automatically determined by the mobile phone 100, or a shooting subject determined by the mobile phone 100 in response to the user's click operation.
[0121] In some other implementations, if the mobile phone 100 does not detect a moving object in the motion focus mode, it can focus on a non-moving object. In the motion focus mode, the priority order of the shooting subjects is: the object indicated by the user operation (which can be called the manually triggered object) > moving object > registered face > non-registered face > human body > cat = dog.
[0122] When the mobile phone 100 focuses on a non-moving object, the mobile phone 100 displays a second detection frame in the display area of the non-moving object. The second detection frame is a rectangular frame. For example, the second detection frame can be a single-line frame, a double-line frame, etc.
[0123] Among them, the registered face is a face pre-registered in the mobile phone 100. For example, the registered face can be a face with the power-on permission pre-registered in the mobile phone 100. For another example, the registered face can be a face pre-registered by the target application. In some other implementations, the registered face is determined by the mobile phone 100 according to the appearance frequency of the face in the stored images. For example, the mobile phone 100 can count the appearance frequency of each face in the images in the gallery application, and thus take the first j faces with the appearance frequency as the registered faces. j is a positive integer.
[0124] In some implementations, in order to make the shooting interface simple and unified and reduce different elements in the shooting interface, at the same time, there can be only one type of detection frame in the shooting interface, such as only displaying the first detection frame including multiple sub-frames.
[0125] As described above, in addition to the motion focus mode, the shooting function of the mobile phone 100 also includes the subject focus mode. In the subject focus mode, the mobile phone 100 identifies a face in the captured image frame and determines the position of the face. Further, it focuses on the face according to the position of the face. For example, the mobile phone 100 can use an autofocus algorithm to adjust the focus of the camera to make the position where the face is located in the image frame clearly imaged.
[0126] It can be understood that considering that the moving speed of the focus subject is usually small, in order to save the power consumption of the mobile phone 100, the mobile phone 100 can adopt a method of focusing on the shooting subject every several image frames. For example, focus is performed once every 4 image frames are collected.
[0127] In the subject focus mode, the priority order of the shooting subject is: the object indicated by the user operation (which can be called the manual trigger object) > the registered face > the unregistered face > the human body > the cat = the dog.
[0128] In the embodiment of the present application, the mobile phone 100 provides the user with a variety of focusing solutions. In the case where no user operation is received, the mobile phone 100 can actively perform automatic tracking focus on a moving object or a face. The user can also lock the tracking focus subject by manually clicking on the screen to meet the diverse shooting needs of the user.
[0129] In some implementation manners, the shooting mode of the mobile phone 100 further includes a snapshot mode. In the snapshot mode, the mobile phone 100 selects the image frame with the best image quality among multiple continuously captured image frames as the photo finally provided to the user. The snapshot mode can also be combined with the focus mode. For example, when the snapshot mode is turned on, the mobile phone 100 defaults to turn on the motion focus mode. When the snapshot mode is not turned on, the mobile phone 100 defaults to turn on the subject focus mode.
[0130] Of course, the user can also set the combination of the snapshot mode and the focus mode according to their own preferences or shooting needs. For example, when the snapshot mode is turned on, select the subject focus mode. When the snapshot mode is not turned on, select the motion focus mode. The mobile phone 100 will automatically record the user's selection. After the mobile phone 100 closes the shooting function under the user's control and then turns on the shooting function again, the focus mode and shooting mode of the camera function are the same as the settings before the shooting function was closed. For example, when the mobile phone 100 turns on the shooting function for the first time, it starts the subject focus mode and the snapshot mode. When the mobile phone 100 turns on the shooting function for the second time, the mobile phone 100 defaults to start the subject focus mode and the snapshot mode. In this way, the mobile phone 100 can maintain the previous user settings for the shooting function, and the user can shoot images in the mode they like or are used to without manual adjustment.
[0131] It should be noted that the personal information (such as face images, etc.) used in the technical solution of the present application is limited to the information for which the individual's separate consent has been obtained, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) before the user uses this function, and signing the agreement (authorization) including authorizing the relevant user information.
[0132] In some other embodiments of the present application, an electronic device is further provided, including: a memory, a camera, and one or more processors. The memory and the camera are respectively coupled to the processor. The camera is used to collect image frames. Computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device can execute each function or step in the above method embodiments. Of course, the electronic device may further include other hardware structures. For example, the electronic device further includes hardware structures such as sensors and communication modules. The structure of the electronic device may refer to Figure 2 the structure of the mobile phone 100 shown.
[0133] An embodiment of the present application further provides a chip system, which is applied to an electronic device. The chip system includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected through a line. For example, the interface circuit can be used to receive signals from other devices (such as a memory). For another example, the interface circuit can be used to send signals to other devices (such as a processor). Exemplarily, the interface circuit can read the instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can execute each step in the above embodiments. Of course, the chip system may further include other discrete devices, and the embodiments of the present application do not make specific limitations thereto.
[0134] An embodiment of the present application further provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above electronic device, the electronic device is enabled to execute each function or step in the above method embodiments.
[0135] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute each function or step in the above method embodiments. For example, the computer may be the above electronic device.
[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0137] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0138] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0140] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0141] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application.
Claims
1. A shooting method, characterized in that, Applied to an electronic device, the electronic device including a camera; the method includes: Receiving a first operation of a user, the first operation being used to indicate turning on the shooting function of a target application; In response to the first operation, displaying a first shooting interface of the target application, the first shooting interface including a first preview window, the first preview window being used to display a first image frame collected by the camera; In response to a second operation of the user, entering a motion focus mode; After detecting a target object, displaying a second shooting interface of the target application, the second shooting interface including a second preview window and a first detection frame, the second preview window being used to display a second image frame collected by the camera, the first detection frame being predicted based on a first image position of the target object in the first image frame, the first detection frame being used to indicate a second image position of the target object in the second image frame.
2. The method according to claim 1, wherein The method further includes: Obtaining a focusing position of the camera based on the second image position of the target object in the second image frame; Adjusting a focus of the camera according to the focusing position of the camera.
3. The method according to claim 2, wherein The second image frame is the i-th frame image collected after the first image frame, where i is a positive integer.
4. The method according to any one of claims 1 to 3, characterized in that The obtaining the focusing position of the camera based on the second image position of the target object in the second image frame includes: Obtaining phase data and / or depth data of the second image frame; wherein, a collection time of the depth data of the second image frame is consistent with a collection time of the second image frame; Obtaining the focusing position according to the phase data and / or depth data of the second image frame, and the second image position of the target object in the second image frame.
5. The method according to claim 4, wherein The electronic device includes a laser device, the laser device being used to provide depth data corresponding to a preset number of image regions, the preset number being 40×30.
6. The method according to claim 5, wherein The electronic device collects depth data of an image frame at a first preset frame rate through the laser device.
7. The method according to claim 6, wherein The first preset frame rate is 60 frames per second.
8. The method according to any one of claims 1-7, characterized in that, When the target object is a moving object, the first detection frame includes a plurality of sub-frames, the plurality of sub-frames covering a display area corresponding to the moving object.
9. The method according to claim 8, wherein The method further includes: Determining, according to depth data of the moving object in the second image frame, a plurality of image regions corresponding to the moving object in a preset number of image regions corresponding to the second image frame, wherein the laser device of the electronic device is used to provide depth data of the preset number of image regions; Generating the plurality of sub-frames included in the first detection frame based on the plurality of image regions corresponding to the moving object.
10. The method according to claim 9, wherein If the moving object includes a human face, the plurality of sub-frames of the first detection frame cover other display areas except the human face of the moving object.
11. The method according to claim 9 or 10, characterized in that, The number of the plurality of sub-frames in the first detection frame is less than or equal to a preset number.
12. The method according to claim 10, characterized in that, The preset number is equal to 1 / 16 of a total number of a preset coordinate grid under a full-frame ratio, and under the full-frame ratio, the electronic device displays the preview window in full screen, and the preset coordinate grid is used to evenly divide a display area of the preview window.
13. The method according to any one of claims 1-11, characterized in that, The camera is a wide-angle camera or a telephoto camera.
14. An electronic device, characterized in that, It includes: a camera, a memory, and one or more processors; the camera and the memory are respectively coupled to the processor; the camera is used to capture image frames; wherein, computer program code is stored in the memory, the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1-13.
15. A computer-readable storage medium, characterized in that, It includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-13.
Citation Information
Patent Citations
Focusing control method and electronic equipment
CN106961552A
Interaction relationship recognition method and apparatus, and device and storage medium
WO2021164662A1
Human behavior detection method and apparatus, electronic device and storage medium
WO2022228252A1
Cited By
Motion decoupling control method and system based on motion mirror big data
CN121888097A