Image processing method and electronic equipment
By using image registration and EIS algorithm model processing, the target offset is determined, which solves the problem of large differences between captured images and preview frames in non-zero-second delay camera mode, and improves the shooting efficiency and image output consistency of electronic devices.
Patent Information
- Application Number
- CN202410445067.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-10-28
AI Technical Summary
In shooting scenarios where non-zero-second delay camera mode is enabled and the zoom level is high, there is a significant difference between the displayed content of the captured image and the preview frame, causing users to repeatedly retake the shot and affecting the human-computer interaction efficiency of electronic devices.
By employing image registration technology in electronic devices, the target offset is determined to ensure consistency between the digitally zoomed image and the preview frame. The EIS algorithm model is used to process the images in the preview stream, and the position of the second image region is determined by combining the pose information. The first field of view is then cropped and the captured image is generated.
It improves the efficiency of human-computer interaction during shooting, reduces unnecessary processing steps, and ensures the quality and consistency of the output images.
Smart Images

Figure CN120856979A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to an image processing method and an electronic device. Background Technology
[0002] With the development of electronic technology, photography functionality has become an essential feature of electronic devices. Users' demands and experiences regarding photography (taking photos and / or videos) on electronic devices are constantly increasing. Electronic devices can capture images using digital zoom. Digital zoom refers to cropping and / or enlarging images through software algorithms.
[0003] In shooting scenarios where non-zero shutter lag (ZSL) camera mode is enabled and the zoom ratio is high, after digital zoom processing, there is a problem that the content displayed in the captured image is significantly different from the content displayed in the last preview frame before taking the picture. This causes users to repeatedly retake the picture, affecting the human-computer interaction efficiency of electronic devices. Summary of the Invention
[0004] This application provides an image processing method and electronic device to solve the problem of human-computer interaction efficiency when taking pictures in non-ZSL camera mode.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, embodiments of this application provide an image processing method applied to an electronic device.
[0007] When the aforementioned electronic device runs an application with shooting function in the foreground, it can take pictures in either ZSL camera mode or non-ZSL camera mode according to actual business needs.
[0008] In certain shooting scenarios, such as shooting night scenes, or when the preview stream and the photo stream output formats are different, electronic devices can only enable non-ZSL camera modes.
[0009] When non-ZSL camera mode is enabled and no action indicating shooting is detected, the electronic device can display a preview stream captured in real time.
[0010] The preview stream consists of multiple continuously captured preview frames, including, for example, a first preview frame. In scenarios where the electronic device has a high zoom ratio (greater than the base zoom), the first preview frame can be an image obtained by digitally zooming the first image; that is, the content displayed in the first preview frame is the same as the first image area in the first image captured by the electronic device. Furthermore, the first image area is located at the first position in the first image.
[0011] Understandably, the first position is used to describe the position information of the first image region in the image coordinate system corresponding to the first image. It can be the image coordinates corresponding to the center point of the first image region, the image coordinates corresponding to the vertices of the first image region, and / or the image coordinates corresponding to the edge points of the first image region, etc.
[0012] In an exemplary scenario, while the electronic device displays the first preview frame in the preview stream, a user-instructed shooting action is detected. In this scenario, the first preview frame is the last preview frame displayed before the shooting action. Then, in response to the user-instructed shooting action, the electronic device captures a second image and stores the first captured image. The first captured image is obtained by digitally zooming the second image; that is, the content displayed in the first captured image is the same as the second image region in the second image. The second image region is located at a second position in the second image. This second position can be the image coordinates corresponding to the center point of the second image region, the image coordinates corresponding to the vertices of the second image region, and / or the image coordinates corresponding to the edge points of the second image region, etc. Furthermore, the first image region and the second image region are of the same size.
[0013] In this embodiment, the offset of the second position relative to the first position is the target offset, which is the offset between image regions presenting the same target object in the first image and the second image.
[0014] The target object mentioned above can be the subject appearing in both the first and second images. It is understood that the positional deviation of the same subject in different image data can characterize the positional transformation relationship between the first and second images. During digital zooming of the second image, by ensuring that the determined second image region and the first image region in the first image also satisfy the above positional transformation relationship, the displayed content in the second image region and the first image region is ensured to be similar. This solves the problem of large differences between the captured image and the last preview frame before taking the picture, ensuring image consistency and improving the efficiency of human-computer interaction during shooting.
[0015] In some embodiments, after the electronic device acquires the second image, it performs image registration based on the first and second images to determine a third image region in the first image and a fourth image region in the second image. The third and fourth image regions are used to display the same target object. The electronic device determines a target offset based on the positions of the third and fourth image regions. The electronic device determines a second position based on the target offset and a first position.
[0016] In the above embodiments, image registration is used to determine the target offset between the first image and the second image. Thus, during digital zoom of the second image, the actual cropping position can be determined using the target offset and the first position, rather than simply using the first position for cropping. This allows the resulting captured image to more closely resemble the content displayed in the first preview frame, improving the efficiency of human-computer interaction during shooting.
[0017] In some embodiments, the target object includes any of the following:
[0018] (1) The subject captured in the first preview frame.
[0019] For example, if the subject being tracked by the camera of the electronic device is in motion, and / or the first preview frame includes only one subject, the target object is the subject included in the first preview frame. It is understood that the subject being tracked by the camera of the electronic device can be the subject appearing in multiple consecutive preview frames in the preview stream. The first preview frame is included in these multiple consecutive preview frames.
[0020] (2) The image area occupied contains the subject at the center point of the first preview frame. For example, if the subject tracked by the lens of the electronic device is in motion, and / or the first preview frame includes only multiple subjects, the target subject is the subject whose image area contains the center point of the first preview frame.
[0021] (3) The subject that occupies the largest image area in the first preview frame. For example, if the subject tracked by the lens of the electronic device is in motion, and / or the first preview frame includes only multiple subjects, the target subject is the subject that occupies the largest image area in the first preview frame.
[0022] (4) A subject whose position remains unchanged. For example, if the subject tracked by the lens of the electronic device is static, the target object is a subject whose position remains unchanged. Such as buildings, facilities, road signs, scenery, mountains, forests, etc.
[0023] (5) The subject appears in both the first and second images and is in a static state. If the subject tracked by the lens of the electronic device is static, the target object is the subject that appears in both the first and second images and is in a static state. For example, a static subject can be a subject whose relative position remains unchanged in multiple consecutive preview frames. Here, relative position refers to the positional relationship with a subject whose position remains fixed.
[0024] In the above embodiments, the electronic device can select a wider variety of target objects. In addition, different target objects can be selected according to different scenarios, so that the selected target objects are more suitable for calculating the target offset and improving the accuracy of the determined target offset.
[0025] In some embodiments, when the electronic device acquires the second image, the electronic device obtains first posture information; the electronic device inputs the first posture information and the second image into the electronic image stabilization (EIS) algorithm model, and combines the first position and the second posture information corresponding to the first image to obtain the second position, wherein the second posture information may be the posture information obtained when acquiring the first image.
[0026] The attitude information mentioned above may include gyroscope information collected from a gyroscope sensor and / or acceleration information collected from an accelerometer sensor.
[0027] In the above embodiment, the second image used to create the photographed image is equivalent to an image in the preview stream that was captured later than the first image. The position of the second image region in the second image is determined by using the EIS algorithm model to process the preview stream.
[0028] Understandably, the EIS algorithm model is an algorithm used to process streaming data. By equating the second image with an image in the preview stream that was acquired later than the first image, the EIS algorithm model can combine relevant information from the first image (second pose information and first position) to determine the second position corresponding to the second image region in the second image, so that the determined second position can be offset from the first position as the target offset value.
[0029] In some embodiments, the electronic device crops a first field of view from a second image based on a second position.
[0030] For example, when the second position includes the edge coordinates of the second image region, the second image region is cropped in the second image according to the edge coordinates in the second position to obtain the first field of view.
[0031] As another example, when the second position includes the coordinates of a specific point in the second image region, the second position also includes the size information of the second image region. The electronic device can locate the second image region in the second image and crop the second image region based on the specific point coordinates and size information in the second position to obtain the first field of view.
[0032] In some embodiments, after cropping the first field of view, the electronic device generates a first photographed image based on the first field of view. For example, the electronic device can enlarge the first field of view to obtain a first photographed image with the same image size as the second image. This enlargement of the first field of view can be achieved by upsampling the first field of view.
[0033] In some embodiments, when ZSL camera mode is enabled, the electronic device acquires a third image and caches the third image. The electronic device displays a second preview frame corresponding to the third image, the content of which is the same as the fifth image area in the third image; in response to a user-instructed shooting operation, the electronic device generates and stores a second captured image based on the cached third image; wherein the content of the second captured image is the same as the fifth image area in the third image.
[0034] In the above embodiments, the electronic device can be compatible with both ZSL camera mode and non-ZSL camera mode for taking pictures, thus improving the adaptability of the solution.
[0035] In some embodiments, before storing the first captured image corresponding to the second image, it is determined that the current zoom ratio is greater than or equal to a preset zoom ratio.
[0036] In some embodiments, after storing the first captured image corresponding to the second image, the electronic device responds to a user operation and configures the zoom ratio to a value less than the preset zoom ratio.
[0037] The electronic device then continues to display the preview stream. While the electronic device displays the third preview frame in the preview stream, the user-instructed shooting action is detected again. The content displayed in the third preview frame is the same as the sixth image area in the fourth image, which is located at the third position in the fourth image.
[0038] In response to the user's instruction to take a picture, the electronic device captures a fifth image and stores the corresponding third image. The content displayed in the third image is the same as the seventh image area in the fifth image, which is located in the third position within the fifth image.
[0039] As can be seen, in scenarios where the zoom ratio is greater than the preset ratio, such as when taking a first photographed image, the second position of the first photographed image in the second image is offset from the first position of the previous preview frame (e.g., the first preview frame) in the first image. This offset ensures that the displayed content of the first photographed image and the first preview frame are similar, thus ensuring the quality of the photographed image.
[0040] In scenarios where the zoom ratio is less than the preset ratio, that is, in scenarios where a second image is taken, the position of the second image on the fifth image is the same as the position of the previous preview frame (e.g., the third preview frame) on the fourth image. Since the zoom ratio is relatively low, the quality of the image is not affected even though no offset processing is performed.
[0041] In the above embodiments, scene-specific offset processing is implemented, which reduces unnecessary processing steps while ensuring shooting quality.
[0042] In some embodiments, before the electronic device acquires the second image, the electronic device also needs to configure target camera parameters for acquiring the captured image into the camera sensor, wherein the target camera parameters are different from the camera parameters for acquiring the preview stream.
[0043] In a second aspect, embodiments of this application provide an electronic device, including a processor and a memory, wherein the memory is used to store code instructions; and the processor is used to execute the code instructions, causing the electronic device to perform the methods described in the first aspect and any of its implementations.
[0044] Thirdly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed, cause a computer to perform the method described in the first aspect and any of its implementations.
[0045] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when run, causes a computer to perform the methods described in the first aspect and any of its implementations.
[0046] It should be understood that the second to fourth aspects of the embodiments of this application correspond to the technical solutions of the first aspect of the embodiments of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be described again. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0048] Figure 2 A schematic diagram of the hardware and software structure of an electronic device provided in an embodiment of this application;
[0049] Figure 3 Example diagram showing the generation and display of the preview stream in ZSL camera mode;
[0050] Figure 4 This is an example diagram showing the generation and display of the photo stream in ZSL camera mode;
[0051] Figure 5 These are example images of previewing and taking photos in non-ZSL camera mode;
[0052] Figure 6 One of the example images of a zoom scene provided in the embodiments of this application:
[0053] Figure 7Example diagram of a zoom scene provided in the embodiments of this application:
[0054] Figure 8 These are example images showing, in some embodiments, how digital zoom causes significant differences between the captured image and the preview image in non-ZSL camera mode;
[0055] Figure 9 This is one of the flowcharts of an image processing method provided in an embodiment of this application;
[0056] Figure 10 Example diagrams for determining offset information between image data a and image data b provided in embodiments of this application;
[0057] Figure 11 The utilization provided in the embodiments of this application Figure 10 Example diagram showing how offset information in the image is used to determine the position of the field-of-view cropping box for cropping the photographed image;
[0058] Figure 12 This is a second flowchart of an image processing method provided in an embodiment of this application. Detailed Implementation
[0059] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0060] The implementation of this embodiment will now be described in detail with reference to the accompanying drawings.
[0061] This application provides an image processing method applied to an electronic device with a shooting function, used to resolve the difference between a preview image and a captured image. The preview image refers to the image captured and displayed by the electronic device before it detects a user-instructed shooting operation. The captured image refers to the image captured by the electronic device in response to the user-instructed shooting operation.
[0062] For example, electronic devices can be desktops, laptops, tablets, handheld computers, mobile phones, laptops, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), televisions, VR devices, AR devices, and other devices with cameras.
[0063] like Figure 1As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0064] The aforementioned sensor module 180 may include sensors such as pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, and bone conduction sensors.
[0065] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include... Figure 1 It can show more or fewer parts, or combine some parts, or split some parts, or arrange different parts. Figure 1 The components shown can be implemented in hardware, software, or a combination of both.
[0066] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0067] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0068] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0069] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0070] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0071] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0072] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a touch layer and a display panel. The touch layer is used to sense user interaction with display screen 194. The display panel can be a liquid crystal display (LCD), organic light-emitting diode (OLED), active-matrix organic light-emitting diode (AMOLED), flexible light-emitting diode (FLED), minimized, microLED, micro-OLED, quantum dot light-emitting diode (QLED), etc. Electronic device 100 can realize shooting functions through ISP, camera 193, video codec, GPU, display screen 194, and application processor.
[0073] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element (image sensor). The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0074] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor, such as a camera sensor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include N cameras 193, where N is a positive integer greater than 1.
[0075] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0076] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0077] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0078] Figure 2 This is a schematic diagram of the architecture (including software system and some hardware) used in the embodiments of this application. Figure 2 As shown, the application architecture is divided into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the application architecture can be divided into five layers, from top to bottom: the application layer, the application framework layer, the hardware abstraction layer (HAL), the driver layer, and the hardware layer.
[0079] like Figure 2 As shown, the application layer includes camera applications and gallery applications. This is understandable. Figure 2 The examples shown are only a portion of the applications; in fact, the application layer can include other applications as well, and this application does not limit this. For example, the application layer may also include applications such as messaging, alarm clock, weather, stopwatch, compass, timer, flashlight, calendar, and Alipay.
[0080] like Figure 2 As shown, the application framework layer includes a camera access interface. The camera access interface includes camera management and camera devices. The hardware abstraction layer includes a camera hardware abstraction layer and a camera algorithm library. The camera hardware abstraction layer includes multiple camera nodes. The camera algorithm library includes post-processing algorithm modules, decision-making modules, image registration algorithm modules, offset processing modules, and electronic image stabilization (EIS) algorithm modules.
[0081] It is understood that the image registration algorithm module, offset processing module, and electronic anti-shake (EIS) algorithm module described above are the algorithm modules required to implement the method provided in the embodiments of this application. Their operating principles and logic will be explained in detail in subsequent embodiments, and will not be repeated here. In addition, the decision module, image registration algorithm module, offset processing module, and EIS algorithm module can also be located at the application layer or application framework layer, and the embodiments of this application do not specifically limit them in this regard.
[0082] The driver layer is used to drive hardware resources. The driver layer can include multiple driver modules. For example... Figure 2 As shown, the driver layer includes camera device drivers, digital signal processor drivers, and graphics processor drivers, etc.
[0083] The hardware layer includes multiple camera sensors, an ISP, a digital signal processor, a graphics processor, and a gyroscope sensor. Figure 2 (Not shown). Understandably, the hardware layer may also include other hardware modules, without specific limitations.
[0084] In this system, each camera node in the camera hardware abstraction layer corresponds to a camera sensor in the hardware layer. For example, the camera application in the application layer can access the camera node through the camera access interface. This camera node is used to manage the corresponding camera sensor, such as setting its configuration parameters and controlling image acquisition.
[0085] In an exemplary scenario, a user can tap the camera application, allowing the electronic device to run the camera application in the foreground. The camera application sends control commands to the camera hardware abstraction layer (HIB) via the camera access interface. These commands may carry information such as the enabled camera mode (e.g., photo mode, video mode, or portrait mode) and zoom level. In response to these control commands, the HIB invokes the camera algorithm library. The decision module within the camera algorithm library determines the target camera sensor (e.g., the target camera sensor) and its configuration parameters (including sensor output method, parameter configuration for the ISP module, and parameter configuration for the post-processing algorithm module) based on the zoom level, camera mode, and ambient light. The decision module can then pass the target camera sensor's identifier and configuration parameters to the HIB.
[0086] The camera hardware abstraction layer can create a target camera node corresponding to the target camera sensor. Through the target camera node, configuration parameters are sent to the camera device driver. The camera device driver then sends these configuration parameters to the hardware layer; for example, it sends the sensor image output method to the target camera sensor and the parameter configurations for each ISP module to the ISP. Subsequently, the target camera sensor can output images based on the sensor image output method. The ISP can then perform corresponding processing based on the parameter configurations for the ISP modules. It is understood that the configuration parameters passed to the target camera sensor can be called camera parameters, and in addition to the sensor image output method exemplified in the previous embodiments, they can also include automatic exposure (AE) strategies, output image size, etc., without specific limitations here.
[0087] The camera algorithm library is also used to send digital signals to the digital signal processor (DSP) driver in the driver layer, so that the DSP driver can call the DSP in the hardware layer to perform digital signal processing. The DSP can then return the processed digital signal to the camera algorithm library through its driver.
[0088] The camera algorithm library is also used to send digital signals to the graphics signal processor (GSProcessor) driver in the driver layer, so that the GSProcessor driver can call the graphics processor in the hardware layer to perform digital signal processing. The graphics processor can then return the processed image data to the camera algorithm library through the graphics processor driver.
[0089] Additionally, when the camera application is running in the foreground, it can send preview or capture commands to the camera hardware abstraction layer via the camera access interface. The capture command is generated by the camera application in response to a user's instruction to take a picture. The preview command is generated periodically by the camera application when it displays the shooting preview interface and no user-instructed shooting is detected.
[0090] After receiving a preview or capture command, the camera hardware abstraction layer (HAL) can transmit the command to the target camera sensor via the camera device driver. The target camera sensor, responding to the command, acquires raw RAW image data and transmits it to the ISP (Image Signal Processor), which converts it into an image visible to the user. The ISP output can then be sent to the camera device driver, which in turn can send it back to the HAL. The HAL can then either send the image to a post-processing algorithm module for further processing or to the camera access interface. The camera access interface can then send the image returned by the HAL to the camera application. The camera application can then render the image and trigger its display.
[0091] To ensure clarity and brevity in the description of the following embodiments, a brief introduction to the relevant technologies is given first.
[0092] (1) Zero shutter lag (ZSL) is a camera mode that reduces the delay time of the camera shutter, making shooting more immediate and responsive, thereby reducing the difference between the preview image and the captured image.
[0093] In an exemplary scenario, when an electronic device is running a camera application in the foreground and ZSL camera mode is enabled, a shooting preview interface provided by the camera application can be displayed. This shooting preview interface is used to display a preview stream from the target camera sensor, which consists of multiple continuously acquired preview frames (or preview images).
[0094] Figure 3 This demonstrates the process of generating and sending a preview stream in ZSL camera mode. An example is provided where the preview stream includes preview frames 1, 2, 3, 4, 5, 6, and 7. Figure 3 As shown, the target camera sensor of the electronic device acquires raw image data (e.g., raw images) of preview frames 1, 2, 3, 4, 5, 6, and 7. Each time the target camera sensor acquires a frame of raw image data (e.g., the raw image data of preview frame 1), it can transmit the raw image data (e.g., the raw image data of preview frame 1) to the ISP. After processing by the ISP, a preview frame (e.g., preview frame 1) that can be displayed is obtained.
[0095] Understandably, the electronic device can transmit the raw image data of preview frames 1, 2, 3, 4, 5, 6, and 7 to the ISP frame by frame, in the order of acquisition. After processing by the ISP, the corresponding preview frames are obtained. In this way, the electronic device can display preview frames 1, 2, 3, 4, 5, 6, and 7 one by one on the shooting preview interface.
[0096] The aforementioned ISP can include an ISP first module, an ISP second module, and an ISP third module. Different modules can perform different processing on image data. For example, the ISP first module may include one or more of the following processing: binning, HDR fusion, etc. The ISP second module may include one or more of the following processing: bad pixel correction (BPC), black level correction (BLC), lens shade correction (LSC), automatic white balance (AWB), Bayer domain noise reduction (NR), Demosaic, etc. The ISP third module may include one or more of the following processing: color correction (CC), YUV domain noise reduction (NR), color enhancer (CE), sharpening, tone mapping, etc.
[0097] For example, after the raw image data is input into the ISP, the first module of the ISP performs remosaic processing on the first image data. The second module of the ISP performs Bayer domain processing and outputs data in RGB format. The third module of the ISP performs either RGB domain or YUV domain processing and outputs data in YUV format.
[0098] The above descriptions of the ISP first module, ISP second module, and ISP third module are merely illustrative and the embodiments of this application are not limited thereto.
[0099] like Figure 3 As shown, the raw image data of the preview frames acquired by the target camera sensor can also be cached in the cache queue. For example, in a scenario where the target camera sensor has acquired raw image data of preview frames 1, 2, 3, 4, 5, 6, and 7, preview frames 1, 2, 3, 4, 5, 6, and 7 can be cached in the cache queue.
[0100] like Figure 4As shown, when the electronic device displays preview frame 5, it detects a user-instructed action to take a picture. The electronic device can retrieve the raw image data of preview frame 5 from the buffer queue. Then, it processes the raw image data of preview frame 5 using an image processing algorithm module (e.g., a post-processing algorithm module) in the ISP and / or camera algorithm library to obtain the captured image corresponding to preview frame 5, and stores it.
[0101] The captured image and the preview frame can reuse the same original image data. This means that the difference between the preview frame displayed when the electronic device detects the user's instruction to take a picture and the captured image obtained by the electronic device in response to the user's instruction is minimal. In other words, this achieves the effect of reducing camera shutter lag.
[0102] ZSL camera mode is a commonly used camera mode, and electronic devices typically enable it by default. However, in some scenarios, electronic devices cannot use ZSL camera mode, meaning a non-ZSL camera mode needs to be used.
[0103] For example, in scenarios where the AE strategy used for previewing and taking photos differs, a non-ZSL camera mode needs to be enabled.
[0104] Taking the scenario of an electronic device shooting a night scene as an example, to ensure the shooting effect, the automatic exposure time when capturing the image is relatively long, for example, 100ms. In addition, the electronic device captures preview frames at a frequency of 30fps, and correspondingly, the automatic exposure time for capturing preview frames does not exceed 33ms. Clearly, the AE strategy required for the target camera sensor of the electronic device differs between the preview and the captured image. In this scenario, the preview frame and the captured image cannot reuse the same original image data.
[0105] For example, in scenarios where the preview output format and the captured image output format are different, the electronic device also needs to enable the non-ZSL camera mode. Understandably, when the preview output format and the captured image output format are different, the camera parameters set when the target camera sensor acquires the raw image data of the captured image and the preview frame are different, and the preview frame and the captured image cannot reuse the same raw image data.
[0106] As another example, the target camera sensor is a specific type of camera sensor, and the electronic device also needs to enable non-ZSL camera mode. The camera parameters required to be configured when acquiring preview frames with the specific type of camera sensor are different from the camera parameters required to acquire captured images.
[0107] In some embodiments, the electronic device runs a camera application in the foreground and enables a non-ZSL camera mode. The target camera sensor acquires raw image data (e.g., raw images) of preview frames 1, 2, 3, 4, 5, 6, and 7. The target camera sensor transmits the acquired raw image data to the ISP (Internet Service Provider). After processing by the ISP, displayable preview frames are obtained and transmitted to the display screen of the electronic device for display. Unlike the ZSL camera mode, in the non-ZSL camera mode, the electronic device does not cache the raw image data of the preview frames.
[0108] like Figure 5 As shown, when the electronic device is running a camera application in the foreground, camera parameter 1 can be set in the target camera sensor. Camera parameter 1 is the camera parameter required to capture preview frames. Then, the electronic device displays preview frames from the target camera component frame by frame. It can be understood that after setting camera parameter 1 on the target camera sensor, the target camera sensor can acquire raw image data based on camera parameter 1. The raw image data acquired by the target camera sensor is processed by image processing modules such as the ISP to obtain displayable preview images, such as preview frame 1, preview frame 2, preview frame 3, preview frame 4, preview frame 5, preview frame 6, and preview frame 7.
[0109] like Figure 5 As shown, when the electronic device displays preview frame 7, a user-instructed shooting operation is detected. In response to the user-instructed shooting operation, camera parameters 2 (also referred to as target camera parameters) are set in the target camera sensor. These camera parameters 2 are the camera parameters required to acquire the captured image. Afterwards, the electronic device stores the acquired image 8. It can be understood that after setting camera parameters 2 in the target camera sensor, the target camera sensor can acquire raw image data based on these parameters. The raw image data acquired by the target camera sensor is then processed by the image processing algorithm module in the ISP and / or camera algorithm library to obtain the captured image 8.
[0110] like Figure 5 As shown, after the electronic device stores the captured image 8, it can re-set the camera parameters 1 in the target camera sensor. Afterward, the electronic device can continue to display preview frames from the target camera component, such as preview frame 9, preview frame 10, etc.
[0111] Additionally, at time point a, the electronic device detects the user's instruction to take a picture, and at time point b, it acquires and generates the captured image 8. Between time points a and b: the electronic device stops acquiring preview frames and changes the camera parameters of the target camera sensor, for example, updating them from camera parameter 1 to camera parameter 2. Then, based on the raw image data acquired by the target camera sensor according to camera parameter 2, the corresponding captured image 8 is generated, for example, by processing the raw image data using an ISP and camera algorithm library to obtain captured image 8, which is then stored. The duration between time points a and b is also known as the shutter delay.
[0112] The aforementioned preview frame 7 is the last preview frame displayed by the electronic device before the captured image 8 is acquired; it can also be referred to as the target preview frame corresponding to the captured image 8. Because the preview frame 7 and the captured image 8 are acquired at different times and with a relatively long interval, the content displayed in the preview frame 7 and the captured image 8 may differ. In particular, in scenarios where both the preview frame 7 and the captured image 8 are images obtained through digital zoom, these differences may be very noticeable.
[0113] (2) Zooming refers to changing the focal length of a lens. Zooming methods include optical zoom and digital zoom, among others.
[0114] Digital zoom refers to cropping and / or enlarging an image using software algorithms to achieve the same effect as changing the focal length. Optical zoom relies on the structure of an optical lens to achieve zooming, changing the focal length of the target camera sensor by moving the lens elements.
[0115] (3) Zoom ratio indicates the range of focal lengths achieved by the target camera sensor under optical zoom and digital zoom functions. With the shooting distance between the target camera sensor and the subject remaining constant, a larger zoom ratio results in a larger area occupied by the subject in the captured image. A smaller zoom ratio results in a smaller area occupied by the subject in the captured image.
[0116] The base magnification is the maximum zoom magnification achievable through optical zoom. For example, if the zoom magnification of an electronic device is less than or equal to the base magnification, the focal length of the electronic device can be adjusted using optical zoom. If the zoom magnification of an electronic device is greater than the base magnification, the focal length of the electronic device can be adjusted using digital zoom.
[0117] In an exemplary scenario, the electronic device's base multiplier is 2.5X. For example... Figure 6As shown, the electronic device displays a main interface 701. This main interface 701 includes an application icon 702 for the camera application. In response to the user clicking the application icon 702, a shooting preview interface 703 is displayed. The shooting preview interface 703 also displays the current zoom level as 1X. The shooting preview interface 703 displays preview frames captured by the target camera sensor at a zoom level of 1X.
[0118] like Figure 6 As shown, the shooting preview interface 703 also includes a zoom bar 704. A sliding window 705 is displayed on this zoom bar 704. It can be understood that different positions within the zoom bar 704 correspond to different zoom ratios. The zoom ratio indicated by the overlapping position of the sliding window 705 and the zoom bar 704 is the currently selected zoom ratio. Additionally, the sliding window 705 can also display the numerical value of the currently selected zoom ratio.
[0119] During the display of the shooting preview interface 703, the electronic device receives a zoom instruction from the user. For example, a slide operation on the zoom bar 704. This slide operation can instruct the slider window 705 to adjust its position overlapping with the zoom bar 704, thereby indicating a change in the selected zoom magnification. After the user's slide operation ends, the electronic device can obtain the changed zoom magnification. For example, Figure 6 As shown, the user instructs to change the selected zoom ratio from 1X to 2X. After confirming that the zoom ratio enabled by the user instruction has been changed to 2X, the electronic device can display a shooting preview interface 706. This shooting preview interface 706 displays the current zoom ratio as 2X.
[0120] Accordingly, the target camera sensor can acquire images at a zoom ratio of 2X, and the raw image data acquired by the target camera sensor is processed to obtain the corresponding preview frame. Thus, the electronic device can display the preview frame from the target camera sensor on the shooting preview interface 706.
[0121] It is understood that sliding on the zoom bar 704 is only an example of zoom operation, and the embodiments of this application do not specifically limit the form of zoom operation.
[0122] like Figure 7 As shown, while the shooting preview interface 706 is displayed, the electronic device can also receive a zoom instruction from the user. Figure 7 As shown, the user instructs to change the selected zoom ratio from 2X to 5X. After confirming that the zoom ratio enabled by the user instruction has been changed to 5X, the electronic device can display a shooting preview interface 801. This shooting preview interface 801 displays the current zoom ratio as 5X.
[0123] in addition, Figure 7The digital zoom process is also shown. For example... Figure 7 As shown, after configuring the zoom ratio to 5X, the target camera sensor can acquire raw image data 1 at a base zoom ratio (2.5X). Then, one or more image processing steps are performed on the raw image data 1 to obtain the zoomable image 1. This zoomable image 1 can also be referred to as the image acquired at the base zoom ratio. Furthermore, the aforementioned one or more image processing steps can include image processing performed by image processing algorithm modules in the ISP and / or camera algorithm library, such as noise reduction, RGB domain or YUV domain processing, etc.
[0124] Next, the position of the crop region in the image to be zoomed is determined. The crop region is used to crop the field of view corresponding to the 5X zoom level, and the image area occupied by the crop region in the image to be zoomed corresponds to the field of view at 5X zoom level.
[0125] The aforementioned location describes the position information of the field-of-view cropping box in the image coordinate system corresponding to the zoomed image 1, and can represent the image area occupied by the field-of-view cropping box on the zoomed image 1. For example, the aforementioned location can be a set of position coordinate information, such as the image coordinates corresponding to a set of edge points of the field-of-view cropping box. As another example, the aforementioned location can include the image coordinates of a specific point of the field-of-view cropping box in the image coordinate system corresponding to the zoomed image 1, as well as the size of the field-of-view cropping box. For example, the specific point can be the center point or vertex of the field-of-view cropping box.
[0126] Alternatively, the crop function can be used to crop the field of view corresponding to 5x zoom from the image to be zoomed (image 1). Then, the field of view corresponding to 5x zoom is enlarged so that the size of the enlarged field of view is the same as that of the image to be zoomed (image 1), thus obtaining a preview frame at 5x zoom. The enlargement of the field of view can be achieved by using an upsampling algorithm.
[0127] In a possible embodiment, the above-mentioned image to be zoomed 1 may also be the raw image data collected by the target camera sensor. After the electronic device crops the field of view at 5X from the image to be zoomed 1, one or more image processing operations are performed on the field of view to obtain the preview frame at 5X.
[0128] In some embodiments, when the zoom level of the electronic device exceeds the base zoom level, if the electronic device enables ZSL camera mode, the electronic device can acquire attitude information indicating the attitude of the electronic device during the acquisition of raw image data of the preview frame. This attitude information may include gyroscope information and / or acceleration information.
[0129] For example, the corresponding gyroscope information can be obtained from the gyroscope sensor. The corresponding gyroscope information may include the deflection angle, deflection speed and deflection time of the electronic device. The corresponding gyroscope information can be used to characterize the attitude of the electronic device when acquiring the original image data of the preview frame.
[0130] For example, the corresponding acceleration information can be obtained from the accelerometer, and this acceleration information can also be used to characterize the posture of the electronic device when it acquires the original image data of the preview frame.
[0131] In addition, during the process of processing the original image data of the preview frame, the electronic device can determine the size of the field of view cropping box based on the zoom ratio, and input the pose information of the original image data into the EIS algorithm module to obtain the position of the field of view cropping box corresponding to the original image data.
[0132] For example, if the original image data is the first frame image acquired at the current zoom level, the EIS algorithm module can determine an offset of 1 based on the pose information of the original image data. This offset of 1 can be the offset of the center point of the field-of-view cropping box corresponding to the original image data relative to the center point of the original image data. Then, based on this offset of 1 and the center point of the original image data, the position of the field-of-view cropping box corresponding to the original image data can be determined.
[0133] If the original image data is not the first frame acquired at the current zoom level, the EIS algorithm module can determine offset 2 based on the pose information of the original image data. This offset 2 can be the offset of the field-of-view cropping box of the original image data relative to the field-of-view cropping box of the adjacent previous frame of original image data. Then, based on offset 2 and the position of the field-of-view cropping box of the adjacent previous frame of original image data, the position of the field-of-view cropping box of the original image data is determined.
[0134] In this way, the electronic device performs one or more image processing steps on the raw image data to obtain a processed image to be zoomed in. Then, based on the field of view cropping frame, a preview frame at the current zoom magnification is obtained from this image to be zoomed in. It is understandable that when the current zoom magnification is greater than the base magnification, the principle of generating a preview frame at the current zoom magnification can be found in [reference needed]. Figure 7 This will not be elaborated upon here.
[0135] In addition, the electronic device can store the field-of-view cropping frame position and the corresponding original image data. When displaying a preview frame, if the electronic device receives an instruction to take a picture, it can retrieve the original image data of the preview frame from the buffer queue, and also retrieve the field-of-view cropping frame position corresponding to the original image data. Then, based on the original image data and the field-of-view cropping frame, it generates a photographic image at the current zoom level.
[0136] For example, when ZSL camera mode is enabled, the electronic device captures a third image and caches it. Then, the electronic device displays a second preview frame corresponding to the third image. The content of the second preview frame is the same as the fifth image region in the third image; in other words, the second preview frame is an image obtained by digitally zooming the third image. The fifth image region is the area determined based on the position of the field-of-view cropping frame.
[0137] Furthermore, the aforementioned third image can be the raw image data acquired by the target camera sensor. Alternatively, the third image can also be an image obtained after performing one or more image processing steps on the raw image data. In response to the user's instruction to take a picture, the electronic device generates and stores a second captured image based on the cached third image. The display content of the second captured image is the same as the fifth image area in the third image. Since the second captured image and the second preview frame are images obtained based on the same raw image data (the third image) and the same field-of-view cropping frame, their display content is identical.
[0138] Understandably, when the current zoom level is greater than the base zoom level, the principle of generating a photograph at the current zoom level can be referenced from the principle of generating a preview frame at the current zoom level, which will not be elaborated here.
[0139] In some embodiments, when the zoom ratio of the electronic device exceeds the base zoom ratio, if the electronic device is in non-ZSL camera mode, the electronic device can also determine the size of the field-of-view cropping box based on the zoom ratio during the processing of the original image data of the preview frame, and input the pose information of the original image data into the EIS algorithm module to obtain the position of the field-of-view cropping box corresponding to the original image data. After receiving a user instruction to take a picture, the electronic device acquires the original image data corresponding to the captured image, and then obtains the position of the field-of-view cropping box of the target preview frame (the preview frame displayed by the electronic device when the shooting operation is detected), which can also be referred to as target position 1. The field-of-view cropping box of the target preview frame can be the field-of-view cropping box corresponding to the original image data of the target preview frame.
[0140] In some embodiments, the electronic device may perform one or more image processing operations on the original image data of the captured image to obtain a zoomable image corresponding to the captured image. Then, based on the zoomable image and the target position 1, a field of view at the current zoom level is cropped, and after being magnified, the captured image at the current zoom level is obtained.
[0141] like Figure 8 As shown, the electronic device displays previews sequentially. Figure 5 Preview Figure 6 and preview Figure 7With preview frame 7 displayed, a user-instructed action to take a picture is detected. In response to the user-instructed action to take a picture, image 8 is captured and stored.
[0142] like Figure 8 As shown, preview frame 5 is based on the field of view. Figure 5 The generated preview frame. The field of view is shown below. Figure 5 It is the image region cropped from the image to be zoomed 5 according to the field of view cropping frame 5. In addition, the field of view cropping frame 5 is located at position a on the image to be zoomed 5.
[0143] For example, the image to be zoomed in 5 can be the original image data acquired by the target camera sensor. For example, the image to be zoomed in 5 can also be image data obtained after processing the original image data, such as an image obtained after processing in the RGB domain or YUV domain.
[0144] like Figure 8 As shown, preview frame 6 is based on the field of view. Figure 6 The generated preview frame. The field of view is shown below. Figure 6 This refers to the image region cropped from the image to be zoomed, based on the field-of-view cropping frame 6. Furthermore, the field-of-view cropping frame 6 is located at position b on the image to be zoomed. The image to be zoomed 6 is similar to the image to be zoomed 5, but was acquired later than the image to be zoomed; details will not be elaborated upon here.
[0145] like Figure 8 As shown, preview frame 7 is based on the field of view. Figure 7 The generated preview frame. The field of view is shown below. Figure 7 This refers to the image region cropped from the image to be zoomed, based on the field-of-view cropping frame 7. Furthermore, the field-of-view cropping frame 7 is located at position c on the image to be zoomed. The image to be zoomed 7 is similar to the image to be zoomed 5, but was acquired later than both the images to be zoomed 5 and 6; details will not be elaborated upon here.
[0146] like Figure 8 As shown, image 8 is based on the field of view. Figure 8 The generated image. Wherein, the field of view... Figure 8 It is the image region cropped from the image to be zoomed in 8 according to the field of view cropping frame 7. The field of view cropping frame 7 is located at position c on the image to be zoomed in 8.
[0147] Understandable. Figure 8 In the scenario shown, the field-of-view cropping frame corresponding to the captured image inherits from the target preview image (i.e., the last preview frame displayed before the shooting operation, such as preview frame 7). In scenarios where non-ZSL camera mode is enabled, due to the long acquisition time interval between the zoomable image 8 and the zoomable image 7, the image content differs significantly. Therefore, the field of view cropped using the field-of-view cropping frame 7 is used. Figure 8and the field of view cropped using the field of view cropping frame 7 Figure 7 Significant differences can lead to discrepancies between the preview and the actual captured image.
[0148] To address the aforementioned issues, this application provides an image processing method applied to an electronic device. When the electronic device is in non-ZSL camera mode and the zoom level exceeds the base zoom level, the above method can improve the difference between the captured image and the target preview frame.
[0149] In some embodiments, as Figure 9 As shown, the above image processing method may include the following steps:
[0150] S101 displays a preview stream from the target camera sensor when the current zoom level is greater than the base zoom level.
[0151] In some embodiments, when a camera application (or other application with shooting capabilities) is running in the foreground of an electronic device, a shooting preview interface provided by that application can be displayed. This shooting preview interface can be used to display a preview stream from the target camera sensor. The preview stream includes multiple continuously acquired preview frames. When the current zoom level is greater than the base zoom level, the preview frames can be images obtained through digital zoom. The electronic device can cache the image to be zoomed for each preview frame, along with corresponding related information.
[0152] For example, the aforementioned relevant information may include pose information when the high-resolution image to be zoomed is acquired. Alternatively, it may include the position information of the field-of-view cropping frame used during the digital zoom processing based on the image to be zoomed.
[0153] S102, while displaying preview frame a, the user's instruction to take a picture is detected.
[0154] For example, the above-mentioned instruction to take a picture could be a user clicking the shooting control in the shooting preview interface. As another example, the above-mentioned instruction to take a picture could also be a user speaking a voice command to take a picture. As yet another example, the above-mentioned instruction to take a picture could also be a user making a gesture to indicate taking a picture, etc.
[0155] Here, preview frame a (such as the first preview frame) is an image obtained by digital zooming based on image data a (the first image). The aforementioned image data a can be the image to be zoomed corresponding to preview frame a.
[0156] For example, the image data 'a' mentioned above can be the raw image data acquired by the target camera sensor, or it can be the image obtained after the raw image data has undergone one or more image processing steps. It is understood that the one or more image processing steps mentioned above can include noise reduction, RAW domain processing, YUV domain processing, etc., and this application embodiment does not specifically limit these steps. In subsequent embodiments, image data 'a' will be described as the image after image processing.
[0157] In addition, preview frame a is the preview frame displayed by the electronic device when the user's instruction to take a picture is detected; it can also be called the target preview frame of the captured image.
[0158] S103, in response to the user's instruction to take a picture, acquires image data b for creating a photographed image.
[0159] In some embodiments, the electronic device may, in response to a user's instruction to take a picture, configure camera parameters suitable for acquiring a photographed image to the target camera sensor. These camera parameters may include parameters such as AE strategy, output format, and output size. Then, the electronic device obtains image data b (also referred to as the second image) through the target camera sensor. Image data b may be a zoomable image used to create the first frame of the photographed image. For example, if image data a is raw image data, image data b may be the first frame of raw image data acquired by the target camera sensor in response to the user's shooting operation. As another example, if image data a is a processed image, image data b may be an image obtained by performing one or more image processing operations on the first frame of raw image data. In subsequent embodiments, image data b will be described as an image obtained through image processing.
[0160] S104, perform image registration between the image data a corresponding to the preview frame a and the acquired image data b to obtain offset information.
[0161] In some embodiments, image registration is used to determine the location of the same object of interest (also known as the target object) between image data a and image data b. Offset information is then determined based on the location of the object of interest in image data a and image data b.
[0162] like Figure 10As shown, through image registration, image region 1101 (also known as the third image region) in image data a and image region 1102 (also known as the fourth image region) in image data b are determined to be image regions presenting the same object of interest. Then, based on the image coordinates of image region 1101 in image data a, image region 1101 is projected into image coordinate system 1103. Based on the image coordinates of image region 1102 in image data b, image region 1102 is projected into image coordinate system 1103. In image coordinate system 1103, the offset information between image region 1101 and image region 1102 is determined. For example, in image coordinate system 1103, the vector 1104 between the center point of image region 1101 and the center point of image region 1102 is used as the corresponding offset information.
[0163] For example, the object of interest (OP) can be a photographed object appearing in preview frame a. For instance, if multiple photographed objects appear in preview frame a, the object occupying the largest image area can be identified as the OOP. Alternatively, if multiple photographed objects appear in preview frame a, the OOP can be a photographed object whose image area includes the center point of preview frame a. Or, if multiple photographed objects appear in preview frame a, the OOP can be a photographed object appearing in multiple consecutive preview frames.
[0164] For example, the object of interest mentioned above can also be a static object captured in image data a. For instance, a static object can be an object whose position remains fixed, such as a building, facility, road sign, scenery, mountain peak, or forest. Another example is a static object whose relative position remains unchanged across multiple consecutive preview frames. Here, relative position refers to the positional relationship with an object whose position remains fixed.
[0165] In some embodiments, if the electronic device identifies the subject being tracked as being in motion, it can identify the subject in preview frame a as the object of interest. Specifically, the electronic device can identify the subject appearing in preview frame a and multiple consecutive preview frames as the subject being tracked. In other embodiments, if the electronic device identifies the subject being tracked as being static, it can identify any static subject in image data a as the object of interest.
[0166] In some embodiments, the image registration algorithm module in the camera algorithm library can be used to perform S104.
[0167] As one implementation method, the image registration algorithm module can perform registration based on features. For example, by detecting feature points (such as corners, edges, etc.) in image data a and image data b, and by matching these feature points, the image regions in image data a and image data b that indicate the same subject can be determined. Then, based on the display positions of the same object of interest in image data a and image data b, offset information, such as target offset, can be determined.
[0168] For example, feature detection can be achieved using algorithms such as SIFT (Scale Invariant Feature Transform), SURF (Speeded Robust Feature Transform), or Oriented FAST and Rotated Brief (ORB).
[0169] As an alternative implementation, the image registration algorithm module can be area-based registration. For example, similarity metrics such as cross-correlation or mutual information are typically used to evaluate the similarity between image regions, thereby determining the image regions that indicate the same subject between image data a and image data b. This method does not require pre-detection of feature points and can operate across the entire image region.
[0170] As another implementation, the image registration algorithm module can be model-based registration. For example, the image is assumed to be an instance of a parametric model, such as an affine transformation or perspective transformation. Image alignment is achieved by optimizing the model parameters.
[0171] As another implementation, the image registration algorithm module can be based on deep learning-based registration.
[0172] As another implementation, the image registration algorithm module can be multi-point registration. When there are multiple corresponding points in image data a and image data b, the multi-point registration method determines the image region between image data a and image data b that indicates the same subject by considering these points simultaneously.
[0173] As another implementation, the image registration algorithm module can employ rigid registration, affine registration, and non-rigid or elastic registration. Rigid registration assumes that changes in the display position of the same object between images only involve rotation and scaling, while affine and non-rigid registration can handle more complex transformations such as translation, rotation, scaling, and distortion. The choice of image registration algorithm depends on the specific application scenario, image characteristics, and required accuracy. In practical applications, it may be necessary to comprehensively consider factors such as image content, quality, and transformation type to select the most suitable registration method from the aforementioned rigid, affine, and non-rigid or elastic registration methods.
[0174] S105, obtain the target position 1 of the field of view cropping box corresponding to preview frame a.
[0175] Understandably, target position 1 (also known as the first position) can be used to indicate the position of the field-of-view cropping box on image data a. Using the field-of-view cropping box at target position 1, the field of view corresponding to preview frame a (target preview frame) can be cropped from image data a. The image area covered by this field-of-view cropping box on image data a can be called the first image area, and the size and display content of the first image area are the same as the field of view of preview frame a. In subsequent embodiments, target position 1 can also be referred to as the position of the field-of-view cropping box used to crop preview frame a.
[0176] In some embodiments, after determining the target location 1 and cropping the field of view corresponding to the preview frame a, the electronic device may cache the target location 1.
[0177] As one implementation, during the creation of preview frames, the electronic device determines the position information of the field-of-view cropping frame corresponding to each preview frame and stores this position information in a preset queue. The method for determining the position of the field-of-view cropping frame can be found in the aforementioned embodiments and will not be repeated here. Furthermore, the position information in the preset queue is arranged in chronological order of its storage time. When a user's shooting operation is detected, the position information with the latest storage time can be retrieved from the preset queue as target position 1.
[0178] As one implementation, during the creation of preview frames, the electronic device determines the position information of the field of view cropping frame corresponding to each preview frame and writes it to a preset storage location. In this preset storage location, the newly written position information replaces the original position information. Thus, when a user's shooting action is detected, the position information can be retrieved from the preset storage location and used as target position 1.
[0179] S106, Determine target position 2 based on target position 1 and offset information.
[0180] In some embodiments, both S105 and S106 can be executed by the offset processing module. Understandably, target position 2, also known as the second position, is the determined position of the field-of-view cropping box on image data b. The image area covered by the field-of-view cropping box located at target position 2 on image data b can be called the second image region.
[0181] For example, if target position 1 includes a set of position coordinate information, each position coordinate information corresponding to target position 1 can be moved according to the offset information to obtain the position coordinate information corresponding to target position 2. Using the position coordinate information contained in target position 2, the edge of the image region to be cropped on image data b can be determined.
[0182] For example, if target position 1 includes center point coordinates (or vertex coordinates) and the size of the field of view cropping frame, the electronic device can move the center point coordinates (or vertex coordinates) of target position 1 according to the offset information to obtain the center point coordinates (or vertex coordinates) corresponding to target position 2. Additionally, target position 2 also includes the size of the field of view cropping frame, which is related to the current zoom level. In some examples, the size of the field of view cropping frame in target position 2 can be the same as the size of the field of view cropping frame in target position 1. Thus, by using the center point coordinates (or vertex coordinates) and the size of the field of view cropping frame contained in target position 2, the image region to be cropped on image data b can be determined.
[0183] like Figure 11 As shown, the target position 1 of the field of view clipping box 1201 corresponding to preview frame a is obtained. The method for obtaining the target position 1 can be found in S104, and will not be described in detail here.
[0184] Understandably, preview frame a is the preview frame displayed when a user's shooting operation is detected, and can also be called the target preview frame. The position of the field-of-view cropping box 1201 corresponding to preview frame a can also be called target position 1. Through the field-of-view cropping box 1201 located at target position 1, the field-of-view map a can be cropped from the image data a. After upsampling the field-of-view map a, a preview frame a with the same size as the image data a is obtained.
[0185] like Figure 11 As stated above, based on the target position 1 and offset information of the field-of-view clipping frame 1201 (such as... Figure 10 The vector 1104 determined in the middle is used to determine the target position 2 of the field of view clipping frame 1202.
[0186] S107, Generate a photographed image based on target location 2 and image data b.
[0187] In some embodiments, the above-described S107 can be performed by the ISP.
[0188] like Figure 11 As shown, a field of view (referred to as the first field of view) is cropped from image data b according to the field of view cropping frame 1202. The size and display content of field of view b are the same as those of the second image region. After upsampling the field of view b, a photographic image with the same size as image data b is obtained, such as the first photographic image.
[0189] As can be seen, after determining the offset information between image data a and image data b through image registration, as... Figure 11 As shown, based on the field-of-view cropping box position and offset information corresponding to image data a, the field-of-view cropping box position for cropping the captured image is determined on image data b, thereby reducing the difference between the final captured image and preview image a (target preview frame).
[0190] Understandably, when preview frame a is displayed, the user instructs to take a picture, indicating that the user expects the captured image to be identical to preview frame a. However, when non-ZSL camera mode is enabled, the original image data a used to create preview frame a and the original image data b used to create the captured image are acquired at different times. During the time interval between acquiring original image data a and original image data b, not only may the position of the subject change (e.g., the subject is moving), but the position of the electronic device may also change (e.g., the user shakes while holding the electronic device). Of course, regardless of whether the position of the subject or the electronic device changes, image registration can ensure the consistency between the captured image and preview frame a, reducing the probability of the user retaking the shot. In non-ZSL camera mode, this improves the efficiency of human-computer interaction during photography and enhances the user experience.
[0191] In other scenarios, when the shooting operation is an instruction to capture a moving image, the electronic device responds to the shooting operation by continuously capturing and generating multiple frames of images to form a moving image.
[0192] In this scenario, firstly, the electronic device generates the first frame of the photographed image according to steps S103 to S107.
[0193] Secondly, after detecting the shooting operation, the electronic device acquires the second frame of raw image data through the target camera sensor. After performing one or more image processing operations on the second frame of raw image data, image data c can be obtained, that is, the zoomable image corresponding to the second frame of captured image. In addition, the electronic device can acquire its own attitude information 1 when acquiring the second frame of raw image data. Similarly, the electronic device also acquires and stores the corresponding attitude information 2 when acquiring the first frame of raw image data. Then, the image data c and the corresponding attitude information 1 are input into the EIS algorithm module. The EIS algorithm module determines the target position 3 of the field of view cropping box on the image data c based on the target position 2, attitude information 2, image data c, and attitude information 1. Then, based on the field of view cropping box located at the target position 3, the field of view map c is cropped from the image data c, and the second frame of captured image is generated based on the field of view map c.
[0194] Furthermore, after detecting the shooting operation, the electronic device continues to acquire the third frame of raw image data. After performing one or more image processing operations on the third frame of raw image data, image data d can be obtained, that is, the zoomable image corresponding to the third frame of the captured image.
[0195] When acquiring the third frame of raw image data, the corresponding pose information 4 is obtained. Image data d and pose information 4 are input into the EIS algorithm module. The EIS algorithm module determines the target position 5 based on image data d, pose information 4, pose information 1, and target position 3. Based on the field-of-view cropping box corresponding to target position 5, a field-of-view map d is cropped from image data d. The third frame of the captured image is obtained based on the field-of-view map d. Similarly, after the electronic device acquires the i-th frame of raw image data, the i-th frame of the captured image can be obtained in the same way, where i can be a positive integer greater than 2. The generated multi-frame captured images can be combined to form a dynamic photograph.
[0196] In some embodiments, as Figure 12 As shown, the above image processing method may include the following steps:
[0197] S201 displays a preview stream from the target camera sensor when the current zoom level is greater than the base zoom level.
[0198] S202, while displaying preview frame a, a user-instructed shooting operation is detected.
[0199] S203, in response to the user's instruction to take a picture, acquires image data b for creating the photographed image.
[0200] In some embodiments, the implementation details of S201 to S203 can be referred to S101 to S103 above, and will not be repeated here.
[0201] S204, Obtain the pose information 2 corresponding to image data b.
[0202] The aforementioned attitude information 2, also known as the first attitude information, includes the gyroscope information collected by the gyroscope sensor and / or the acceleration information collected by the accelerometer sensor when the target camera sensor acquires the original image data of image data b.
[0203] S205, input image data b and pose information 2 into the EIS algorithm module, and combine them with the relevant information corresponding to image data a to obtain the target position 4 of the field of view cropping box used for cropping the captured image.
[0204] In some embodiments, in S205 above, image data b is equivalent to the next frame image in the preview stream that was acquired later than image data a. Thus, even if there is a certain acquisition time interval between image data b and image data a, the EIS algorithm module can determine the target position 4 of the field-of-view clipping box based on image data b, pose information 2, and relevant information corresponding to image data a.
[0205] For example, the aforementioned related information may include the pose information 3 (also referred to as the second pose information) of image data a and the target position 1. The pose information 2 can characterize the pose 1 of the electronic device when acquiring image data b, and the pose information 3 can characterize the pose 2 of the electronic device when acquiring image data a. The EIS algorithm module determines the jitter (e.g., jitter amplitude) of the electronic device from capturing image data a to capturing image data b based on the difference between pose 1 and pose 2. Then, based on the target position 1 and the jitter, the EIS algorithm module can predict the target position 4 (i.e., the second position determined by the EIS algorithm module).
[0206] S206, Generate a photographed image based on the target location 4 and image data b.
[0207] In some embodiments, the implementation details of S206 can be referred to S107 in the foregoing embodiments, and will not be repeated here.
[0208] In addition, in scenarios where the shooting operation is an instruction to capture a dynamic image, after detecting the shooting operation, the first frame of raw image data is acquired, and one or more image processing operations are performed on the first frame of raw image data to obtain image data b. The electronic device generates the first frame of captured image based on image data b according to steps S204-S205. The electronic device acquires the second frame of raw image data through the target camera sensor, and after performing one or more image processing operations on the second frame of raw image data, image data c can be obtained. The electronic device can acquire attitude information 1, indicating the electronic device's posture, while acquiring the second frame of raw image data. The electronic device inputs image data c and the corresponding attitude information 1 into the EIS algorithm module. The EIS algorithm module determines the target position 3 corresponding to the field of view cropping box based on the target position 4, attitude information 2, image data c, and attitude information 1. Then, based on the field of view cropping box located at target position 3, the field of view map c is cropped from image data c, and the second frame of captured image is generated based on the field of view map c. This process is repeated to obtain multiple frames of captured images, forming a dynamic photograph.
[0209] In some embodiments, the electronic device is pre-configured with a preset magnification, which is greater than a base magnification. When the electronic device is in a non-ZSL camera mode, and the current zoom magnification is greater than or equal to the preset magnification, the following can be performed: Figure 9 or Figure 12 The method shown.
[0210] When the current zoom level is less than the preset zoom level, after acquiring image data b, the target position 1 of the field-of-view cropping frame for cropping preview frame a is obtained. The field-of-view cropping frame located at target position 1 is determined on image data b. Based on the field-of-view cropping frame at target position 1, a field-of-view image b is cropped from image data b. Then, the field-of-view image b is processed to obtain a captured image with the same image size as image data b.
[0211] For example, in scenarios where non-ZSL camera mode is enabled, the electronic device can reconfigure its zoom level in response to user-instructed zoom operations. If the zoom level is less than a preset level, the electronic device acquires a fourth image and displays the corresponding third preview frame. The content displayed in the third preview frame is the same as the sixth image region in the fourth image. The sixth image region is located at the third position in the fourth image; in other words, the field of view corresponding to the third preview frame is cropped from the fourth image based on the field of view cropping box at the third position.
[0212] During the display of the third preview frame, in response to the user's instruction to take a picture, the electronic device captures the fifth image and stores the corresponding third photographed image. The content displayed in the third photographed image is the same as the seventh image area in the fifth image, which is located at the third position in the fifth image. That is, the field of view corresponding to the third photographed image is cropped from the fifth image based on a field of view cropping frame at the third position.
[0213] As can be seen, when the zoom level is less than the preset zoom level, the position of the field of view cropping frame corresponding to the target preview frame can be used to obtain the captured image from the fifth image.
[0214] Understandably, the larger the zoom ratio, the smaller the field-of-view cropping frame. Conversely, the smaller the zoom ratio, the larger the field-of-view cropping frame. In non-ZSL camera mode, a smaller field-of-view cropping frame means its position has a greater impact on the content displayed in the captured image, potentially leading to a significant difference between the target preview frame and the captured image, exceeding what is acceptable to the user. Conversely, a larger field-of-view cropping frame means its position has less impact on the content displayed in the captured image, resulting in a smaller difference between the target preview frame and the captured image. Each of the aforementioned preset zoom ratios corresponds to a target size. If the field-of-view cropping frame is smaller than the target size, its position will significantly affect the content displayed in the captured image. If the field-of-view cropping frame is larger than or equal to the target size, its position will have an acceptable impact on the content displayed in the captured image. These preset zoom ratios can be empirical values obtained through testing; different electronic devices with different base zoom ratios may have different preset zoom ratios, and this application does not specifically limit this.
[0215] The foregoing embodiments described Figure 9 and Figure 12 The method shown is applied to scenarios involving capturing single photographic images and capturing animated images. This method can also be applied to video recording scenarios; the specific implementation process can be found in the previously described embodiments for capturing and generating animated images, and will not be elaborated upon here.
[0216] Some embodiments of this application also provide an electronic device, which may include a memory and one or more processors. The memory and processors are coupled. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the electronic device in the above method embodiments.
[0217] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.
[0218] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.
[0219] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0220] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0221] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0222] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0223] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0224] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An image processing method, characterized in that, The method further includes: When the non-zero-second delay ZSL camera mode is enabled, the electronic device displays a preview stream; When the electronic device displays the first preview frame in the preview stream, it detects a user-instructed shooting operation; the content displayed in the first preview frame is the same as the first image area in the first image captured by the electronic device, and the first image area is located at the first position of the first image; In response to a user-instructed shooting operation, the electronic device captures a second image and stores a first photographed image corresponding to the second image; wherein the display content of the first photographed image is the same as the second image area in the second image, and the second image area is located at a second position in the second image; the offset between the second position and the first position is a target offset, which is the offset between image areas presenting the same target object in the first image and the second image.
2. The method according to claim 1, characterized in that, After the electronic device acquires the second image, the method further includes: The electronic device performs image registration based on the first image and the second image to determine a third image region in the first image and a fourth image region in the second image. The third image region and the fourth image region are used to display the same target object. The electronic device determines the target offset based on the positions of the third image region and the fourth image region; The electronic device determines the second position based on the target offset and the first position.
3. The method according to claim 1 or 2, characterized in that, The target object includes any of the following: The subject captured in the first preview frame; The image area occupied includes the captured object at the center point of the first preview frame; The subject that occupies the largest image area in the first preview frame; The photographed object whose position remains fixed; The photographed object appears in both the first and second images and is in a static state.
4. The method according to claim 3, characterized in that, If the camera tracking object of the electronic device is in motion, and the first preview frame contains only one camera object, the target object is the camera object contained in the first preview frame; If the subject tracked by the lens of the electronic device is in motion, and the first preview frame contains multiple subjects, the target subject is the subject whose image area includes the center point of the first preview frame, or the subject that occupies the largest image area in the first preview frame. If the subject tracked by the lens of the electronic device is static, the target object is the subject whose position is fixed, or the subject appears in both the first image and the second image and is in a static state.
5. The method according to claim 1, characterized in that, When the electronic device acquires the second image, the method further includes: The electronic device acquires first attitude information; The electronic device inputs the first posture information and the second image into the electronic image stabilization (EIS) algorithm model, and combines the first position and the second posture information corresponding to the first image to obtain the second position. The second posture information can be the posture information obtained when the first image is captured.
6. The method according to claim 2 or 5, characterized in that, The method further includes: The electronic device crops a first field of view from the second image based on the second position; The electronic device generates the first photographed image based on the first field of view.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: When ZSL camera mode is enabled, the electronic device acquires a third image and caches the third image; The electronic device displays a second preview frame corresponding to the third image, and the content of the second preview frame is the same as the fifth image area in the third image; In response to a user's instruction to take a picture, the electronic device generates and stores a second photographed image based on the cached third image; wherein the display content of the second photographed image is the same as the fifth image area in the third image.
8. The method according to any one of claims 1-6, characterized in that, Before storing the first captured image corresponding to the second image, the method further includes: Ensure the current zoom level is greater than or equal to the preset zoom level.
9. The method according to claim 8, characterized in that, After storing the first captured image corresponding to the second image, the method further includes: The electronic device responds to user operation and configures the zoom ratio to a value less than the preset zoom ratio; The electronic device displays a preview stream; When the electronic device displays the third preview frame in the preview stream, a user-instructed shooting operation is detected; wherein, the content displayed in the third preview frame is the same as the sixth image region in the fourth image captured by the electronic device, and the sixth image region is located at the third position in the fourth image; In response to a user's instruction to take a picture, the electronic device captures a fifth image and stores a third image corresponding to the fifth image; wherein the display content of the third image is the same as the seventh image area in the fifth image, and the seventh image area is located at the third position in the fifth image.
10. The method according to claim 1, characterized in that, Before the electronic device acquires the second image, the method further includes: The electronic device configures target camera parameters in the camera sensor for acquiring captured images, wherein the target camera parameters are different from the camera parameters used for acquiring preview streams.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-10.
12. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-10.
13. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-10.