Image processing method and related apparatus

CN122845935APending Publication Date: 2026-09-29HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396990.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]但目前对图像的去模糊处理往往难以区分上述的虚焦模糊、运动模糊和抖动模糊,在针对图像进行去模糊处理时,不仅会使得画面变得清晰,也会因为去除了虚焦模糊而导致画面失去立体感和层次感

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845935A_ABST
    Figure CN122845935A_ABST
Patent Text Reader

Abstract

This application provides an image processing method and related apparatus, relating to the field of terminals. The method includes: an electronic device capturing a first image using a camera and recording the capturing parameters of the first image; the electronic device determining the depth information of the first image; the electronic device determining the focus tolerance range and focus position coordinates of the first image based on the capturing parameters of the first image; the electronic device determining the focus area of ​​the first image based on the depth information, focus tolerance range, and focus position coordinates of the first image; the electronic device determining the motion state information of the first image; and the electronic device performing deblurring processing on the first image based on the depth information, focus area, and motion state information of the first image to generate a fifth image. The fifth image is an image after removing motion blur and jitter blur from the first image, retaining the out-of-focus blur of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminals, and more particularly to an image processing method and related apparatus. Background Technology

[0002] With the development of the terminal field, it has become a trend for users to use electronic devices to photograph objects such as landscapes, people, and items. However, under different shooting conditions, the subject often appears blurred in the image. For example, when the subject is out of focus, it will appear blurred (called out-of-focus blur); another example is when the subject is moving while the camera is stationary, the subject will also appear blurred (called motion blur); yet another example is when the subject is stationary while the camera is moving, the subject will also appear blurred (called shake blur).

[0003] However, current image deblurring processes often struggle to distinguish between out-of-focus blur, motion blur, and jitter blur. When deblurring an image, not only does it make the image sharper, but removing out-of-focus blur can also cause the image to lose its sense of depth and dimension. Therefore, how to improve image sharpness while preserving its sense of depth and dimension has become a pressing issue. Summary of the Invention

[0004] This application provides an image processing method and related apparatus that can improve the clarity of an image while preserving its three-dimensionality and sense of depth.

[0005] In a first aspect, this application provides an image processing method, comprising: an electronic device capturing a first image via a camera, and recording the capturing parameters of the first image; wherein the first image includes a first object and a second object, the first object being displayed as a first blur, and the second object being displayed as a second blur, the first blur being the blur of the first object in the first image relative to the camera being in motion, and the second blur being the blur of the second object in the first image being in a defocused area; the electronic device acquiring depth information of the first image; wherein the depth information of the first image includes the distance from the actual object in three-dimensional space corresponding to each pixel in the first image to the camera; the electronic device based on the first image... The electronic device determines the focus tolerance range and focus position coordinates by shooting parameters; wherein, the focus tolerance range includes the distance between each object in the focus area and the camera; the electronic device determines the focus area of ​​the first image based on the focus tolerance range, focus position coordinates, and depth information of the first image; the electronic device determines the motion state information of the first image; wherein, the motion state information of the first image indicates the motion state of each subject in the first image; the electronic device generates a fifth image based on the first image, the depth information of the first image, the focus area of ​​the first image, and the motion state information of the first image; wherein, in the fifth image, the first object does not display a first blur, and the second object displays a second blur.

[0006] In one possible implementation, the first blur refers to the blurring of the first object in the first image relative to when the camera is in motion, including: the first blurring is the blurring of the first object in the first image when the first object is stationary and the camera is in motion; or, the first blurring is the blurring of the first object in the first image when the first object is in motion and the camera is stationary; or, the first blurring is the blurring of the first object in the first image when the first object is in motion and the camera is in motion.

[0007] In one possible implementation, the focus area in the first image is the region within the focus tolerance range centered on the focus position coordinates.

[0008] In one possible implementation, the electronic device indicates the depth information of the first image using a depth image; wherein the grayscale value of each pixel in the depth image is the distance from the actual object in the three-dimensional space corresponding to each pixel to the camera.

[0009] In one possible implementation, the electronic device indicates the focus area of ​​the first image using a focus area image; wherein pixels in the focus area of ​​the focus area image are marked with a first value, and pixels in the out-of-focus area of ​​the focus area image are marked with a second value; wherein the out-of-focus area is the region in the first image excluding the focus area.

[0010] In one possible implementation, the electronic device indicates the motion state information of the first image using a motion state image; wherein, objects in motion in the motion state image are marked with a third value, and objects in stationary state in the motion state image are marked with a fourth value.

[0011] In one possible implementation, the electronic device generates a fifth image based on the first image, the depth information of the first image, the focus area of ​​the first image, and the motion state information of the first image, including: the electronic device inputting the first image to a first encoder and outputting a first feature; the electronic device inputting the depth image, the focus area image, and the motion state image to a second encoder and outputting a second feature, a third feature, and a fourth feature; wherein the second encoder is different from the first encoder; the electronic device fusing the first feature, the second feature, the third feature, and the fourth feature and inputting them to a decoder; and the electronic device generating the fifth image through the decoder.

[0012] In a second aspect, this application provides an electronic device comprising: one or more processors and a memory. The memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, which the one or more processors invoke to cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0013] Thirdly, this application provides a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0014] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform a method as described in any of the possible implementations of the first aspect above.

[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, causes the processor to perform a method as described in any of the possible implementations of the first aspect above. Attached Figure Description

[0016] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0017] Figure 2 A schematic diagram illustrating a specific example of image deblurring processing provided in this application embodiment;

[0018] Figure 3 The implementation flow of an image processing method provided in this application embodiment;

[0019] Figure 4 A schematic diagram illustrating an example of a focus tolerance range provided in an embodiment of this application;

[0020] Figure 5 A schematic diagram illustrating a deblurring network for generating a fifth image, provided in an embodiment of this application;

[0021] Figure 6 A schematic diagram of the device architecture of an electronic device 100 provided in an embodiment of this application;

[0022] Figure 7 This is a schematic diagram illustrating the module interaction of an image processing method provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0024] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0025] Figure 1 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.

[0026] Figure 1 The electronic device 100 shown is the electronic device described in the embodiments of this application.

[0027] Electronic device 100 may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device and / or smart city device. The embodiments of this application do not impose any special restrictions on the specific type of electronic device 100.

[0028] like Figure 1 As shown, the electronic device 100 may include a processor 101, a memory 102, a wireless communication module 103 (optional), a display screen 104, and a camera 105, wherein:

[0029] Processor 101 may include one or more processor units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0030] The processor 101 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 101 is a cache memory. This memory can store instructions or data that the processor 101 has just used or that are used repeatedly. If the processor 101 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 101, and thus improves the efficiency of the system.

[0031] In some embodiments, the processor 101 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.

[0032] The memory 102 is coupled to the processor 101 and is used to store various software programs and / or multiple sets of instructions. In specific implementations, the memory 102 may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as ROM, flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 102 may also include combinations of the above types of memory. The memory 102 may also store some program code so that the processor 101 can call the program code stored in the memory 102 to implement the implementation method of the present application embodiment in the electronic device 100. The memory 102 may store an operating system, such as uCOS, VxWorks, RTLinux, or other embedded operating systems.

[0033] The wireless communication module 103 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 103 can be one or more devices integrating at least one communication processing module. The wireless communication module 103 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 101. The wireless communication module 103 can also receive signals to be transmitted from the processor 101, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via the antenna. In some embodiments, the electronic device 100 can also communicate via the Bluetooth module in the wireless communication module 103 (… Figure 1 (not shown), WLAN module ( Figure 1 (Not shown) The device transmits signals to detect or scan devices near electronic device 100 and establishes wireless communication connections with those devices to transmit data. The Bluetooth module can provide solutions for one or more Bluetooth communication methods, including basic rate / enhanced data rate (BR / EDR) or Bluetooth Low Energy (BLE), and the WLAN module can provide solutions for one or more WLAN communication methods, including Wi-Fi direct, Wi-Fi LAN, or Wi-Fi softAP.

[0034] The display screen 104 can be used to display images, videos, etc. The display screen 104 may include a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 104, where N is a positive integer greater than 1.

[0035] Camera 105 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an image signal processor (ISP) to be converted into a digital image signal. The ISP outputs the digital image signal to a digital signal processor (DSP) for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 105, where N is a positive integer greater than 1.

[0036] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0037] In some applications, users can use the camera 105 on an electronic device to photograph objects such as landscapes, people, and items. Under different shooting conditions, the subject often appears blurred in the image. For example, when the subject is out of focus, it will appear blurred (called out-of-focus blur); similarly, when the subject is moving while the camera is stationary, it will also appear blurred (called motion blur); and again, when the subject is stationary while the camera is moving, it will appear blurred (called jitter blur). However, in some implementations, the image deblurring process of electronic devices often struggles to distinguish between out-of-focus blur, motion blur, and jitter blur. When deblurring the image, not only does it make the image clearer, but removing out-of-focus blur can also cause the image to lose its sense of depth and dimension.

[0038] Figure 2 This is a schematic diagram illustrating a specific example of image deblurring processing provided in an embodiment of this application.

[0039] like Figure 2 As shown, image A captured by the electronic device through camera 105 may include a background area 201 and a toy car 202. The background area 201 is out of focus and therefore appears blurred in image A. The toy car 202 is in motion, thus appearing as motion blur in image A. In some implementations, the electronic device performs deblurring on image A, removing both the blurred background area 201 and the motion blur of the toy car 202, resulting in image B. It can be seen that both the background area 201 and the toy car 202 are clearer in image B than in image A. However, because the blurred background area 201 is removed, the image in image B loses its sense of depth and dimension compared to image A.

[0040] Therefore, this application provides an image processing method in which an electronic device can capture a first image (which may be an RGB format image) using a camera 105 and record the shooting parameters of the first image (e.g., focal length, aperture, pixel size, etc.). The first image may include motion blur, and / or jitter blur, and / or out-of-focus blur. The electronic device can determine the depth information of the first image. Then, the electronic device can determine the allowable focus range and the focus position coordinates of the first image based on the shooting parameters of the first image. Based on the depth information, the allowable focus range, and the focus position coordinates of the first image, the electronic device determines the focus area of ​​the first image. Next, the electronic device can determine the motion state information of the first image. Based on the depth information, the focus area, and the motion state information of the first image, the electronic device can perform deblurring processing on the first image to generate a fifth image. The fifth image is an image after removing the motion blur and jitter blur of the first image, retaining the out-of-focus blur. In this way, while improving the clarity of the first image, the three-dimensionality and layering of the image can also be preserved.

[0041] Figure 3 The following is an implementation flow of an image processing method provided in an embodiment of this application.

[0042] like Figure 3 As shown, the implementation process of this image processing method may include:

[0043] S301. Electronic device 100 captures a first image using a camera and records the capture parameters of the first image.

[0044] The first image can be an RGB format image, meaning it consists of three color channels: red (R), green (G), and blue (B). A pixel in the first image can be represented by a (R, G, B) triplet. For example, when the values ​​of the three color channels (i.e., the values ​​of the triplet) of pixel A are (255, 0, 0), that is, the R channel is 255, the G channel is 0, and the B channel is 0, the pixel is pure red; when the values ​​of the three color channels of pixel B are (0, 0, 255), that is, the R channel is 0, the G channel is 0, and the B channel is 255, the pixel is pure blue.

[0045] The shooting parameters of the first image can refer to the various settings used by the electronic device 100 to control the image quality of the first image captured by the camera when the electronic device 100 captures the first image. The shooting parameters of the first image may include, but are not limited to, focal length, aperture, pixel size, etc.

[0046] For example, electronic device 100 can capture a first image including a first object and a second object using one or more cameras. The cameras can focus on the first object, which is in motion while the camera is stationary; the second object is also stationary. When the second object is out of focus, it appears blurred in the first image, while the first object, being in motion, appears motion blurred in the first image.

[0047] Among them, motion blur and jitter blur can be called the first blur, and out-of-focus blur can be called the second blur.

[0048] S302. Electronic device 100 acquires depth information of the first image.

[0049] The depth information of the first image includes the depth value of each pixel in the first image. The depth value of a pixel can be used to indicate the distance from the camera to the actual object in the three-dimensional space corresponding to that pixel.

[0050] In one possible implementation, the depth information of the first image can be represented by a depth image of the first image. In the depth image of the first image, the depth value of each pixel can be encoded as the grayscale value of each pixel. That is, in the depth image of the first image, the grayscale value of each pixel is used to represent the distance from the actual object in the three-dimensional space corresponding to each pixel to the camera.

[0051] Specifically, the electronic device 100 can input the first image into a depth estimation network to obtain a depth image of the first image. The depth estimation network can be composed of neural networks such as convolutional neural networks (CNN), recurrent neural networks (RNN), and long short-term memory networks (LSTM); alternatively, the electronic device 100 can emit light pulses through a time-of-flight (TOF) sensor. These light pulses can interact with objects in the first image and then be reflected back to the TOF sensor. The TOF sensor can measure the round-trip time of the light pulses and calculate the depth value of each pixel corresponding to each object in the first image based on this round-trip time. Then, the electronic device 100 can generate a depth image of the first image based on the depth values ​​of each pixel in the first image calculated by the TOF sensor. It is understood that this application does not impose any restrictions on how the electronic device 100 obtains the depth information of the first image or how it generates the depth image of the first image.

[0052] S303. The electronic device 100 determines the focus tolerance range of the first image and the focus position coordinates of the first image based on the shooting parameters of the first image.

[0053] The focus tolerance range refers to the distance range between the closest and farthest points in three-dimensional space where the camera can capture a clear image. It's understandable that the "distance" in the description of the focus tolerance range refers to the distance between the camera and the viewfinder.

[0054] For example, if the focus tolerance range of the first image is 15cm to 1.5m, it means that when the camera captures the first image, the closest distance in three-dimensional space where the camera can capture a clear image is 15cm from the camera, and the farthest distance is 1.5m from the camera. In other words, for objects within a 15cm to 1.5m range from the camera in three-dimensional space, the camera can capture a clear image of them in the first image. The specific calculation method for the focus tolerance range will be described in detail in subsequent embodiments and will not be repeated here.

[0055] The focus position coordinates refer to the position coordinates of the focal point on the first image. These focus position coordinates can be the coordinates of the focal point in the pixel coordinate system. The origin of the pixel coordinate system is located at the upper left corner of the image, and the u-axis and v-axis are parallel to the two vertical edges of the image plane (for example, the u-axis can be horizontal to the right, and the v-axis can be vertically downward).

[0056] Specifically, the electronic device 100 can obtain the coordinates (x1, y1) of the focus point in the camera coordinate system from the shooting parameters of the first image. Then, the electronic device 100 can determine the coordinates (u1, v1) of the focus point on the first image based on the camera's internal parameters and the coordinates (x1, y1), which are the focus position coordinates.

[0057] In this system, the camera coordinate system has the optical center of the camera as its origin, the optical axis as the z-axis, and the x-axis and y-axis parallel to the two perpendicular sides of the image plane (for example, the x-axis can be perpendicular to the z-axis horizontally to the right, and the y-axis can be perpendicular to the z-axis vertically upwards). The camera's internal parameters refer to parameters describing the camera's internal properties, which may include focal length, optical center coordinates, distortion coefficients, etc.

[0058] S304. The electronic device 100 determines the focus area of ​​the first image based on the focus tolerance range of the first image, the focus position coordinates of the first image, and the depth information.

[0059] The focus area of ​​the first image refers to the area in the first image that can be clearly imaged (i.e., present a clear picture).

[0060] In one possible implementation, the electronic device 100 can use a focus area image (also called a focus map) to indicate the focus area of ​​the first image. The focus map can be a binary image, where each pixel is represented by a value of 0 or 1. A value of 1 in the focus map can be used to represent pixels in the focus area, and a value of 0 can be used to represent pixels in the out-of-focus area. Alternatively, the focus map can also use a value of 1 to represent pixels in the out-of-focus area and a value of 0 to represent pixels in the focus area; this application is not limited to this. It is understood that all areas in the first image other than the focus area are out-of-focus areas.

[0061] Specifically, the electronic device 100 can generate a focus map based on the focus tolerance range of the first image, the focus position coordinates of the first image, and the depth image of the first image (hereinafter referred to as the depth image).

[0062] Since the focal point coordinates on the first image and the focal point coordinates on the depth image are the same, the focal point coordinates of the first image are also the focal point coordinates of the depth image. Therefore, after determining the focal point coordinates of the first image, the electronic device 100 can look up multiple pixels in the depth image whose depth values ​​are within the focal tolerance range of the first image, centered on the focal point coordinates. The area formed by these multiple pixels is the focus area, and the area outside the focus area is the out-of-focus area. When the electronic device 100 determines the focus area and out-of-focus area of ​​the first image, it can generate a focus map.

[0063] S305. Electronic device 100 determines the motion state information of the first image. The motion state information of the first image is used to indicate the motion state of each subject in the first image.

[0064] The motion state of the subject in the first image can be divided into: the subject is in motion relative to the camera (e.g., the subject is moving while the camera is stationary, or the subject is stationary while the camera is moving, etc.) and the subject is stationary relative to the camera (e.g., the subject is stationary and the camera is also stationary, etc.).

[0065] In one possible implementation, the motion state information of the first image can be represented by a motion state image (also called a motion map). The motion state image can be a binary image, where each pixel is represented by a value of 0 or 1. A value of 1 in the motion state image can be used to represent a moving object region, and a value of 0 can be used to represent a stationary object region. Alternatively, the motion state image can also use a value of 1 to represent a stationary object region and a value of 0 to represent a moving object region; this application does not impose any limitations.

[0066] Specifically, the electronic device 100 can determine the motion state information of the first image based on optical flow, inter-frame difference algorithm, background subtraction algorithm, or feature matching algorithm. When the electronic device 100 determines the motion state information of the first image, it can generate a motion state image of the first image.

[0067] It is understood that the implementation method of electronic device 100 in determining the motion state information of the first image is not limited in this application.

[0068] S306. The electronic device 100 generates a fifth image based on the first image, the depth information of the first image, the focus area of ​​the first image, and the motion state information of the first image. The fifth image is the first image with motion blur and jitter blur removed, while retaining the out-of-focus blur.

[0069] For example, referring to the example shown in step S301, if the first object in the first image exhibits motion blur and the second object exhibits out-of-focus blur, then the pixels of the first object can be represented by a value of 1 in the focus map (the first object is in focus) and by a value of 1 in the motion state image (the first object is in motion); the pixels of the second object can be represented by a value of 0 in the focus map (the second object is out of focus) and by a value of 0 in the motion state image (the second object is stationary). The electronic device 100 can perform deblurring processing on the first image based on the depth information, the focus area, and the motion state information of the first image to generate a fifth image including the first and second objects. In this fifth image, the motion blur of the first object is removed, so the first object appears clear in the fifth image, but the out-of-focus blur of the second object is still retained.

[0070] Thus, by implementing the image processing method provided in this application, when the electronic device 100 performs deblurring processing on the first image, it can not only improve the clarity of the first image, but also retain the three-dimensionality and layering of the image.

[0071] Furthermore, the calculation method for the focus tolerance range in step S304 is explained in detail:

[0072] Figure 4 This is a schematic diagram illustrating an example of a focus tolerance range provided in an embodiment of this application.

[0073] like Figure 4 As shown, light rays from the subject are refracted by the camera and converge at the focal point of the imaging plane, resulting in a very clear image of the subject. Furthermore, objects within a certain distance of the subject can also be clearly imaged on the imaging plane. Here, "far point" refers to the farthest distance in three-dimensional space from which the camera can capture a clear image, and "near point" refers to the closest distance in three-dimensional space from which the camera can capture a clear image. "Near point distance" is the distance from the near point to the camera, and "far point distance" is the distance from the far point to the camera. "Subject distance" can also be called the focusing distance. The range between the near point distance and the far point distance is the aforementioned permissible focusing range.

[0074] Specifically, Formula 1 for calculating the nearest point distance can be as follows:

[0075]

[0076] Where D1 is the near point distance, F is the aperture value of the first image, f is the focal length of the first image, and B is the diameter of the permissible circle of confusion (PCOC).

[0077] Formula 2 for calculating the distance to the far point can be as follows:

[0078]

[0079] Where D2 is the distance to the far point, and the description of other parameters can be found in the calculation formula 1.

[0080] In one possible implementation, the formulas for calculating the near distance and the far distance can be simplified.

[0081] Specifically, the simplified formula for calculating the nearest point distance (Formula 3) can be as follows:

[0082]

[0083] Where p is the pixel size of the first image, and n is the number of pixels at the focus tolerance edge. The focus tolerance edge pixels refer to the pixels in the transition region between the focused and out-of-focus areas in the first image. Other parameters can be explained using Formula 1.

[0084] After simplification, Formula 4 for calculating the distance to the far point can be as follows:

[0085]

[0086] The explanation of the relevant parameters in calculation formula 4 can be found in the description of calculation formula 3.

[0087] In some application scenarios, the electronic device 100 can calculate the allowable focusing range based on the calculation formula 3 for the near point distance and the calculation formula 4 for the far point distance, which can improve the calculation efficiency of the electronic device 100.

[0088] In some application scenarios, the electronic device 100 can determine whether the first image was captured from a distance. If it was captured from a distance, the electronic device 100 calculates the allowable focus range based on formula 3 for near distance and formula 4 for far distance; if it was not captured from a distance, the electronic device 100 calculates the allowable focus range based on formula 1 for near distance and formula 2 for far distance. Regarding the determination of whether an image was captured from a distance, the electronic device 100 can base its judgment on the ratio of focal length to focusing distance. For example, if the focal length of the first image is greater than a first threshold, or the focusing distance of the first image is greater than a second threshold, the electronic device 100 determines that the first image was captured from a distance; if the focal length of the first image is less than the first threshold, or the focusing distance of the first image is less than the second threshold, the electronic device 100 determines that the first image was not captured from a distance. However, the electronic device 100 can also determine whether the first image was captured from a distance using other methods, and this application does not impose any limitations on this. In this way, the electronic device 100 can calculate the allowable focus range more accurately.

[0089] Furthermore, the implementation method of step S306, in which the electronic device 100 generates the fifth image, is described in detail:

[0090] Figure 5 This is a schematic diagram of a deblurring network used to generate a fifth image, as provided in an embodiment of this application.

[0091] Specifically, the electronic device 100 can input the first image, the focus map, the depth image, and the motion state image into the deblurring network to generate the fifth image. The deblurring network can be composed of neural networks based on CNN, RNN, LSTM, etc.

[0092] The first image has dimensions w*h*3, where w is the width, h is the height, and 3 is the number of channels. Since the first image is in RGB format with three color channels, it has 3 channels. The depth image, focus map, and motion image all have dimensions w*h*1. The depth image, focus map, and motion image all have the same dimensions as the first image, as do their corresponding heights. Because the depth image, focus map, and motion image are all binarized images, they each have 1 channel.

[0093] The first image is input to encoder 1 (which can be called the first encoder), and outputs feature 1; the depth image is input to encoder 2, and outputs feature 2; the focus map is input to encoder 2, and outputs feature 3; the motion state image is input to encoder 2 (which can be called the second encoder), and outputs feature 4. Encoder 1 and encoder 2 are different, but features 1 through 4 have the same number of channels.

[0094] Then, features 1 through 4 can be fused using a concat (i.e., cat) method, and the fused features 1 through 4 are then input into the decoder. Next, the decoder can output a fifth image. This fifth image is the first image with motion blur and jitter blur removed, while retaining the out-of-focus blur.

[0095] In this system, the depth image can be used by the deblurring network to maintain a natural out-of-focus blur effect for pixels with different depth values ​​in the first image during deblurring; the motion state image can indicate areas with motion blur in the first image, and the focus map can indicate areas with out-of-focus blur in the first image. It can be understood that the combination of the motion state image and the focus map can indicate whether there are areas in the first image where both out-of-focus blur and motion blur coexist.

[0096] Among them, feature 1 can be called the first feature, feature 2 can be called the second feature, feature 3 can be called the third feature, and feature 4 can be called the fourth feature.

[0097] Figure 6 This is a schematic diagram of the device architecture of an electronic device 100 provided in an embodiment of this application.

[0098] like Figure 6 As shown, the device architecture of electronic device 100 can be divided into five layers, from top to bottom: application layer, application framework layer, system library layer, kernel layer, and hardware layer.

[0099] like Figure 6 As shown, the application layer can include a series of application packages, such as camera, calendar, notes, browser, and gallery.

[0100] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, a content provider, a resource manager, a notification manager, etc., where:

[0101] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0102] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0103] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0104] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0105] In this embodiment of the application, the application framework layer may further include a depth information acquisition module, a motion state determination module, a focus area determination module, and a deblurring module.

[0106] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0107] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0108] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0109] The system library can include multiple functional modules. For example: surface manager, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.

[0110] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0111] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0112] A 2D graphics engine is a graphics engine for 2D drawing.

[0113] The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers and camera drivers. The display driver controls the display module in the electronic device 100 to display images captured by the camera, and the camera driver controls the camera to capture images.

[0114] The hardware layer may include hardware modules on the electronic device 100, such as a display module and a camera. The display module can be used to display images (e.g., a first image, a fifth image, etc.), and the camera can be used to capture images.

[0115] Figure 6 The device structure shown is for illustrative purposes only and does not constitute any limitation on this application.

[0116] Figure 7 This is a schematic diagram illustrating the module interaction of an image processing method provided in an embodiment of this application.

[0117] like Figure 7 As shown, the module interactions of this image processing method may include:

[0118] S701. The camera captures a first image, records the capture parameters of the first image, and copies the first image into three copies.

[0119] For instructions on this step, please refer to the description in S301, which will not be repeated here.

[0120] S702. The camera sends the first image to the depth information acquisition module.

[0121] S703. The depth information acquisition module generates a depth image based on the first image.

[0122] For instructions on this step, please refer to the description in S302, which will not be repeated here.

[0123] S704. The depth information acquisition module sends a depth image to the focus area determination module.

[0124] S705. The camera sends the shooting parameters of the first image to the focus area determination module.

[0125] S706. The focus area determination module determines the allowable focus range and the focus position coordinates of the first image based on the shooting parameters of the first image.

[0126] For instructions on this step, please refer to the description in S303, which will not be repeated here.

[0127] S707. The focus area determination module generates a focus map based on the focus tolerance range of the first image, the focus position coordinates of the first image, and the depth image.

[0128] For instructions on this step, please refer to S304; it will not be repeated here.

[0129] S708. The camera sends the first image to the motion state determination module.

[0130] S709. The motion state determination module generates a motion state image corresponding to the first image based on the first image.

[0131] For instructions on this step, please refer to S305; it will not be repeated here.

[0132] S710. The camera sends the first image to the deblurring module.

[0133] S711. The depth information acquisition module sends a depth image to the deblurring module.

[0134] S712. The focus area determination module sends the focus map to the deblurring module.

[0135] S713. The motion state determination module sends the motion state image to the deblurring module.

[0136] S714. The deblurring module generates a fifth image based on the first image, the depth image, the focus map, and the motion state image.

[0137] For instructions on this step, please refer to S306; it will not be repeated here.

[0138] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0139] This application also provides a computer program product, including a computer program that, when run on a processor, can implement the steps executed by the electronic device in the above-described method embodiments.

[0140] This application also provides a chip system, which includes a processing circuit interface circuit. The interface circuit receives instructions and transmits them to the processing circuit, which executes the instructions to cause the chip system to perform the steps executed by the electronic device in any of the method embodiments of this application. The chip system can be a single chip or a chip module composed of multiple chips.

[0141] The term "user interface (UI)" used in the specification and accompanying drawings of this application refers to the medium through which an application or operating system interacts and exchanges information with the user. It converts the internal form of information into a form acceptable to the user. The user interface of an application is source code written in a specific computer language such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the terminal device, ultimately presenting user-recognizable content, such as images, text, buttons, and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, images, and text. The attributes and content of controls in the interface are defined through tags or nodes, such as XML. <textview> 、 <imgview> 、 <videoview>Nodes define the controls contained in the interface. A node corresponds to a control or property in the interface, and after parsing and rendering, the node is presented as the content visible to the user. In addition, many applications, such as hybrid applications, often contain web pages within their interfaces. A web page, also known as a webpage, can be understood as a special control embedded in the application interface. Web pages are source code written in a specific computer language, such as Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript (JS), etc. Web page source code can be loaded and displayed as user-readable content by a browser or a web page display component with browser-like functionality. The specific content contained in a webpage is also defined through tags or nodes in the webpage source code; for example, HTML uses tags or nodes to define the content. 、 、 <video> 、 <canvas>Used to define the elements and attributes of a webpage.

[0142] The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0143] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0144] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0145] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.< / canvas> < / video> < / videoview> < / imgview> < / textview>

Claims

1. An image processing method, characterized in that, include: An electronic device captures a first image using a camera and records the capture parameters of the first image; wherein, the first image includes a first object and a second object, the first object is displayed as a first blur, and the second object is displayed as a second blur, the first blur being the blur of the first object in the first image when the camera is in motion, and the second blur being the blur of the second object in the first image when the second object is in a defocused area; The electronic device acquires the depth information of the first image; wherein, the depth information of the first image includes the distance from the actual object in the three-dimensional space corresponding to each pixel in the first image to the camera; The electronic device determines the focus tolerance range and focus position coordinates based on the shooting parameters of the first image; wherein, the focus tolerance range includes the distance between each object in the focus area and the camera; The electronic device determines the focus area of ​​the first image based on the focus tolerance range, the focus position coordinates, and the depth information of the first image. The electronic device determines the motion state information of the first image; wherein, the motion state information of the first image indicates the motion state of each subject in the first image; The electronic device generates a fifth image based on the first image, the depth information of the first image, the focus area of ​​the first image, and the motion state information of the first image; wherein, in the fifth image, the first object does not display a first blur, and the second object displays a second blur.

2. The method according to claim 1, characterized in that, The first blur refers to the blurring of the first object in the first image relative to when the camera is in motion, including: The first blur refers to the blurring of the first object in the first image when the first object is stationary and the camera is moving; Alternatively, the first blur is the blurring of the first object in the first image when the first object is moving and the camera is stationary; Alternatively, the first blur is the blurring of the first object in the first image when the camera moves.

3. The method according to claim 1, characterized in that, The focus area in the first image is the region within the allowable focus range centered on the focus position coordinates.

4. The method according to any one of claims 1-3, characterized in that, The electronic device indicates the depth information of the first image using a depth image; wherein, the grayscale value of each pixel in the depth image is the distance from the actual object in the three-dimensional space corresponding to each pixel to the camera.

5. The method according to claim 4, characterized in that, The electronic device indicates the focus area of ​​the first image using a focus area image; wherein, the pixels of the focus area in the focus area image are marked with a first value, and the pixels of the out-of-focus area in the focus area image are marked with a second value; wherein, the out-of-focus area is the area in the first image other than the focus area.

6. The method according to claim 5, characterized in that, The electronic device indicates the motion state information of the first image using a motion state image; wherein, objects in motion in the motion state image are marked with a third value, and objects in stationary state in the motion state image are marked with a fourth value.

7. The method according to claim 6, characterized in that, The electronic device generates a fifth image based on the first image, the depth information of the first image, the focus area of ​​the first image, and the motion state information of the first image, including: The electronic device inputs the first image to the first encoder and outputs the first feature; The electronic device inputs the depth image, the focus area image, and the motion state image into the second encoder and outputs a second feature, a third feature, and a fourth feature; wherein the second encoder is different from the first encoder; The electronic device fuses the first feature, the second feature, the third feature, and the fourth feature, and inputs them into the decoder; The electronic device generates the fifth image through the decoder.

8. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-7.

9. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-7.

11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, causes the processor to perform the method as described in any one of claims 1-7.