Image deblurring method and device, electronic equipment and readable storage medium

By obtaining the jitter amount to determine the target deblurring strategy and using the trained deblurring network to process the image, the problem of poor deblurring effect in the existing technology is solved, and the image clarity and user experience are improved.

CN120282025BActive Publication Date: 2026-08-25HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311871057.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2026-08-25
Estimated Expiration
2043-12-28

AI Technical Summary

Technical Problem

In the existing technology, electronic devices are prone to processing clear image content as blurry during the image deblurring process, or blurry images are not completely restored to clarity, resulting in poor deblurring effect and poor user visual experience.

Method used

By acquiring the amount of jitter in electronic devices, a target deblurring strategy (local or global) is determined, and the trained deblurring network is used for image processing. The training samples are then optimized by combining motion mask images to improve the deblurring effect.

Benefits of technology

It improves the image deblurring effect, provides users with a better visual experience, and enhances image clarity and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282025B_ABST
    Figure CN120282025B_ABST
Patent Text Reader

Abstract

The application discloses an image deblurring method and device, electronic equipment and readable storage medium. The method comprises the following steps: acquiring a to-be-processed image; acquiring a shaking amount of the electronic equipment, wherein the shaking amount is used to represent the shaking degree of the electronic equipment within a time period when a camera of the electronic equipment collects the to-be-processed image; determining a target deblurring strategy according to the shaking amount, wherein the target deblurring strategy matches the blurring degree of the to-be-processed image; and performing deblurring processing on the to-be-processed image according to the target deblurring strategy to obtain a clear image. The image deblurring method provided by the application can improve the image deblurring effect, thereby better meeting the visual experience of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and more specifically, to an image deblurring method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] When using the camera of an electronic device (e.g., a mobile phone) to photograph a subject, if the electronic device moves relative to the subject (e.g., the electronic device shakes or the subject moves), the image captured by the camera will be blurry. Currently, blurry images can be restored to clear images (i.e., unblurred images) using deblurring techniques. In related technologies, the electronic device inputs the acquired blurry image into a deblurring network for deblurring processing. This can easily lead to the deblurring network processing originally clear image content into blurry image content, or it can easily lead to the deblurring network failing to restore all blurry image content to clear content. In other words, traditional techniques suffer from poor deblurring effects, resulting in a poor user visual experience.

[0003] Therefore, improving the image deblurring effect has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides an image deblurring method, apparatus, electronic device, and readable storage medium, which can improve the image deblurring effect and thus better meet the user's visual experience.

[0005] In a first aspect, this application provides an image deblurring method applied in an electronic device. The method includes: acquiring an image to be processed; acquiring the jitter amount of the electronic device, wherein the jitter amount represents the degree of jitter of the electronic device during the time period in which the camera of the electronic device acquires the image to be processed; determining a target deblurring strategy based on the jitter amount, wherein the target deblurring strategy matches the degree of blur of the image to be processed; and performing deblurring processing on the image to be processed according to the target deblurring strategy to obtain a clear image.

[0006] The image to be processed is an image captured by the camera of an electronic device onto a subject. For example, the image to be processed can be, but is not limited to, a locally blurred image or a globally blurred image. It is understood that when the blur in an image is localized, the image can be called a locally blurred image, meaning that pixels in a localized area of ​​the image are blurred. When the blur in an image is globalized, the image can be called a globally blurred image, meaning that pixels in all areas of the globally blurred image are blurred.

[0007] Matching the target deblurring strategy with the degree of blurring of the image to be processed means that the target deblurring strategy produces a clear image after deblurring the image to be processed, and that the clear image can meet the user's deblurring requirements.

[0008] For example, when the image to be processed is a globally blurred image, the target deblurring strategy is a global deblurring strategy. Conversely, when the image to be processed is a locally blurred image, the target deblurring strategy is a local deblurring strategy.

[0009] In the above technical solution, after the electronic device acquires the image to be processed, it first determines the deblurring strategy (i.e., the target deblurring strategy) based on the jitter level. Since the jitter level represents the degree of shaking of the electronic device during the time period when the camera acquires the image, the target deblurring strategy determined based on the jitter level matches the degree of blur in the image. Then, the electronic device performs deblurring processing on the image according to the target deblurring strategy determined based on the jitter level, resulting in a clear image. This method avoids the problem in traditional technologies where the deblurring strategy and the degree of blur in the acquired image are mismatched, leading to poor deblurring results. Therefore, this method can improve the image deblurring effect, thereby better satisfying the user's visual experience.

[0010] In one possible implementation, the target deblurring strategy includes a local deblurring strategy or a global deblurring strategy, and the target deblurring strategy is determined based on the jitter amount, including: if the jitter amount meets a first jitter level, the target deblurring strategy is determined to be a local deblurring strategy; if the jitter amount meets a second jitter level, the target deblurring strategy is determined to be a global deblurring strategy, wherein the second jitter level is greater than the first jitter level.

[0011] Satisfying the first level of jitter means that the electronic device experiences minimal jitter during the time the camera captures the image to be processed. In this case, the image captured by the camera is locally blurred. Satisfying the second level of jitter means that the electronic device experiences significant jitter during the time the camera captures the image to be processed. In this case, the image captured by the camera is globally blurred.

[0012] The presentation format of the first and second jitter levels is not specifically limited. For example, at least one of the first and second jitter levels can be a specific jitter level. For example, at least one of the first and second jitter levels can be a jitter level within a preset range, without specific limitations.

[0013] In the above technical solution, when the electronic device experiences significant shaking during the time its camera captures the image to be processed, it determines that the deblurring strategy for the image to be processed is a global deblurring strategy. When the electronic device experiences relatively little shaking during the time its camera captures the image to be processed, it determines that the deblurring strategy for the image to be processed is a local deblurring strategy. In summary, this method can achieve the purpose of deblurring both locally and globally blurred images, thereby improving the deblurring effect of both globally and locally blurred images and better meeting the user's visual experience.

[0014] In another possible implementation, the image to be processed is deblurred according to a target deblurring strategy to obtain a clear image. This includes: using a target deblurring network to deblurr the image to be processed to obtain a clear image. The target deblurring network is a neural network model obtained by training an initial deblurring network using training samples. The training samples include training images, label images, and motion mask images. The training images are blurred images, the label images are clear images corresponding to the training images, and the motion mask images are images that include features of the blurred regions of the training images.

[0015] The initial deblurring network can have a decoder-decoder structure; for example, the initial deblurring network can be, but is not limited to, a Unet network.

[0016] The number of training samples can be one or more, without specific limitation. Similarly, when there are multiple training samples, there is no specific limitation on whether any two training samples are identical.

[0017] In one example, when the training image is a locally blurred image, that is, a local area in the training image is blurred content and the area in the training image excluding the local area is sharp content, the motion mask image includes the features of the local area in the training image. For example, the area with a pixel value of "1" in the motion mask image corresponds to the features of the local area (i.e., the blurred area) in the training image, and the area with a pixel value of "0" in the motion mask image corresponds to the area in the training image excluding the local area (i.e., the sharp area).

[0018] In one example, when the training image is a globally blurred image, meaning that all regions in the training image are blurred, the motion mask image includes features of all regions in the training image, for example, the pixel value of all regions in the motion mask image is "1".

[0019] In the above technical solution, the training samples used by the electronic device to train the initial deblurring network include not only the training image and the corresponding label image, but also the corresponding motion mask image. Unlike traditional methods that only input the training image and label image into the initial network model, this image deblurring method utilizes a target deblurring network trained on the initial deblurring network using these training samples (i.e., the training image, label image, and motion mask image). This improves the deblurring accuracy of the trained target deblurring network. Subsequently, by performing deblurring processing on the image to be processed based on this trained target deblurring network, the image deblurring effect can be improved, thereby better satisfying the user's visual experience.

[0020] In another possible implementation, the initial deblurring network includes an encoder and a decoder, wherein the decoder includes a first network module and a second network module, and the target deblurring network is specifically a neural network model obtained by training the encoder, the adjusted first network module, and the second network module according to a first loss value, wherein the first loss value is determined based on the difference between the predicted image and the training image, the predicted image is the image obtained by the second network module decoding the second feature image, the second feature image is the image obtained by the adjusted first network module extracting features from the first feature image, the adjusted first network module adjusts the parameters of the first network module according to the second loss value, the first feature image is the image obtained by the encoder extracting features from the training image, and the second loss value is determined based on the difference between the first feature image and the motion mask image.

[0021] In the above technical solution, the first network module in the decoder obtains the first feature image obtained by the encoder after feature extraction processing of the training image. Then, it adjusts the weights (i.e., parameters) of the first network module based on the difference between the first feature image and the motion mask image corresponding to the pre-acquired training image. Next, the second network module in the decoder obtains the predicted image based on the optimized first feature image (i.e., the second feature image) output by the adjusted first network module. Finally, the weights of the encoder, the adjusted first network module, and the second network module are adjusted according to the difference between the predicted image and the training image to obtain the trained network (i.e., the target deblurring network). During this training process, because the feature image (i.e., the first feature image) extracted by the encoder is optimized using the motion mask image corresponding to the pre-acquired training image, the optimized first feature image (i.e., the second feature image) can meet the deblurring requirements. Therefore, this method can improve the image deblurring effect, thereby better satisfying the user's visual experience.

[0022] In another possible implementation, when the training image is a globally blurred image, the entire region of the training image is blurred, the motion mask image includes the features of the entire region, and the target deblurring network is a global deblurring network; when the training image is a locally blurred image, a portion of the training image is blurred, the motion mask image includes the features of the portion of the region, and the target deblurring network is a local deblurring network.

[0023] In the above technical solution, since the blur types of the training samples (e.g., global blur or local blur) are different, the corresponding motion mask images are also different. Thus, when training the initial deblurring network based on the corresponding training samples, the network can be guided to learn the corresponding type of blur processing capability. This results in a globally deblurring network with strong global deblurring capability and a locally deblurring network with strong local deblurring capability. Subsequently, the locally deblurring network can be used to deblur locally blurred images, and the globally deblurring network can be used to deblur globally blurred images. This improves both the deblurring effect of globally blurred images and the deblurring effect of locally blurred images, better meeting the user's visual experience.

[0024] In another possible implementation, obtaining the jitter of the electronic device includes: obtaining angular acceleration data, wherein the angular acceleration data is used to represent the pose data of the electronic device during the time period in which the camera acquires the image to be processed; and determining the jitter based on the angular acceleration data.

[0025] For example, a gyroscope sensor located in an electronic device can detect and record the pose data of the electronic device at different times. In this case, the aforementioned angular acceleration data can be obtained by the electronic device from the gyroscope sensor located in the electronic device.

[0026] In the above technical solution, the electronic device determines the amount of jitter based on the pose data (i.e., angular acceleration data) of the electronic device during the time period of the image to be processed acquired by the camera. This implementation process is relatively simple and can improve the efficiency of image deblurring.

[0027] In another possible implementation, the angular acceleration data includes multiple angular accelerations, wherein the multiple angular accelerations correspond to multiple target points in the image to be processed, each angular acceleration is the pose data of the electronic device when the camera captures the corresponding target point, the multiple target points correspond to multiple first two-dimensional positions, the position of each target point in the image to be processed is the corresponding first two-dimensional position, and the jitter amount is determined based on the angular acceleration data, including: obtaining multiple second two-dimensional positions based on the multiple angular accelerations and the multiple first two-dimensional positions; obtaining multiple blur values ​​corresponding to the multiple target points based on the multiple second two-dimensional positions and the multiple first two-dimensional positions; and determining the jitter amount based on the multiple blur values.

[0028] In one example, multiple target points correspond one-to-one with multiple image patches in the image to be processed, with each target point being a point within its corresponding image patch. For example, each target point could be the center point of its corresponding image patch. Or, for example, each target point could be the bottom-left corner point of its corresponding image patch.

[0029] In the above technical solution, the electronic device measures the blur level of each image block by the pixels in each image block. Then, it determines the jitter of the electronic device based on multiple blur values ​​corresponding to multiple image blocks. This method can improve the accuracy of the obtained jitter of the electronic device. Subsequently, the target deblurring strategy determined based on the jitter is more compatible with the image to be processed, thereby improving the deblurring effect of globally blurred images and locally blurred images, and better meeting the user's visual experience.

[0030] In another possible implementation, the jitter amount is determined based on multiple fuzzy values, including: if the smallest fuzzy value among the multiple fuzzy values ​​is less than a preset threshold, the jitter amount is determined to meet a first jitter level; if the smallest fuzzy value among the multiple fuzzy values ​​is greater than or equal to a preset threshold, the jitter amount is determined to meet a second jitter level, wherein the second jitter level is greater than the first jitter level.

[0031] In the above technical solution, the electronic device determines the degree of jitter by comparing the relationship between multiple blur values ​​and a preset threshold. This implementation process is relatively simple and can improve the efficiency of image deblurring.

[0032] In another possible implementation, before acquiring the image to be processed, the method further includes: displaying a shooting interface; acquiring the image to be processed, including: obtaining the image to be processed through a camera in response to a shooting operation on the shooting interface; and after performing deblurring processing on the image to be processed according to a target deblurring strategy to obtain a clear image, the method further includes: displaying an interface including the clear image.

[0033] In the above technical solution, the electronic device acquires the image to be processed in response to the user's shooting operation. Then, the electronic device processes the image using a target deblurring strategy. Finally, the electronic device displays an interface with a clear image to the user. The process of the electronic device performing deblurring on the image to be processed is performed in the background, that is, the user is unaware of the process, which can improve the user's visual experience.

[0034] Secondly, this application provides an image deblurring apparatus applied in an electronic device. The apparatus includes a processing unit, wherein the processing unit is configured to: acquire an image to be processed, wherein the image to be processed is an image captured by a camera of the electronic device of a photographed object; acquire the jitter amount of the electronic device, wherein the jitter amount is used to represent the degree of jitter of the electronic device during the time period in which the camera of the electronic device acquires the image to be processed; determine a target deblurring strategy based on the jitter amount; and perform deblurring processing on the image to be processed according to the target deblurring strategy to obtain a clear image.

[0035] Thirdly, this application provides an electronic device including a unit for performing any of the methods in the first aspect. The device may be a terminal device or a chip within a terminal device. The device may include an input unit and a processing unit.

[0036] When the device is a terminal device, the processing unit may be a processor, and the input unit may be a communication interface; the terminal device may also include a memory for storing computer program code, which, when the processor executes the computer program code stored in the memory, causes the terminal device to execute any of the image deblurring methods in the first aspect.

[0037] When the device is a chip within a terminal device, the processing unit can be an internal processing unit of the chip, and the input unit can be an output interface, pin, or circuit, etc.; the chip may also include a memory, which can be an internal memory of the chip (e.g., registers, cache, etc.) or an external memory (e.g., read-only memory, random access memory, etc.); the memory is used to store computer program code, and when the processor executes the computer program code stored in the memory, the chip performs any of the image deblurring methods in the first aspect.

[0038] In one possible implementation, the memory is used to store computer program code; the processor executes the computer program code stored in the memory, and when the computer program code stored in the memory is executed, the processor is used to perform any of the image deblurring methods in the first aspect.

[0039] Fourthly, this application provides a computer-readable storage medium storing computer program code that, when executed by an image deblurring device, causes the image deblurring device to perform any of the image deblurring methods in the first aspect.

[0040] Fifthly, this application provides a computer program product comprising: computer program code, which, when run by an image deblurring device, causes the image deblurring device to perform any of the image deblurring methods described in the first aspect.

[0041] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0042] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the user interface and shooting results of an electronic device when it takes a picture of the subject.

[0044] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0045] Figure 3 This is a schematic diagram of an image deblurring method provided in an embodiment of this application.

[0046] Figure 4 yes Figure 3 The provided image deblurring method involves a schematic diagram of a locally blurred image and a sharp image.

[0047] Figure 5 yes Figure 3 The provided image deblurring method involves a schematic diagram of a globally blurred image and a sharp image.

[0048] Figure 6 yes Figure 3 A schematic diagram of multiple image blocks corresponding to the blurred image involved in the provided image deblurring method.

[0049] Figure 7 This is a schematic diagram of the system to which the training method provided in the embodiments of this application is applicable.

[0050] Figure 8 This is a block diagram of the hardware structure of the neural network processor provided in the embodiments of this application.

[0051] Figure 9 This is a schematic diagram of the architecture of a deblurred network model provided in an embodiment of this application.

[0052] Figure 10 Is training the above Figure 9 This is a schematic diagram of a training sample of a deblurring network model.

[0053] Figure 11 Is training the above Figure 9 This is a schematic diagram of another training sample for the deblurring network model.

[0054] Figure 12 This is a schematic diagram of a model training method provided in an embodiment of this application.

[0055] Figure 13 This is a schematic diagram of an image deblurring method provided in an embodiment of this application.

[0056] Figure 14 This is a schematic diagram of a user interface displayed by an electronic device when the electronic device executes the image deblurring method provided in the embodiments of this application.

[0057] Figure 15 This is a schematic diagram of another user interface displayed by an electronic device when the electronic device executes the image deblurring method provided in the embodiments of this application. Detailed Implementation

[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0059] With the development of mobile terminals and the maturity of image processing technology, people's requirements for terminal photography are also gradually increasing. In practical applications, when there is relative motion between the camera of an electronic device and the subject being photographed, it can lead to blurred images and a poor user visual experience. For example, if the camera of the electronic device is stationary while the subject moves during shooting, the resulting image will be locally blurred, meaning that the area where the moving subject is located will be blurred. Conversely, if the camera of the electronic device moves during shooting, regardless of whether the subject is moving or not, the resulting image will be globally blurred, meaning that the entire image will be blurred.

[0060] For example, taking the use of a mobile phone's camera to photograph a subject, please refer to... Figure 1 As shown in Figure (1), the shooting interface S1 provided by the mobile phone presents subjects including people and buildings. In response to the user triggering the camera control 10 in the shooting interface S1, the mobile phone takes a picture of the subject to obtain a captured image. During the shooting process, if the mobile phone shakes, regardless of whether the subject moves, the captured image obtained by the mobile phone is a globally blurred image. For example, Figure 1 The image shown in the photo album interface S2 provided by the mobile phone in Figure (2) is a globally blurred image, meaning that the entire area of ​​the image shown in the photo album interface S2 is blurred. During the shooting process, if the mobile phone is stationary (i.e., there is no shaking) and the subject being photographed (e.g., a person walking during the shooting process) moves, the image obtained by the mobile phone from photographing the subject will be a partially blurred image. For example, Figure 1 The image shown in the album interface S3 provided by the mobile phone in (3) is a partially blurred image. Since the subject is in motion during the shooting process, some areas of the image shown in the album interface S3 are blurred (i.e., the area of ​​the person is blurred), and the remaining areas of the image are clear (i.e., the area excluding the area of ​​the person).

[0061] In related technologies, electronic devices directly input the acquired blurry image into a deblurring network for deblurring processing. This can easily lead to the deblurring network processing the originally clear image content in the blurry image into blurry image content, or it can easily lead to the deblurring network failing to restore all the blurry image content in the blurry image into a clear image. In other words, traditional technologies have the problem of poor deblurring effect, resulting in a poor user visual experience.

[0062] As mentioned earlier, to address the poor image deblurring effect in traditional technologies, the image deblurring method provided in this application adopts the following technical solution: After acquiring the image to be processed, the electronic device first determines the deblurring strategy (i.e., the target deblurring strategy) based on the jitter level. Since the jitter level represents the degree of shaking of the electronic device during the time period when the camera acquires the image to be processed, the target deblurring strategy determined based on the jitter level matches the degree of blur in the image to be processed. Then, the electronic device performs deblurring processing on the image to be processed according to the target deblurring strategy determined based on the jitter level to obtain a clear image. This method avoids the problem in traditional technologies where the electronic device directly performs deblurring processing on the acquired image to be processed, resulting in a mismatch between the deblurring strategy and the degree of blur in the image. Therefore, this method can improve the image deblurring effect, thereby better satisfying the user's visual experience.

[0063] The image deblurring method provided in this application can be applied to electronic devices, such as mobile phones, smart screens, tablets, wearable electronic devices, in-vehicle electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), projectors, in-vehicle devices, etc. In other words, the embodiments of this application do not impose any restrictions on the specific type of electronic device.

[0064] The structure of the electronic device applicable to the image deblurring method provided in this application will now be described with reference to the accompanying drawings.

[0065] Figure 2 This is a schematic diagram illustrating the structure of an electronic device provided in an embodiment of this application. See also... Figure 2The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0066] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more than Figure 1 The diagram shows more or fewer components, or combinations of some components, or splitting of some components, or different arrangements of components. Figure 1 The components shown can be implemented in hardware, software, or a combination of both.

[0067] The processor 110 is used to execute the image deblurring method provided in the embodiments of this application. The processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0068] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0069] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0070] In some embodiments, the processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0071] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0072] Display screen 194 is used to display user interfaces, images, and videos, etc. For example, display screen 194 can display the following... Figure 4 The image shown (e.g., a clear image) and / or Figure 5 The image shown (e.g., a clear image), and Figure 14 and / or Figure 15The user interface is shown. The display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or L displays 194, where L is an integer greater than 1.

[0073] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular acceleration of the electronic device 100 about three axes (i.e., the x-axis, y-axis, and z-axis). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in scenarios such as navigation and motion-sensing games.

[0074] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K is used to detect touch operations applied to or near it. For an example, please see below. Figure 14 In Figure (1), after the user triggers the camera control 10 in the camera interface S4 of the mobile phone, the touch sensor 180K can detect the user's trigger operation. For example, please refer to the following text. Figure 15 In Figure (2), after the user triggers the deblurring control 20 in the photo album interface S8 of the mobile phone, the touch sensor 180K can sense the user's trigger operation. The touch sensor 180K can pass the detected touch operation to the application processor to determine the touch event type (e.g., click type, press type, or swipe type). In addition, visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K may also be disposed on the surface of the electronic device 100, in a different position than the display screen 194.

[0075] Figure 3 This is a schematic diagram of an image deblurring method provided in an embodiment of this application. The image deblurring method can be derived from the above... Figure 2 The illustrated electronic device performs this action; for example, the electronic device may be, but is not limited to, a mobile phone, tablet, or computer. Figure 3 As shown, the image deblurring method includes steps S310 to S380. Steps S310 to S380 will be described in detail below.

[0076] S310: The first electronic device acquires a blurred image to be processed, wherein the blurred image to be processed is an image obtained by the camera of the second electronic device from the object being photographed.

[0077] The blurred image to be processed can be either a locally blurred image or a globally blurred image; there is no specific limitation on this. It can be understood that when the blur in an image is localized, then that image can be called a locally blurred image, meaning that the pixels in a localized region of the image are blurred. When the blur in an image is globalized, then that image can be called a globally blurred image, meaning that the pixels in all regions of the globally blurred image are blurred.

[0078] For example, when the blurred image to be processed is a partially blurred image, the blurred image to be processed can be... Figure 4 The partial blurred image shown in (1) is a partially blurred image where the area of ​​people in the partially blurred image is blurred content, and the area of ​​people in the partially blurred image is clear (i.e. not blurred) content.

[0079] For example, when the blurred image to be processed is a globally blurred image, the blurred image to be processed can be Figure 5 The global blurred image shown in (1) is a global blurred image in which all regions are blurred content.

[0080] The first electronic device and the second electronic device can be the same electronic device or two different electronic devices; there is no specific limitation in this regard. For example, the first electronic device and the second electronic device can be the same mobile phone. Or, the first electronic device can be a computer, and the second electronic device can be a mobile phone.

[0081] The image acquisition method for the first electronic device to acquire the blurred image to be processed in S310 is not specifically limited, and can be selected according to the actual situation.

[0082] As an example of this application, the first electronic device and the second electronic device are different electronic devices. The blurred image to be processed is an image obtained by the second electronic device from the subject being photographed. The first electronic device executes S310, that is, the first electronic device acquires the blurred image to be processed, which includes: the first electronic device acquiring the second blurred image from the second electronic device. In this implementation, before the first electronic device executes S310, the second electronic device can also take a picture of the subject to obtain the blurred image to be processed.

[0083] As another example of this application, the blurred image to be processed is an image obtained by the first electronic device from the subject being photographed. The first electronic device executes S310, that is, the first electronic device acquires the blurred image to be processed, including: the first electronic device photographs the subject being photographed to obtain the blurred image to be processed.

[0084] S320: The first electronic device divides the blurred image to be processed into P image blocks including P target points, and obtains the first two-dimensional position of the i-th target point among the P target points. The P image blocks and the P target points correspond one-to-one. The i-th target point is a point in the i-th image block. P is a positive integer greater than 1, and i = 1, 2, ..., P.

[0085] The first two-dimensional position of the i-th target point out of P target points refers to its two-dimensional coordinate position within the i-th image patch of the P-th image patch. It can be understood that the magnitude of the first two-dimensional position of the i-th target point is determined in the image coordinate system. The i-th target point out of P target points, i = 1, 2, ..., P, refers to each of the P target points. For ease of description, the first two-dimensional position of the i-th target point will be denoted as... .

[0086] The number of P image patches, the number of P target points, and the shape of the i-th image patch are not specifically limited and can be set according to the actual situation. For example, P can be 2, 5, 10, or 12. For example, the shape of each image patch can be, but is not limited to, any of the following shapes: rectangle, square, or triangle, etc.

[0087] In one example, the target point of the i-th image patch is the center point of the i-th image patch. For example, consider the blurred image to be processed as follows: Figure 6 For example, please refer to the image shown. Figure 6 The blurred image to be processed is divided into P=18 image blocks by multiple dashed lines. Each image block is rectangular in shape, and the center point of each image block is the center point of that image block.

[0088] In another example, the target point of the i-th image patch can be any point other than the center point of the i-th image patch. For example, the target point of the i-th image patch can also be the lower left corner or the upper right corner of the i-th image patch.

[0089] S330: The first electronic device acquires P first three-dimensional positions of the second electronic device, and obtains the i-th second two-dimensional position based on the i-th first three-dimensional position and the first two-dimensional position of the i-th target point, so as to obtain P second two-dimensional positions, wherein the i-th first three-dimensional position is the pose parameter of the second electronic device when the camera of the second electronic device captures the i-th target point.

[0090] The i-th first three-dimensional position is the pose parameter of the second electronic device when its camera captures the i-th target point. During the process of capturing images of the subject using the camera of the second electronic device, the gyroscope sensor of the second electronic device detects and records the angular acceleration of the second electronic device at each moment (i.e., multiple moments) during image acquisition. At each moment, the second electronic device captures information about the corresponding point in the subject. Thus, the second electronic device can obtain the pose parameter (i.e., angular acceleration) of the first electronic device at any given moment from the gyroscope sensor. For ease of description, the i-th first three-dimensional position obtained by the second electronic device from the gyroscope sensor will be denoted as... .

[0091] As an example of this application, the first electronic device obtains the i-th second two-dimensional position based on the i-th first three-dimensional position and the i-th target point's first two-dimensional position, which may include the following steps: the first electronic device can determine the horizontal coordinate of the i-th first two-dimensional position. Perform the following calculations to obtain the x-coordinate of the i-th second two-dimensional position. :

[0092] (3.1)

[0093] In the above formula (3.1), Let represent the i-th rotation matrix. This refers to the internal and external parameters of the camera in the first electronic device. (Right now ) represents the i-th transformation matrix.

[0094] Furthermore, the first electronic device can determine the ordinate in the i-th first two-dimensional position. Perform the following calculations to obtain the ordinate of the i-th second two-dimensional position. :

[0095] (3.2)

[0096] In the above formula (3.2), This represents the i-th rotation matrix, for example, the rotation matrix is ​​the Rodrigues matrix. This refers to the internal and external parameters of the camera in the first electronic device. Indicates to The result of finding the inverse. (Right now ) represents the i-th transformation matrix.

[0097] The first electronic device can be calculated using the following formula to obtain the value in formula (3.1) above. :

[0098] (3.3)

[0099] In the above formula (3.3), Represents the identity matrix. for The representation of the rotation axis. Yes The vector obtained after processing. express The components of a vector along the x-axis. express The y-components of the vector. express The components of the vector along the z-axis.

[0100] The first electronic device is calculated using the following formula. :

[0101] (3.4)

[0102] In the above formula (3.4), Yes The resulting vector after processing This can be expressed by the following mathematical formula:

[0103] (3.5)

[0104] In the above formula (3.5), .

[0105] S340: The first electronic device determines the fuzzy value of the i-th target point based on the first two-dimensional position and the i-th second two-dimensional position of the i-th target point, so as to obtain P fuzzy results for P target points. The i-th fuzzy result includes the fuzzy value of the i-th target point (denoted as ). ).

[0106] Optionally, the i-th blurring result may also include the blurring direction of the i-th target point (denoted as ). ), where the blur direction of the i-th target point represents the shaking direction of the second electronic device when the camera of the second electronic device captures the i-th target point.

[0107] As an example of this application, the first electronic device obtains the blurred result of the i-th target point based on the first two-dimensional position and the i-th second two-dimensional position of the i-th target point, which can be expressed by the following formula:

[0108] (3.6)

[0109] In the above formula (3.6), This represents the fuzzy value of the i-th target point. This represents the fuzzy direction of the i-th target point. This represents the i-th second two-dimensional position. This indicates that the i-th target point is located in the first two-dimensional position in the blurred image to be processed.

[0110] Thus, after the first electronic device executes the above S340, it can obtain the blur values ​​of P target points corresponding to P image blocks in the blurred image to be processed. The blur value of each target point can measure the blur value of the image block to which each target point is located.

[0111] S350: The first electronic device determines a first degree of ambiguity and a second degree of ambiguity based on the P ambiguity values ​​of the P target points, wherein the first degree of ambiguity is the minimum degree of ambiguity among the P degrees of ambiguity, and the second degree of ambiguity is the maximum degree of ambiguity among the P degrees of ambiguity.

[0112] The first degree of ambiguity is less than the preset degree of ambiguity, and the second degree of ambiguity is greater than or equal to the preset degree of ambiguity.

[0113] The first degree of blur is the blur level of the first image block out of P image blocks, and the second degree of blur is the blur level of the second image block out of P image blocks.

[0114] For example, take the blurred image to be processed as... Figure 6 Taking the partially blurred image shown as an example, the first and second image blocks mentioned above can be found in [reference needed]. Figure 6 The first and second image blocks are shown in the figure.

[0115] As an example of this application, the electronic device performs S350, which includes: the first electronic device determining the smallest blur level among a plurality of blur levels as the second blur level; and the first electronic device determining the largest blur level among a plurality of blur levels as the second blur level.

[0116] Thus, after the first electronic device executes the above S350, it can obtain the minimum and maximum blur values ​​among the blur values ​​of the P target points corresponding to the P image blocks in the blurred image to be processed.

[0117] S360: The first electronic device determines whether the first degree of ambiguity is less than a preset threshold.

[0118] As an example of this application, after the first electronic device executes S360, the first electronic device determines that the first ambiguity is less than a preset threshold (denoted as ). In other words, when the first electronic device considers that the second electronic device experiences very little jitter when capturing the blurred image to be processed (i.e., the jitter level is less than the preset jitter level), the first electronic device will consider the captured blurred image to be processed to be a partially blurred image. Subsequently, the first electronic device executes S370, that is, it calls the local deblurring network to process the blurred image to be processed, so as to obtain a clear image after deblurring the blurred image to be processed.

[0119] As another example of this application, after the first electronic device executes S360, the first electronic device determines that the first degree of ambiguity is greater than or equal to a preset threshold. In other words, if the first electronic device believes that the second electronic device experiences significant jitter when capturing the blurred image to be processed (i.e., the jitter is greater than or equal to a preset jitter level), the first electronic device will consider the captured blurred image to be processed to be a globally blurred image. Subsequently, the first electronic device executes S380, which calls the global deblurring network to process the blurred image to be processed, so as to obtain a clear image after deblurring the blurred image to be processed.

[0120] For the preset threshold ( The size of the threshold is not specifically limited; for example, the preset threshold can be, but is not limited to, 0.8 or 1.

[0121] Thus, the first electronic device executes steps S350 and S360, namely, determining whether the blurred image to be processed is a globally blurred image or a locally blurred image. Then, if it is determined that the blurred image to be processed is a globally blurred image, the first electronic device subsequently uses a global deblurring network to perform deblurring processing on the blurred image to obtain a clear image. If it is determined that the blurred image to be processed is a locally blurred image, the first electronic device subsequently uses a local deblurring network to perform deblurring processing on the blurred image to obtain a clear image.

[0122] S370: The first electronic device calls a local deblurring network to deblur the blurred image to be processed, and obtains a clear image.

[0123] Local deblurring networks are used to deblur local blurred regions in a blurred image to obtain a clear image. No specific limitations are placed on the structure of the local deblurring network or the training method for obtaining it.

[0124] As an example of this application, a local deblurring network can be described below. Figure 12 The local deblurring network obtained by the provided model training method can be found below for the training method of the neural network model to obtain the local deblurring network. Figure 12 The relevant descriptions and the structure of the local deblurring network can be found below. Figure 9 The network shown is not described in detail here.

[0125] As another example of this application, the local deblurring network can be a network already proposed in related technologies, and the training method for obtaining the local deblurring network can be a training method already proposed in related technologies. For example, the training method includes training the encoder-decoder using multiple training samples to obtain the local deblurring network, wherein each training sample includes a blurred image and a label image (i.e., a sharp image corresponding to the blurred image). The sharp image in S370 above is the image obtained after performing local deblurring processing on the blurred image to be processed.

[0126] For example, take the blurred image to be processed in S370 as... Figure 4 Taking the partially blurred image shown in Figure (1) as an example, the clear image obtained after the first electronic device executes S370 can be found in [reference 1]. Figure 4 The clear image shown in (2) of the diagram.

[0127] In this way, the first electronic device can use a local deblurring network that matches the local blurring type of the blurry image to be processed to process the blurry image, which can improve the deblurring effect of the blurry image to be processed, so that the clear image obtained can better meet the user's needs.

[0128] S380: The first electronic device calls the global deblurring network to deblur the blurred image to be processed, and obtains a clear image.

[0129] Global deblurring networks are used to deblur globally blurred regions in a blurred image to obtain a sharp image. No specific limitations are placed on the structure of the global deblurring network or the training method for obtaining it.

[0130] As an example of this application, a global deblurring network can be described below. Figure 12 The global deblurring network obtained by the provided model training method can be found below for the training method of the neural network model. Figure 12 The relevant descriptions and the structure of the global deblurring network can be found below. Figure 9 The network shown is not described in detail here.

[0131] As another example of this application, the global deblurring network can be a network already proposed in the related art, and the training method to obtain the global deblurring network can be a training method already proposed in the related art. For example, the training method includes training the encoder-decoder using multiple training pairs to obtain the global deblurring network, wherein each training data pair includes a blurred image and a label image (i.e., a sharp image corresponding to the blurred image).

[0132] The clear image in S380 above is the deblurred image obtained after performing global deblurring on the blurry image to be processed.

[0133] For example, take the blurred image to be processed in S380 as... Figure 5 Taking the globally blurred image shown in Figure (1) as an example, the clear image obtained after the first electronic device executes S380 can be found in [reference 1]. Figure 5 The clear image shown in (2) of the diagram.

[0134] In this way, the first electronic device can use a global deblurring network that matches the global blur type of the blurred image to be processed to process the blurred image, which can improve the deblurring effect of the blurred image to be processed, so that the clear image obtained can better meet the user's needs.

[0135] It should be understood that the above Figure 3 The image deblurring method shown is for illustrative purposes only and does not constitute any limitation on the image deblurring method provided in the embodiments of this application. For example, the degree of shaking of the electronic device during the time period in which the camera of the electronic device captures the blurred image to be processed can also be determined based on other existing methods.

[0136] In this embodiment, after acquiring a blurred image to be processed, the electronic device first determines a deblurring strategy (i.e., a target deblurring strategy) based on the jitter level. Since the jitter level represents the degree of shaking of the electronic device during the time period when the camera acquires the blurred image, the target deblurring strategy determined based on the jitter level matches the degree of blurriness of the blurred image. Then, the electronic device performs deblurring on the blurred image according to the target deblurring strategy determined based on the jitter level to obtain a clear image. This method avoids the problem in traditional technologies where the deblurring strategy and the degree of blurriness of the acquired blurred image are mismatched, resulting in poor deblurring performance. Therefore, this method can improve the image deblurring effect, thereby better satisfying the user's visual experience.

[0137] As mentioned above, the image deblurring method provided in this application involves a global deblurring network and a local deblurring network. Therefore, this application also provides a training method for training a neural network model to obtain the global and local deblurring networks. Before introducing the training method provided in this application, the system architecture and neural network processor hardware structure to which the training method provided in this application is applicable will be described below with reference to the accompanying drawings.

[0138] See Figure 7 , Figure 7 This is a schematic diagram of a system 700 to which the training method provided in the embodiments of this application applies. It should be understood that the following description should not be construed as limiting any examples of this disclosure. Figure 7In the illustrated system 700, labeled training data can be stored in a database 730. Database 730 can be located on a server or in a data center, or it can be provided as a service by a cloud computing service provider. In the context of this disclosure, labeled training data refers to the training data used to learn the training weights of a deblurring neural network 701 (also referred to as deblurring network 701 for simplicity). In one example, labeled training data includes training images, label images, and motion mask images, where the training images are blurred images, the label images are sharp images corresponding to the training images, and the motion mask images are images that include features of the blurred regions of the training images. Labeled training data differs from application-time training data, which can be unlabeled real-world data (e.g., real-world images captured by application device 710, discussed below) or unlabeled test data. Application-time data includes unlabeled images and motion mask images. As will be discussed further below, the application-time training of the deblurring network 701 can be performed using a single input real-world image to obtain the application-time training weights of the deblurring network 701. The deblurring network 701, including the application-time training weights, can be used to predict the corresponding single deblurred output image based on a single input real-world image.

[0139] Database 730 may contain, for example, previously collected labeled training data that is typically used to train models related to image tasks (e.g., image recognition). The input images for the labeled training data stored in database 730 may, or additionally, be images optionally collected from application device 710 (which may be a user device) (e.g., with user consent). For example, images captured by the camera of application device 710 and stored on application device 710 may be optionally anonymized and uploaded to database 730 for storage as input images for labeled training data. The labeled training data stored in database 730 may include the training images, labeled images, and motion mask images described above.

[0140] As will be discussed further below, training device 720 can be used to train deblurring network 701 based on training data stored in database 730. Alternatively, training device 720 can use training data obtained from other sources (e.g., distributed storage (or cloud storage platform)) to train deblurring network 701. The trained deblurring network 701 (i.e., the result of training by training device 720) has a set of training weights. According to the examples disclosed herein, application device 710 can further train the trained deblurring network 701 to deblur specific blurred real-world images. Application-time training of application device 710 can be performed using images (e.g., digital photographs) captured by the camera (not shown) of application device 710. Application device 710 may not have access to the training data stored in database 730.

[0141] In the examples disclosed herein, training the deblurring network 701 can be implemented in the processing unit 711 of the application device 710. For example, the deblurring network 701 can be encoded and then stored as instructions in the memory (not shown) of the application device 710, and the processing unit 711 executes the stored instructions to implement the deblurring network 701. In some examples, the deblurring network 701 can be encoded and then stored as instructions in the memory of the processing unit 711 (e.g., the weights of the deblurring network 701 can be stored in a corresponding weight memory of the processing unit 711, which can be embodied as follows). Figure 8 The neural network processor 800 shown is illustrated. In some examples, the deblurring network 701 can be implemented in the integrated circuit of the application device 710 (as software and / or hardware). Although Figure 7 An example is shown where the training device 720 and the application device 710 are separate. It should be understood that this disclosure is not limited to this embodiment. In some examples, separate training device 720 and application device 710 may not exist. That is, the training of the deblurring network 701 and the application training of the deblurring network 701 can be performed on the same device (e.g., application device 710).

[0142] Application device 710 can be a user device, such as a client terminal, mobile terminal, tablet computer, laptop computer, augmented reality (AR) device, virtual reality (VR) device, or in-vehicle terminal, etc. Application device 710 can also be a server, cloud computing platform, etc., which users can access through their user devices. Figure 7 In this application device 710, an I / O interface 712 is included for data interaction with external devices. For example, the application device 710 can provide uploaded data (e.g., image data, such as photos and / or videos captured by the application device 710) to the database 730 via the I / O interface 712. Although Figure 7 An example of direct interaction between a user and application device 710 is shown. It should be understood that this disclosure is not limited to this embodiment. In some examples, the user device may be separate from the application device 710, with the user interacting with the user device, and the user device instead exchanging data with the application device 710 through I / O interface 712.

[0143] In this example, application device 710 includes a data storage device 714, which may be system memory (e.g., random access memory (RAM), read-only memory (ROM), etc.) or mass storage device (e.g., solid-state drives and hard disk drives, etc.). Data storage device 714 can store data accessible to processing unit 711. For example, data storage device 714 may be separate from processing unit 711, storing captured and / or repaired images on application device 710.

[0144] In some examples, the application device 710 may optionally call data and code from an external data storage system 750 for processing, or may store data and instructions obtained through the corresponding processing in the data storage system 750.

[0145] It is important to note that Figure 7 This is merely a schematic diagram of an example system architecture 700 according to an embodiment of this disclosure. Figure 7 The relationships and interactions between the devices, components, and processing units shown are not intended to limit this disclosure.

[0146] Figure 8 This is a block diagram of the hardware structure of the neural network processor provided in the embodiments of this application.

[0147] The neural network processor 800 can be mounted on an integrated circuit (also known as a computer chip). Figure 7 In the application device 710 shown, the processing unit 711 performs calculations and implements the deblurring network 701 (including training during the application of the deblurring network). Alternatively, the neural network processor 800 may be configured... Figure 7 The training device 720 shown is used to train the deblurring network 701. All algorithms of the layers in the neural network (e.g., the layers of the deblurring network 701, which are discussed further below) can be implemented in the neural network processor 800.

[0148] The neural network processor 800 can be any processor capable of performing the computations required in a neural network (e.g., computations involving numerous XOR operations). For example, the neural network processor 800 can be a neural processing unit (NPU), a tensor processing unit (TPU), or a graphics processing unit (GPU). The neural network processor 800 can be a coprocessor of an optional host central processing unit (CPU) 820. For example, the neural network processor 800 and the host CPU 820 can be mounted on the same package. The host CPU 820 can be responsible for performing the core functions of the application device 710 (e.g., execution of the operating system (OS), management communication, etc.). The host CPU 820 can manage the operation of the neural network processor 800, for example, by assigning tasks to the neural network processor 800.

[0149] The neural network processor 800 includes an arithmetic circuit 803. The controller 804 of the neural network processor 800 controls the arithmetic circuit 803 to retrieve data (e.g., matrix data) from the input memory 801 and weight memory 802 of the neural network processor 800, and to perform data operations (e.g., addition and multiplication operations).

[0150] In some examples, the arithmetic circuit 803 internally includes multiple processing units (also called process engines, PEs). In some examples, the arithmetic circuit 803 is a two-dimensional pulsating array. In other examples, the arithmetic circuit 803 can be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some examples, the arithmetic circuit 803 is a general-purpose matrix processor.

[0151] In one example operation, the arithmetic circuit 803 retrieves the weight data of the weight matrix B from the weight memory 802 and caches the weight data in each PE of the arithmetic circuit 803. The arithmetic circuit 803 retrieves the input data of the input matrix A from the input memory 801 and performs matrix operations based on the input data of matrix A and the weight data of matrix B. The resulting partial or final matrix result is stored in the accumulator 808 of the neural network processor 800.

[0152] In this example, the neural network processor 800 includes a vector computation unit 807. The vector computation unit 807 includes multiple computation processing units. If needed, the vector computation unit 807 further processes the output from the computation circuit 803 (which can be retrieved from the accumulator 808 by the vector computation unit 807), such as vector multiplication, vector addition, exponentiation, logarithmic operations, or magnitude comparisons. The vector computation unit 807 can be primarily used for operations in non-convolutional or fully connected layers of the neural network. For example, the vector computation unit 807 can perform processing such as pooling or normalization on the operations. The vector computation unit 807 can apply a nonlinear function, such as a vector of accumulated values, to the output of the computation circuit 803 to generate activation values. These activation values ​​can be used by the computation circuit 803 as activation inputs for the next layer of the neural network. In some examples, the vector computation unit 807 generates normalized values, combined values, or a combination of normalized and combined values.

[0153] In this example, the neural network processor 800 includes a memory access controller 805 (also known as direct memory access control, DMAC). The memory access controller 805 is used to access external memory (e.g., data memory 714 of application device 710) via a bus interface unit 810. The memory access controller 805 can access data from external memory and directly transfer data to one or more memories of the neural network processor 800. For example, the memory access controller 805 can directly transfer weight data to a weight memory 802, or directly transfer input data to a unified memory 806 and / or an input memory 801. The unified memory 806 is used to store input and output data (e.g., processing vectors from a vector computation unit 807).

[0154] The bus interface unit 810 is also used for interaction between the memory access controller 805 and the instruction fetch memory (also known as the instruction fetch cache) 809. The bus interface unit 810 is also used to enable the instruction fetch memory 809 to fetch instructions from memory external to the neural network processor 800 (e.g., the data memory 714 of the application device 710). The instruction fetch memory 809 is used to store instructions for use by the controller 804.

[0155] Typically, the unified memory 806, input memory 801, weight memory 802, and instruction fetch memory 809 are all part of the neural network processor 800's memory (also known as on-chip memory). The data memory 714 is independent of the neural network processor 800's hardware architecture.

[0156] The image deblurring method provided in this application involves a first electronic device calling a deblurring network model (e.g., a global deblurring network or a local deblurring network) to deblur a blurred image. The deblurring network model provided in this application will be described in detail below with reference to the accompanying drawings. It should be noted that the training and application of the deblurring network model will be further described below.

[0157] Figure 9 This is a schematic diagram of the architecture of a deblurring network model provided in an embodiment of this application. For simplicity, the neural network layers (or blocks) of the deblurring network will be simply referred to as layers in the following text. Figure 9 As shown, the deblurring network model is based on a fully convolutional neural network. This deblurring network model includes an encoder, a decoder, and skip connections between layers of the same size (i.e., ...). Figure 9 (as shown by the dashed arrow in the diagram) and the connection portion, wherein the connection portion includes multiple convolutional layers for connecting the encoder and decoder together.

[0158] In this embodiment of the application, the input to the deblurring network model includes a training dataset, and the output of the deblurring network model includes the predicted deblurred image (abbreviated as...). The predicted deblurred image is also known as the sharp image.

[0159] The training dataset consists of N training samples, where each training sample includes a blurred image (abbreviated as N). ), labeled images and motion mask images (abbreviated as The training dataset consists of N blurred images, N labeled images, and N motion mask images, where N is a positive integer.

[0160] The above N blurred images correspond one-to-one with N label images, where each label image represents the sharp image (i.e., the unblurred image) corresponding to the blurred image. In one example, the blurred images in the training dataset can be represented as a two-dimensional matrix that encodes multiple channels (e.g., red-green-blue, RGB channels) (multiple individual pixels of the input image).

[0161] For example, when performing local blur training to obtain a local deblurring network, the training images can be... Figure 10 The partially blurred image shown in (2) of the figure can be a label image. Figure 10 The clear image shown in (1) is as follows. For example, when performing global blur training to obtain a global deblurring network, the training image can be... Figure 11 The globally blurred image shown in (2) can be a label image. Figure 11The clear image shown in (1) is shown in the figure.

[0162] The above N motion mask images correspond one-to-one with N blurred images, and each motion mask image ( ) including the corresponding blurred image ( The region where the moving object is located in each motion mask image. In other words, each motion mask image ( The region of interest (ROI) in the blurred image includes the blurred image ( The region where the moving object is located in the image, and each motion mask image ( The remaining area included is the background area. In one example, each motion mask image ( The corresponding blurred image included The pixel value of the region containing the moving object (i.e., the region of interest) in the motion mask image can be set to "1", meaning that the region of interest will appear white in the motion mask image. Each motion mask image ( The pixel values ​​of the remaining area (i.e., the background area) of the motion mask image can be set to "0", meaning that the background area is within the motion mask image. It appears as black in the image.

[0163] For example, with Figure 10 (2) In the figure, the locally blurred image shown is a blurred image in the training data. The moving object in the locally blurred image is a person. Therefore, the motion mask image corresponding to the locally blurred image can be... Figure 10 Figure (3) shows a motion mask image, which includes white areas with a pixel value of "1". Figure 10 (2) shows the region where the moving object (i.e., the person) is located in the blurred image, and the motion mask image includes a black area with a pixel value of "0", which corresponds to Figure 10 (2) The blurred region shown in the figure excludes the area where the moving object is located. For example, using... Figure 11 (2) In the figure, the globally blurred image shown is taken as a blurred image in the training data. The moving objects in this locally blurred image include people and buildings. The motion mask image corresponding to this globally blurred image can be... Figure 11 Figure (3) shows a motion mask image, which includes white areas with a pixel value of "1". Figure 11 (2) shows the area where the moving object is located in the blurred image (i.e., the area of ​​people and the area of ​​buildings), and the black area with a pixel value of "0" corresponding to it. Figure 11 (2) The blurred area shown in the figure is the area excluding the area where the moving object is located.

[0164] As mentioned earlier, the input to the deblurring network model includes a training dataset. The deblurring network model processes each of the N training samples in the training dataset in the same way. Therefore, the following description will use the training of the deblurring network model with a single training sample as an example.

[0165] like Figure 9 As shown, the encoder of the deblurring network model includes an input layer 910, convolutional layers 920 (i.e., convolutional layers 920a1 and 920a2), and downsampling layers 930 (i.e., downsampling layers 930a and 930b). Each downsampling layer is preceded by a convolutional layer; for example, downsampling layer 930b is preceded by convolutional layer 920a2, and downsampling layer 930a is preceded by convolutional layer 920a1. The kernel size of convolutional layer 920a1 is larger than that of convolutional layer 920a2, and the kernel size of downsampling layer 930a is larger than that of downsampling layer 930b. For example, the kernel size of convolutional layer 920a1 is... The kernel size of convolutional layer 920a2 is .

[0166] The encoder is used to process the blurred image input to the deblurring network model. Feature encoding is performed to obtain a blurred image. The feature map of ). Specifically, the input layer 910 is used to receive the blurred image () input to the deblurring network model. ), labeled images and motion mask images ( ), and output a blurred image ( ), labeled images and motion mask images ( Convolutional layer 920a1 obtains the blurred image from input layer 910. ), labeled images and motion mask images ( ), and for blurred images ( Perform a convolution operation to obtain a blurred image. The features encoded by the feature encoding Figure 1 (i.e., feature representation). Features output by convolutional layer 920a1 Figure 1 Labeled images and motion mask images ( This is used as the input to the downsampling layer 930a. The downsampling layer 930a processes the features... Figure 1 Perform downsampling (also known as pooling) to obtain features. Figure 2 (Feature representation), then, the downsampling layer 930a outputs features. Figure 2 Labeled images and motion mask images ( Next, convolutional layer 920a2 processes the features. Figure 2Features are obtained by performing convolution operations. Figure 3 (i.e., feature representation), then the convolutional layer 920a2 outputs the features. Figure 3 Labeled images and motion mask images ( The downsampling layer 930b is used for features. Figure 3 Features are obtained by performing downsampling processing. Figure 4 (i.e., feature representation), then the downsampling layer 930b outputs features. Figure 4 Labeled images and motion mask images ( Features output from downsampling layer 930b Figure 4 Labeled images and motion mask images ( The data is processed by multiple convolutional layers in the connection part and then used as input data for the decoder.

[0167] like Figure 9 As shown, the decoder of the deblurring network model includes deblurring feature reconstruction layers 940 (i.e., deblurring feature reconstruction layers 940a and 940b), upsampling layers 950 (i.e., upsampling layers 950a and 950b), convolutional layers 920b (i.e., convolutional layers 920b1 and 920b2), and an output layer 960. Each upsampling layer is preceded by a deblurring feature reconstruction layer. The size of the convolutional kernel in deblurring feature reconstruction layer 940a is smaller than that in deblurring feature reconstruction layer 940b. The size of the convolutional kernel in upsampling layer 950a is smaller than that in upsampling layer 950b.

[0168] Deblurring feature reconstruction layer 940 is used to reconstruct features based on the motion mask image ( ) for the acquired blurred image ( The corresponding feature map is optimized to generate an optimized feature map (i.e., feature representation). Below, taking the defuzzification feature reconstruction layer 940a as an example, the function of the defuzzification feature reconstruction layer 940 provided in this application embodiment is explained. Figure 9 As shown, the size of the input feature map (i.e., feature representation) of the defuzzification feature reconstruction layer 940a is... Where h represents the length of the feature map, w represents the width of the feature map, and n represents the number of layers in the feature map. In the deblurring feature reconstruction layer 940a, it is used to determine the feature map based on the motion mask image ( The difference between the input feature image and the deblurred feature reconstruction layer 940a is used to adjust the parameters of the deblurred feature reconstruction layer 940a so that the feature image output by the adjusted deblurred feature reconstruction layer 940a matches the motion mask image. The difference between the two images is less than the preset difference. For example, with a blurred image ( Taking a locally blurred image as an example, this motion mask image ( Please see Figure 9The motion mask image in the image, and the feature map corresponding to a feature layer obtained by the deblurred feature reconstruction layer 940a are shown below. Figure 9 The image shown is a layer 1 feature image.

[0169] Skip connections (also known as short connections) can supplement the same-scale features extracted in the downsampling layer 930 in the upsampling layer 950, recovering features lost during downsampling. Skip connections can be used in residual neural networks to facilitate faster learning of weights for deblurring networks. For example, Figure 9 One skip connection is shown to supplement the same-scale features extracted in downsampling layer 930a in upsampling layer 950b, and another skip connection supplements the same-scale features extracted in downsampling layer 930b in upsampling layer 950a.

[0170] Upsampling layer 950 concatenates the feature map output from the corresponding downsampling layer with the feature map output from the previous layer (i.e., merging deep and shallow features to enrich the information). Then, it decodes the concatenated feature map to generate an output feature map (i.e., feature representation). For example, taking upsampling layer 950a as an example, upsampling layer 950a concatenates the feature map (i.e., feature representation) output from deblurred feature reconstruction layer 940a and the feature map (i.e., feature representation) output from downsampling layer 930b. Then, it decodes the concatenated feature map (i.e., feature representation) to generate an output feature map (i.e., feature representation). For example, taking upsampling layer 950b as an example, upsampling layer 950b concatenates the feature map output from deblurred feature reconstruction layer 940b and the feature map (i.e., feature representation) output from downsampling layer 930a. Then, it decodes the concatenated feature map (i.e., feature representation) to generate an output feature map (i.e., feature representation).

[0171] The convolutional layer 920 connected after the upsampling layer 950 is used to perform convolution operations on the feature map (i.e. feature representation) output by the upsampling layer 950b.

[0172] The output layer 960 performs a convolution operation on the feature map (i.e., feature representation) output by the convolutional layer 920b connected to it to generate the predicted deblurred image. ). Subsequently, based on the predicted deblurred image ( The difference between the image and the label image is used to adjust the parameters of the encoder, the adjusted deblurred feature reconstruction layer 940, the upsampling layer 950, and the output layer 960 until the preset training conditions are met, and then the model training is stopped, thus obtaining the trained deblurred network.

[0173] It should be understood that the above Figure 9The architecture of the deblurring network shown is for illustrative purposes only and does not constitute any limitation on the architecture of the deblurring network applicable to the image deblurring method provided in this application. That is to say, Figure 9 The architecture of the deblurred network shown can be modified (e.g., with fewer or more neural network layers). For example, the convolutional layer 920 in the encoder can also include more convolutional layers, the downsampling layer 930 can also include more downsampling layers, and the upsampling layer 950 can also include more downsampling layers.

[0174] Below, based on the above Figure 9 Taking the architecture of the deblurring network shown as an example, combined with Figure 12 This application introduces a training method for obtaining a global deblurring network from the deblurring network, and a training method for obtaining a local deblurring network from the deblurring network, both provided in the embodiments of this application. It should be noted that the model structures of the global and local deblurring networks in the embodiments of this application are the same; the difference lies in the model parameters of the global and local deblurring networks, and the training datasets for obtaining the global and local deblurring networks are different.

[0175] Figure 12 This is a schematic diagram illustrating a model training method provided in an embodiment of this application. It should be understood that... Figure 12 The model training method shown is for illustrative purposes only and does not constitute any limitation on the model training method provided in this application.

[0176] Figure 12 This is a schematic diagram of a model training method provided in an embodiment of this application. The model training method can be derived from the above... Figure 7 The training device 720 shown is used for execution. For example... Figure 12 As shown, the model training method includes S1210 to S1250. Below, S1210 to S1250 will be described in detail.

[0177] S1210: Obtain the training dataset, which includes N training samples. The i-th training sample includes the i-th blurred image, the i-th motion mask image, and the i-th label image. The i-th motion mask image includes the blurred region of the i-th training image. The i-th label image is the clear image corresponding to the i-th blurred image. i=1,2,…,N, where N is a positive integer.

[0178] The blurred images in any two training samples out of N training samples can be the same or different.

[0179] The i-th motion mask image includes the blurred region of the i-th training image, also known as the feature that the i-th motion mask image includes the blurred region of the i-th training image. In the embodiments of this application, the i-th motion mask image includes the region where the moving object (which appears as blurred content in the i-th blurred image) is located in the i-th blurred image. It can be understood that the moving object moves during the time period of acquiring the i-th blurred image, causing the corresponding region of the i-th blurred image to appear blurred. In one example, the pixel value of the region where the moving object is located in the i-th blurred image included in the i-th motion mask image can be set to "1", that is, the region appears white in the i-th motion mask image, and the pixel value of the remaining region of the i-th motion mask image can be set to "0", that is, the region appears black in the i-th motion mask image.

[0180] As an example of this application, in the context of... Figure 9 The deblurring network shown performs local deblurring training to obtain a scene in which the i-th training sample in the N training samples includes the i-th blurred image, the i-th motion mask image includes the blurred region of the local blurred image (i.e. the region of the moving object), and the i-th label image is the clear image corresponding to the local blurred image.

[0181] For example, the training images in one training sample of the training dataset can be Figure 10 The partially blurred image shown in (2) is characterized by a blurred image of the person (i.e., the motion region). The motion-blurred image in one of the training samples can be... Figure 10 Figure (3) shows a motion mask image where the pixel value of the person area (i.e., the blurred area) is "1", and the pixel value of the remaining area (i.e., the unblurred area) is "0". The label image in a training sample can be... Figure 10 The clear image shown in (1) is shown in the figure.

[0182] As an example of this application, in the context of... Figure 9 The deblurring network shown performs global deblurring training to obtain a scene in which the i-th training sample in the N training samples includes the i-th blurred image, which is the global blurred image, the i-th motion mask image includes the blurred region of the global blurred image (i.e. the region of the moving object), and the i-th label image is the clear image corresponding to the global blurred image.

[0183] For example, the training images in one training sample of the training dataset can be Figure 11Figure (2) shows a globally blurred image where the human figure region (i.e., the motion region) is blurred. The motion-blurred image in one of the training samples can be... Figure 11 Figure (3) shows a motion mask image where all regions (i.e., blurred regions) have a pixel value of "1". The label image in a training sample can be... Figure 11 The clear image shown in (1) is shown in the figure.

[0184] There are no specific limitations on the method for obtaining N training samples; it can be set according to the actual scenario.

[0185] As an example of this application, a high frame rate camera can be used to capture video, and blurry and clear images in consecutive frames of the video can be found as a set of data (i.e., the i-th training image and the i-th label image in the i-th training sample).

[0186] For example, the following steps can be used to obtain the i-th training image and the i-th label image in the i-th training sample: obtain N video streams, where the i-th video stream includes multiple consecutive frames; perform weighted averaging on the multiple consecutive frames included in the i-th video stream to obtain the i-th training image; determine one clear frame from the multiple consecutive frames included in the i-th video stream as the i-th label image to obtain the i-th label image.

[0187] As another example of this application, the i-th clear image is blurred using a known or randomly generated motion blur kernel to generate a corresponding set of data (i.e., the i-th training image and the i-th label image in the i-th training sample).

[0188] Thus, by performing the above processing on the i-th video stream among the N video streams, we can obtain N training images and N label images from the N training samples.

[0189] As an example of this application, the i-th motion mask image in the i-th training sample can be obtained through the following steps: including: using a background subtraction algorithm to obtain the moving foreground region in the i-th training image; and using a connected component extraction algorithm to extract the mask region from the moving foreground region to obtain the i-th motion mask image. Optionally, after obtaining the i-th motion mask image, an image filter can be used to smooth the i-th motion mask image, or a morphological filtering operation can be used to filter the i-th motion mask image to remove smaller noise regions in the i-th motion mask image.

[0190] Thus, by performing the above processing on the i-th blurred image in the N training samples, N motion mask images in the N training samples can be obtained.

[0191] S1220: Input N training samples from the training dataset into the encoder. The encoder performs blur feature extraction on the N training images corresponding to the N training samples to obtain N feature images #1 corresponding to the N training images. The initial deblurring network includes the encoder.

[0192] As an example of this application, the structure of the initial deblurring network is shown in [reference]. Figure 9 The network shown above, including the structure of the initial deblurring network and the steps of the encoder extracting features from the i-th training image out of N training images, are described in the relevant sections above and will not be repeated here. It is understood that the encoder is... Figure 9 The encoder shown in the image.

[0193] S1230: Input the N feature images #1 and N motion mask images output by the encoder into the first network module of the decoder. Adjust the parameters of the first network module according to the loss value #1 so that the adjusted first network module outputs N feature images #2. The loss value #1 is determined based on the difference between the i-th feature image #1 and the i-th motion mask image. The initial deblurring network also includes the decoder.

[0194] In S1230 above, the parameters of the first network module are adjusted using the loss value #1, that is, the weights (i.e., parameters) of the first network module are adjusted using the loss value #1. As an example of this application, the structure of the initial deblurring network is shown below. Figure 9 The network shown has a decoder that is Figure 9 The decoder shown in the figure includes a first network module. Figure 9 The two unfuzzy feature reconstruction layers 940 shown in the figure can be found in the relevant description above for details not elaborated here.

[0195] The loss value #1 is based on the difference between the i-th feature image #1 and the i-th motion mask image. The loss value #1 can be obtained by calculating the difference between the i-th feature image #1 and the i-th motion mask image using a loss function.

[0196] During the training of a neural network model, because the output of the model should be as close as possible to the desired predicted value, the weight vector of each layer can be updated by comparing the current prediction with the target value and considering the difference between the two (of course, there is usually an initialization process before the first update, i.e., pre-configuring parameters for each layer in the neural network model). For example, if the model's prediction is too high, the weight vector is adjusted to predict a lower value, and this adjustment continues until the neural network model can predict the target value or a value very close to it. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training the neural network model becomes a process of minimizing this loss. In this embodiment, the loss function is not specifically limited and can be selected according to user needs. For example, the loss function can be, but is not limited to, the mean squared error (MSE) loss function or the mean absolute error (MAE) loss function.

[0197] For example, taking the loss value #1 obtained by calculating using the MSE loss function, this loss value #1 can be expressed by the following formula:

[0198] (12.1)

[0199] In the above formula (12.1), This represents the loss value #1. This represents the i-th training image. This represents the i-th predicted image. This represents the difference between the i-th training image and the i-th prediction image.

[0200] In step S1230 above, a loss value #1 can be calculated based on N feature images #1 and N motion mask images. Then, the parameters of the first network module in the decoder are adjusted based on the loss value #1 to minimize the loss value #1 corresponding to the N feature images #1, thereby better meeting the requirements for image deblurring.

[0201] Thus, in the decoder, the parameters of the first network module are adjusted using the differences between the N motion mask images corresponding to the N training images from the N training samples and the N feature images #1 corresponding to the N training images obtained by the encoder. Afterwards, the adjusted first network module can output N feature images #2 obtained by optimizing the N feature images #1; that is, the N feature images #2 output by the adjusted first network module can more accurately represent the blurred features in the corresponding training images.

[0202] S1240: The second network module of the decoder processes the N feature images #2 and outputs N predicted label images.

[0203] As an example of this application, the structure of the initial deblurring network is shown in [reference]. Figure 9 The network shown includes a second network module. Figure 9 The two upsampling layers 950 shown here are not detailed here; please refer to the relevant descriptions above.

[0204] S1250: Adjust the parameters of the encoder, the adjusted parameters of the first network module, and the parameters of the second network module according to the loss value #2 to obtain the trained deblurring network, wherein the loss value #2 is determined based on the difference between the i-th training image and the i-th label image.

[0205] In the above S1250, the parameters of the encoder, the adjusted parameters of the first network module, and the parameters of the second network module are adjusted using the loss value #2, that is, the weights of the encoder, the adjusted weights of the first network module, and the weights of the second network module are adjusted using the loss value #2.

[0206] For example, loss value #2 can be, but is not limited to, a loss value calculated by the electronic device based on the difference between the i-th training image and the i-th label image using any of the following loss functions: MSE loss function or MAE loss function.

[0207] As an example of this application, if the i-th training image among the N training samples in the above S1210 is a locally blurred image, the trained deblurring network obtained after executing the above S1250 is called a locally deblurring network.

[0208] As another example of this application, if the i-th training image among the N training samples in the above S1210 is a globally blurred image, the trained deblurring network obtained after executing the above S1250 is called a global deblurring network.

[0209] In one example, the parameters of the encoder, the adjusted parameters of the first network module, and the parameters of the second network module are adjusted according to the loss value #2 to obtain a trained deblurred network. This includes: adjusting the parameters of the encoder, the adjusted parameters of the first network module, and the parameters of the second network module according to the loss value #2 until a preset training condition is met, at which point the model training ends, resulting in a trained deblurred network. The iterative termination condition for training the initial deblurred network can be, but is not limited to, at least one of the following conditions: the loss value of the loss function is less than a preset threshold, the current number of iterations meets the preset number of iterations requirement, or the current model training time meets the preset training duration requirement.

[0210] Thus, adjusting the encoder parameters, the adjusted parameters of the first network module, and the parameters of the second network module based on the loss value #2 determined from the N blurred images and the N predicted clear images is to minimize the loss value #2 corresponding to the N blurred images, so as to better meet the requirements of image deblurring.

[0211] It should be understood that the above Figure 12 The training method shown is for illustrative purposes only and does not constitute any limitation on the training method provided in this application. For example, the trained image segmentation model can also be used to process N training images to obtain N motion mask images corresponding to the N training images.

[0212] In this embodiment, the training samples used by the electronic device to train the initial deblurring network include not only the training image and the corresponding label image, but also the motion mask image corresponding to the training image. Since the blur types of the training samples (e.g., global blur or local blur) are different, the corresponding motion mask images are also different. Thus, when training the initial deblurring network based on the corresponding training samples, the network can be guided to learn the corresponding type of blur processing capability, resulting in a globally deblurring network with strong global deblurring capability and a locally deblurring network with strong local deblurring capability. Unlike traditional methods that only input the training image and label image into the initial network model, this image deblurring method utilizes a target deblurring network trained on the initial deblurring network based on the training samples (i.e., the training image, label image, and motion mask image). This improves the deblurring accuracy of the trained target deblurring network. Subsequently, deblurring processing is performed on the image to be processed based on this trained target deblurring network, improving the image deblurring effect and better satisfying the user's visual experience.

[0213] Figure 13This is a schematic diagram of an image deblurring method provided in an embodiment of this application. The image deblurring method provided in this embodiment can be executed by an electronic device. It is understood that the electronic device can be implemented as software, or a combination of software and hardware. For example, the electronic device in the embodiments of this application can be, but is not limited to, the one described above. Figure 2 The electronic device 100 is shown. (For example...) Figure 13 As shown, the image deblurring method provided in this application includes steps S1310 to S1340. Steps S1310 to S1340 will be described below.

[0214] S1310: Electronic device acquires image to be processed.

[0215] The image to be processed is an image captured by the camera of an electronic device onto a subject; there is no specific limitation on the image to be processed. For example, the image to be processed could be... Figure 4 The partially blurred image shown in (1) is an example. For instance, the image to be processed could be... Figure 5 The global blurred image shown in (1) is shown in the figure.

[0216] In the embodiments of this application, the method for acquiring the image to be processed by the electronic device is not specifically limited.

[0217] As an example of this application, an electronic device acquires an image to be processed, including: the electronic device taking a picture of a subject using a camera to obtain the image to be processed. In this implementation, the electronic device that takes the picture of the subject to obtain the image to be processed is the same electronic device as the electronic device that performs the image deblurring method provided in the embodiments of this application.

[0218] For example, the image to be processed in the above implementation is the one described above. Figure 3 The provided image deblurring method is used to process the blurred image.

[0219] S1320: The electronic device acquires the amount of jitter of the electronic device, wherein the amount of jitter is used to represent the degree of jitter of the electronic device's camera during the time period of acquiring the image to be processed.

[0220] As an example of this application, the electronic device obtains the jitter amount of the electronic device by: the electronic device acquiring angular acceleration data, wherein the angular acceleration data is used to represent the pose data of the electronic device during the time period when the camera acquires the image to be processed; and the electronic device determining the jitter amount based on the angular acceleration data.

[0221] In one example, a gyroscope sensor located in an electronic device can detect and record the pose data of the electronic device at different times. In this case, the aforementioned angular acceleration data can be obtained by the electronic device from the gyroscope sensor located in the electronic device.

[0222] In the embodiments of this application, the implementation method of determining the jitter amount of electronic devices based on angular acceleration data is not specifically limited.

[0223] In one example, the angular acceleration data in the above implementation includes multiple angular accelerations, wherein the multiple angular accelerations correspond to multiple target points in the image to be processed, each angular acceleration is the pose data of the electronic device when the camera captures the corresponding target point, the multiple target points correspond to multiple first two-dimensional positions, the position of each target point in the image to be processed is the corresponding first two-dimensional position, and the electronic device determines the jitter amount based on the angular acceleration data, including: the electronic device obtains multiple second two-dimensional positions based on the multiple angular accelerations and multiple first two-dimensional positions; the electronic device obtains multiple blur values ​​corresponding to the multiple target points based on the multiple second two-dimensional positions and multiple first two-dimensional positions; and the first electronic device determines the jitter amount based on the multiple blur values.

[0224] In one example, multiple target points correspond one-to-one with multiple image patches in the image to be processed, with each target point being a point within its corresponding image patch. For example, each target point could be the center point of its corresponding image patch. Or, for example, each target point could be the bottom-left corner point of its corresponding image patch.

[0225] For example, the image to be processed can be Figure 6 The image shown is partially blurred and divided into P=18 image blocks by multiple dashed lines. Each image block is rectangular in shape and has its center point as the center point of that image block.

[0226] The step of determining the jitter amount based on multiple fuzzy values ​​in the above-mentioned electronic device may, for example, include the following steps: if the smallest fuzzy value among the multiple fuzzy values ​​is less than a preset threshold, determine that the jitter amount meets a first jitter level, wherein the first jitter level is less than a preset jitter level; if the smallest fuzzy value among the multiple fuzzy values ​​is greater than or equal to the preset threshold, determine that the jitter amount meets a second jitter level, wherein the second jitter level is greater than the first jitter level.

[0227] The presentation format of the first and second jitter levels is not specifically limited. For example, at least one of the first and second jitter levels can be a specific jitter level. For example, at least one of the first and second jitter levels can be a jitter level within a preset range, without specific limitations.

[0228] For example, the multiple target points in the above implementation are those mentioned above. Figure 3 The provided image deblurring method uses P target points, with angular acceleration data as described above. Figure 3 The provided image deblurring method uses P first 3D positions and multiple first 2D positions as described above. Figure 3 The provided method includes P first two-dimensional positions and multiple second two-dimensional positions as described above. Figure 3 The provided method includes P second-dimensional positions.

[0229] There is no specific restriction on the execution order of S1320 and S1310. For example, S1320 can be executed first, and then S1310 can be executed.

[0230] S1330: The electronic device determines a target deblurring strategy based on the amount of jitter, wherein the target deblurring strategy is matched with the degree of blurring of the image to be processed.

[0231] Matching the target deblurring strategy with the degree of blurring of the image to be processed means that the target deblurring strategy produces a clear image after deblurring the image to be processed, and that the clear image can meet the user's deblurring requirements.

[0232] For example, when the image to be processed is a globally blurred image, the target deblurring strategy is a global deblurring strategy. Conversely, when the image to be processed is a locally blurred image, the target deblurring strategy is a local deblurring strategy.

[0233] As an example of this application, the target deblurring strategy includes a local deblurring strategy or a global deblurring strategy. The electronic device determines the target deblurring strategy based on the jitter amount, including: if the jitter amount meets a first jitter level, determining the target deblurring strategy as a local deblurring strategy; if the jitter amount meets a second jitter level, determining the target deblurring strategy as a global deblurring strategy, wherein the second jitter level is greater than the first jitter level.

[0234] The descriptions of the first and second jitter levels can be found in the description in S1320 above, and will not be repeated here.

[0235] For example, the target deblurring strategy in the above implementation is as described above. Figure 3 The provided image deblurring method utilizes a global deblurring network or a local deblurring network to perform deblurring processing on the blurred image to be processed. The specific implementation process of the above method can be found in the above description. Figure 3 S350 to S380 in the illustrated embodiment will not be described again here.

[0236] Thus, the electronic device determines the deblurring strategy to be applied to the image based on the degree of jitter of the electronic device during the time period of capturing the image to be processed, thereby ensuring that the determined target deblurring strategy matches the image to be processed. Specifically, when the degree of jitter of the electronic device during the time period of capturing the image to be processed is small, a local deblurring strategy is applied to the image to be processed. When the degree of jitter of the electronic device during the time period of capturing the image to be processed is large, a global deblurring strategy is applied to the image to be processed.

[0237] S1340: The electronic device performs deblurring processing on the image to be processed according to the target deblurring strategy to obtain a clear image.

[0238] After the electronic device determines the target deblurring strategy based on the jitter amount, such as a global deblurring strategy or a local deblurring strategy, the implementation method of the electronic device performing deblurring processing on the image to be processed according to the global deblurring strategy or the local deblurring strategy is not specifically limited.

[0239] As an example of this application, an electronic device performs deblurring processing on an image to be processed according to a target deblurring strategy to obtain a clear image, including: the electronic device uses a target deblurring network to perform deblurring processing on the image to be processed to obtain a clear image, wherein the target deblurring network is a neural network model obtained by training an initial deblurring network using training samples, the training samples include training images, label images, and motion mask images, the training images are blurred images, the label images are clear images corresponding to the training images, and the motion mask images are images that include features of the blurred regions of the training images.

[0240] In one example, the initial deblurring network in the above steps includes an encoder and a decoder, wherein the decoder includes a first network module and a second network module, and the target deblurring network is specifically a neural network model obtained by training the encoder, the adjusted first network module, and the second network module according to a first loss value, wherein the first loss value is determined based on the difference between the predicted image and the training image, the predicted image is the image obtained by the second network module decoding the second feature image, the second feature image is the image obtained by the adjusted first network module extracting features from the first feature image, the adjusted first network module adjusts the parameters of the first network module according to the second loss value, the first feature image is the image obtained by the encoder extracting features from the training image, and the second loss value is determined based on the difference between the first feature image and the motion mask image.

[0241] Specifically, in the above steps, when the training image is a locally blurred image, a portion of the training image contains blurred content, the motion mask image includes features of a portion of the region, and the target deblurring network is a local deblurring network.

[0242] For example, the training images can be Figure 10 The image shown in Figure (1) can be a motion mask image. Figure 10 The image shown in (3) is shown in the figure.

[0243] Specifically, when the training image in the above steps is a globally blurred image, the entire region of the training image is blurred content, the motion mask image includes features of the entire region, and the target deblurring network is a globally blurred network.

[0244] For example, the training images can be Figure 11 The image shown in Figure (1) can be a motion mask image. Figure 11 The image shown in (3) is shown in the figure.

[0245] For example, the training samples in the above implementation are those mentioned above. Figure 3 The provided image deblurring method uses a training dataset containing N training samples, and the encoder is as described above. Figure 3 The first network module in the provided encoder and decoder is as described above. Figure 3 The first network module provided, and the second network module in the decoder are as described above. Figure 3 The provided second network module has a first loss value as described above. Figure 3 The provided image deblurring method uses loss value #2, and the second loss value is as described above. Figure 3 The loss value #1 in the provided method, and the N first feature images are as described above. Figure 3 The provided method uses N feature images #1 and N second feature images as described above. Figure 3 The provided method uses N feature images #2. The specific implementation process for training the initial deblurring network to obtain the target deblurring network in the above implementation method can be found in the above description. Figure 12 The training methods provided will not be elaborated here.

[0246] As another example of this application, an electronic device performs deblurring on an image to be processed according to a target deblurring strategy to obtain a clear image. Exemplarily, it may include the following steps: the electronic device uses a target deblurring network to perform deblurring on the image to be processed to obtain a clear image, wherein the target deblurring network is a neural network model obtained by training an initial deblurring network using training samples. The training samples include training images and label images, where the training image is a blurred image and the label image is a clear image corresponding to the training image.

[0247] In the above steps, the initial deblurring network includes an encoder and a decoder. The target deblurring network is specifically a neural network model obtained by training the initial deblurring network using the difference between the label image and the prediction image. The prediction image is obtained by the decoder after decoding the feature image obtained from the encoder. The feature image obtained from the encoder is the image obtained by the encoder after extracting features from the training image.

[0248] As an example of this application, before the electronic device acquires the image to be processed, the method further includes: displaying a shooting interface; the electronic device acquiring the image to be processed includes: in response to a shooting operation on the shooting interface, the electronic device obtains the image to be processed through a camera; after the electronic device performs deblurring processing on the image to be processed according to a target deblurring strategy to obtain a clear image, the method further includes: the electronic device displaying an interface including the clear image.

[0249] In this way, in response to the user's shooting operation, the electronic device acquires the image to be processed, then the electronic device processes the image using a target deblurring strategy, and finally the electronic device displays an interface with a clear image to the user. The process of the electronic device performing deblurring on the image to be processed is performed in the background, that is, the user is unaware of the process, which can improve the user's visual experience.

[0250] It should be understood that the above Figure 13 The image deblurring method shown is for illustrative purposes only and does not constitute any limitation on the image deblurring method provided in this application. It should be noted that the above... Figure 13 The illustrated image deblurring method uses an electronic device to capture an image of a subject, and the electronic device performing the image deblurring method as an example. Optionally, the electronic device performing the image deblurring method and the electronic device capturing the image of the subject can be two different electronic devices. In this case, after obtaining the image of the subject, the electronic device capturing the image can transmit it to the electronic device performing the image deblurring method, so that the electronic device performing the image deblurring method can acquire the image of the subject.

[0251] In this embodiment, after acquiring an image to be processed, the electronic device first determines a deblurring strategy (target deblurring strategy) based on the jitter level. Since the jitter level represents the degree of shaking of the electronic device during the time period when the camera acquires the image, the target deblurring strategy determined based on the jitter level matches the degree of blur in the image. Then, the electronic device performs deblurring on the image according to the target deblurring strategy determined based on the jitter level to obtain a clear image. This method avoids the problem of mismatch between the deblurring strategy and the degree of blur in the acquired image when performing deblurring on the image in traditional technologies. Therefore, this method can improve the image deblurring effect, thereby better satisfying the user's visual experience.

[0252] In this application, the application form of the image deblurring method provided above is not specifically limited on the terminal device. The following is a schematic diagram of the user interface of the mobile phone, taking the application of the image deblurring method provided in this application on a mobile phone as an example.

[0253] As an example, in response to a user triggering the phone's camera function, the image deblurring method provided in this application is called to deblur the blurred image captured by the phone. In this case, the user is unaware of the image deblurring process performed by the phone.

[0254] For example, with Figure 14 Taking Figure (1) as an example, the shooting interface S4 of the mobile phone displays objects including people and buildings. In response to the user triggering the camera control 10 of the mobile phone, the mobile phone can then take a picture of the objects to obtain an image. Although the mobile phone shakes during the shooting process, because the mobile phone performs the image deblurring method provided in this application on the acquired blurry image, the final image obtained after the mobile phone takes a picture of the objects is a clear image. For example, Figure 14 The clear image shown in the mobile phone's photo album interface S5 is illustrated in (2) of the figure.

[0255] As another example, in response to the user triggering the deblurring control 20 of the phone's photo album interface, the image deblurring method provided in this application is called to deblur the blurred image stored in the photo album. In this case, the user can perceive the phone performing the deblurring process on the blurred image.

[0256] For example, with Figure 15Taking Figure (1) as an example, the shooting interface S7 of the mobile phone displays subjects including people and buildings. In response to the user triggering the camera control 10 of the mobile phone, the mobile phone can then take a picture of the subject to obtain an image. During the shooting process, due to the shaking of the mobile phone, the final image obtained after the mobile phone takes a picture of the subject is a blurry image. For example, Figure 15 The blurry image shown in the mobile phone's photo album interface S8 is illustrated in (2). Afterwards, the user can perform deblurring on the blurry image in S8 by triggering the deblurring control 20 in the photo album interface S8 to obtain a clear image. Please refer to [link to relevant documentation]. Figure 15 The clear image shown in the camera interface S9 of the mobile phone is illustrated in (2).

[0257] This application also provides a computer program product that, when executed by a processor, implements the image deblurring method described in any of the method embodiments of this application.

[0258] The computer program product can be stored in memory, for example, it is a program. The program is eventually converted into an executable object file that can be executed by the processor after processes such as preprocessing, compilation, assembly and linking.

[0259] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer, implements the image deblurring method described in any of the method embodiments of this application. The computer program may be a high-level language program or an executable object program.

[0260] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0261] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0262] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0263] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0264] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and other division methods may be used in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0265] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0266] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

Claims

1. An image deblurring method, characterized in that, When applied to electronic devices, the method includes: Obtain the image to be processed; The jitter of the electronic device is obtained, wherein the jitter is used to represent the degree of jitter of the electronic device during the time period in which the camera of the electronic device acquires the image to be processed; Based on the jitter amount, a target deblurring strategy is determined, wherein the target deblurring strategy is matched with the degree of blurring of the image to be processed; The image to be processed is deblurred using a target deblurring network to obtain a clear image. The target deblurring network is a neural network model obtained by training an initial deblurring network using training samples. The training samples include training images, label images, and motion mask images. The training images are blurred images, the label images are clear images corresponding to the training images, and the motion mask images are images that include features of the blurred regions of the training images. The initial deblurring network includes an encoder and a decoder, wherein the decoder includes a first network module and a second network module, and the target deblurring network is specifically a neural network model obtained by training the encoder, the adjusted first network module, and the second network module according to a first loss value, wherein the first loss value is determined based on the difference between the predicted image and the training image, the predicted image is the image obtained by the second network module after decoding the second feature image, the second feature image is the image obtained by the adjusted first network module after extracting features from the first feature image, the adjusted first network module is obtained by adjusting the parameters of the first network module according to the second loss value, the first feature image is the image obtained by the encoder after extracting features from the training image, and the second loss value is determined based on the difference between the first feature image and the motion mask image.

2. The method according to claim 1, characterized in that, The target deblurring strategy includes a local deblurring strategy or a global deblurring strategy, and determining the target deblurring strategy based on the jitter amount includes: If the jitter amount meets the first jitter level, the target deblurring strategy is determined to be the local deblurring strategy; If the jitter amount meets the second jitter level, the target deblurring strategy is determined to be the global deblurring strategy, wherein the second jitter level is greater than the first jitter level.

3. The method according to claim 1, characterized in that, In the case where the training image is a globally blurred image, the entire region of the training image is blurred content, the motion mask image includes features of the entire region, and the target deblurring network is a globally blurred network; In the case where the training image is a locally blurred image, a portion of the training image is blurred content, the motion mask image includes features of the portion of the image, and the target deblurring network is a local deblurring network.

4. The method according to claim 1, characterized in that, The step of obtaining the jitter amount of the electronic device includes: Acquire angular acceleration data, wherein the angular acceleration data is used to represent the pose data of the electronic device during the time period in which the camera acquires the image to be processed; The jitter amount is determined based on the angular acceleration data.

5. The method according to claim 4, characterized in that, The angular acceleration data includes multiple angular accelerations, wherein the multiple angular accelerations correspond to multiple target points in the image to be processed, each angular acceleration is the pose data of the electronic device when the camera captures the corresponding target point, the multiple target points correspond to multiple first two-dimensional positions, and the position of each target point in the image to be processed is the corresponding first two-dimensional position. Furthermore, determining the jitter amount based on the angular acceleration data includes: Based on the multiple angular accelerations and the multiple first two-dimensional positions, multiple second two-dimensional positions are obtained; Based on the plurality of second two-dimensional positions and the plurality of first two-dimensional positions, a plurality of fuzzy values ​​corresponding to the plurality of target points are obtained; The jitter amount is determined based on the plurality of fuzzy values.

6. The method according to claim 5, characterized in that, Determining the jitter amount based on the plurality of fuzzy values ​​includes: If the smallest of the plurality of fuzzy values ​​is less than a preset threshold, the jitter amount is determined to meet the first jitter level. If the smallest of the plurality of fuzzy values ​​is greater than or equal to the preset threshold, the jitter amount is determined to satisfy a second jitter level, wherein the second jitter level is greater than the first jitter level.

7. The method according to any one of claims 1 to 6, characterized in that, Before acquiring the image to be processed, the method further includes: Display the shooting interface; The process of acquiring the image to be processed includes: In response to a shooting operation on the shooting interface, the image to be processed is obtained through the camera; After performing deblurring processing on the image to be processed according to the target deblurring strategy to obtain a clear image, the process further includes: The interface displays the clear image.

8. An image deblurring device, characterized in that, The image deblurring device, used in electronic devices, includes a processing unit, which is used for: Obtain the image to be processed; The jitter of the electronic device is obtained, wherein the jitter is used to represent the degree of jitter of the electronic device during the time period in which the camera of the electronic device acquires the image to be processed; Based on the jitter amount, a target deblurring strategy is determined, wherein the target deblurring strategy is matched with the degree of blurring of the image to be processed; The image to be processed is deblurred using a target deblurring network to obtain a clear image. The target deblurring network is a neural network model obtained by training an initial deblurring network using training samples. The training samples include training images, label images, and motion mask images. The training images are blurred images, the label images are clear images corresponding to the training images, and the motion mask images are images that include features of the blurred regions of the training images. The initial deblurring network includes an encoder and a decoder, wherein the decoder includes a first network module and a second network module, and the target deblurring network is specifically a neural network model obtained by training the encoder, the adjusted first network module, and the second network module according to a first loss value, wherein the first loss value is determined based on the difference between the predicted image and the training image, the predicted image is the image obtained by the second network module after decoding the second feature image, the second feature image is the image obtained by the adjusted first network module after extracting features from the first feature image, the adjusted first network module is obtained by adjusting the parameters of the first network module according to the second loss value, the first feature image is the image obtained by the encoder after extracting features from the training image, and the second loss value is determined based on the difference between the first feature image and the motion mask image.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program from the memory, such that the processor performs the image deblurring method according to any one of claims 1 to 7.

10. A chip system, characterized in that, The chip system includes a processor, which, when executing instructions, performs the image deblurring method as described in any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the image deblurring method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Deblurring method and electronic equipment

    CN116993620A