Image processing method and apparatus, and program product
By acquiring image pairs and performing image alignment and fusion processing, the problem of low image matching degree and consistency is solved, the effect of image deblurring is improved, and high-quality image output is achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2025-10-16
- Publication Date
- 2026-05-21
AI Technical Summary
In existing image processing methods, the matching degree and consistency of unprocessed image pairs are not high, which affects the deblurring effect.
By acquiring a set of image pairs, image alignment and fusion are performed to improve the matching degree between image pairs, and image processing models are used for deblurring.
It improves the effect of image deblurring, enhances image clarity and readability, and achieves high-quality image output.
Smart Images

Figure CN2025128174_21052026_PF_FP_ABST
Abstract
Description
Image processing methods, devices and software products
[0001] This disclosure claims priority to Chinese patent application No. 202411617425.7, filed on November 12, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus and program product. Background Technology
[0003] With the development of science and technology, mobile devices have become increasingly popular, serving as primary tools for recording life. Consequently, users' demands for the photography functions of mobile devices have also increased. However, during the shooting process, issues such as insufficient lighting, poor depth-of-field control, lens distortion, and aberrations may arise, leading to problems such as blurred image edges and poor overall sharpness, thus affecting the user's visual experience.
[0004] Currently, image processing methods based on image processing models can utilize supervised learning of blurred-sharp image pairs collected in a large number of scenarios to learn feature maps for image deblurring, thereby improving image sharpness. Summary of the Invention
[0005] On one hand, this disclosure provides an image processing method, which includes: acquiring an image pair set, the image pair set including at least one image pair, each image pair including a first image and a second image, wherein the image sharpness of the second image is greater than that of the first image;
[0006] Image alignment is performed on each image pair in the image pair set to obtain the aligned image in each image pair;
[0007] Image fusion processing is performed on the aligned image in each image pair and the first image to obtain a processed image pair set; wherein the processed image pair set includes at least one processed image pair, and the matching degree between the second image and the first image in the processed image pair is greater than the matching degree between the second image and the first image in the unprocessed image pair.
[0008] On the other hand, this disclosure also provides an image processing apparatus, comprising: an acquisition module for acquiring an image pair set, the image pair set including at least one image pair, each image pair including a first image and a second image, the image sharpness of the second image being greater than the image sharpness of the first image; a processing module for performing image alignment processing on each image pair in the image pair set respectively to obtain an aligned image in each image pair; the processing module is further configured to perform image fusion processing on the aligned image in each image pair and the first image to obtain a processed image pair set; wherein the processed image pair set includes at least one processed image pair, the matching degree between the second image and the first image in the processed image pair is greater than the matching degree between the second image and the first image in the unprocessed image pair.
[0009] In another aspect, an electronic device is provided, comprising: a processor and a memory; the memory storing processor-executable instructions; when the processor is configured to execute the instructions, causing the communication device to implement any of the methods described above.
[0010] In another aspect, a computer-readable storage medium is provided, including a non-transitory computer-readable storage medium that stores computer instructions that, when executed on a computer, cause the computer to perform any of the methods described above.
[0011] On the other hand, a computer program product containing computer instructions is provided, which, when run on a computer, cause the computer to perform any of the methods described above. Attached Figure Description
[0012] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.
[0013] Figure 1 is a schematic diagram of an image processing system according to some embodiments;
[0014] Figure 2 is a flowchart of an image processing method according to some embodiments;
[0015] Figure 3 is a flowchart of another image processing method according to some embodiments;
[0016] Figure 4 is a flowchart of another image processing method according to some embodiments;
[0017] Figure 5 is a flowchart of another image processing method according to some embodiments;
[0018] Figure 6 is a flowchart of another image processing method according to some embodiments;
[0019] Figure 7 is a block diagram of an electronic device according to some embodiments;
[0020] Figure 8 is a block diagram of an electronic device according to some embodiments. Detailed Implementation
[0021] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this disclosure.
[0022] Unless the context otherwise requires, throughout the specification and claims, the term "comprise" and its other forms, such as the third-person singular "comprises" and the present participle "comprising," are interpreted as open-ended and encompassing, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiments," "example," "specific example," or "some examples," etc., are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples.
[0023] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0024] In this disclosure, the terms "exemplarily" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplarily" or "for example" in this disclosure should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the terms "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0025] In addition, the use of “based on” implies openness and inclusivity, because processes, steps, calculations or other actions “based on” one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated.
[0026] Current image deblurring methods typically rely on training image processing models with large datasets, and then performing image deblurring based on the trained image processing models. However, the matching degree and consistency of unprocessed image pairs are not high, which affects the deblurring effect.
[0027] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0028] As shown in Figure 1, this application provides a schematic diagram of an image processing system. The image processing system 100 includes a terminal device 10 and a server 20.
[0029] Terminal device 10 can be a device with data processing capabilities. Examples include cameras, camcorders, scanners, video recorders, mobile phones, tablets, desktop computers, laptops, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, cellular phones, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc. This disclosure does not impose any special limitations on the specific form of terminal device 10. In some embodiments, terminal device 10 can interact with the user through one or more methods such as a keyboard, touchpad, touchscreen, remote control, voice interaction, or handwriting device, for example, performing image capture or image deblurring based on user instructions.
[0030] In some embodiments, the terminal device 10 may have its own image acquisition function. Alternatively, the terminal device may be connected to an image acquisition device.
[0031] In some embodiments, the image pair set obtained based on the image processing method provided in this disclosure can be used to train an image processing model, which is used for image deblurring. Exemplarily, the image processing model can be deployed on the terminal device 10, so that the terminal device 10 can use the image processing model to deblur images with low clarity. In some embodiments, after a user takes an image using the terminal device 10, the terminal device 10 can automatically deblur the image to obtain a clear image.
[0032] For example, the terminal device 10 may also have an image display function, or be connected to an image display device, for displaying the acquired image. For instance, if the acquired image has low resolution, the terminal device 10 may automatically perform deblurring processing on the low-resolution image using an image processing model to display the deblurred image to the user. Alternatively, if the acquired image has low resolution, the terminal device 10 may directly display the acquired image to the user, and in response to a deblurring command input by the user, the terminal device may further perform deblurring processing on the low-resolution image using an image processing model to display the deblurred image to the user.
[0033] In some embodiments, the terminal device 10 can execute the image processing method provided in this disclosure to obtain an image pair set. Based on this image pair set, a model is trained to obtain a trained image processing model.
[0034] In one example, terminal device 10 can acquire image pairs based on its own captured images for training an image processing model. For instance, terminal device 10 can capture a first image and a second image from the same target, i.e., the aforementioned image pair, where the second image has higher image clarity than the first image. In this case, terminal device 10 can acquire a large number of image pairs under various suitable conditions, such as different focal lengths, different lighting conditions, different degrees of camera shake, different shooting angles, different depths of field, or dynamic / static image types.
[0035] In another example, the number of terminal devices 10 in the image processing system 100 described above can be multiple. Each terminal device 10 can acquire an image pair set based on images acquired by itself and / or other terminal devices 10 for training an image processing model.
[0036] In another example, terminal device 10 can also obtain image pairs from server 20 for training an image processing model. It should be understood that server 20 can store or retrieve image pairs uploaded by each terminal device 10 in real time, and send them to terminal device 10 when needed for image processing model training. In this case, terminal device 10 can connect to server 20. The connection method can be wireless, such as Bluetooth or Wi-Fi; or it can be wired, such as fiber optic, etc., without limitation.
[0037] In some embodiments, the image processing method provided in this disclosure can be applied to the server 20 described above, or it can be jointly executed by the server 20 and the terminal device 10. This disclosure does not impose any special limitations on this.
[0038] For example, server 20 can execute the image processing method provided in this disclosure to obtain an image pair set. Based on this image pair set, a model can be trained to obtain a trained image processing model, which is then sent to terminal device 10. Alternatively, server 20 and terminal device 10 can collaboratively train the image processing model and deploy the trained model on terminal device 10.
[0039] In some embodiments, server 20 may be a server that provides various services, such as a backend server that supports terminal device 10. It can analyze and process received requests and other data, and feed the processing results back to terminal device 10. Server 20 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data servers.
[0040] It should be noted that Figure 1 is merely an exemplary framework diagram. The number of devices or nodes included in Figure 1, and the names of each device, are not limited. For example, the system 100 may include more or fewer nodes or devices. For instance, in addition to the devices shown in Figure 1, the system 100 may also include an image acquisition device. Alternatively, the system 100 shown in Figure 1 may also include only the terminal device 10.
[0041] The system architecture and business scenarios described in the embodiments of this disclosure are intended to more clearly illustrate the technical solutions of the embodiments of this disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this disclosure are also applicable to similar technical problems.
[0042] This disclosure also provides an image processing model (or electronic device or other possible name) for executing an image processing method provided in this disclosure. For example, the image processing device may be the terminal device 10 in the system 100 described above, or it may be a functional module of the terminal device 10. Alternatively, the image processing device may be the server 20 in the system 100 described above, or it may be a functional module of the server 20. This disclosure does not limit the scope of the disclosure in this regard.
[0043] The method provided in this disclosure will now be described in detail with reference to the accompanying drawings.
[0044] As shown in Figure 2, this disclosure provides an image processing method, which includes:
[0045] S101. Obtain the image pair set.
[0046] Here, the image pair set includes at least one image pair, and each image pair includes a first image and a second image, wherein the image sharpness of the second image is greater than that of the first image. In some embodiments, the image pair set may also include only one image pair set, that is, the image processing method provided in this disclosure can be applied to image pair sets with multiple image pairs, or to image pair sets with a single image pair.
[0047] In some embodiments, the acquisition target corresponding to the first image in each image pair is the same as the acquisition target corresponding to the second image.
[0048] In some embodiments, the image pair provided in this disclosure may also be referred to as a blurry-sharp image pair, that is, the first image may also be referred to as a blurry image and the second image may also be referred to as a sharp image.
[0049] In some embodiments, the first image refers to an image or video frame that is not clear in detail, has blurred edges, or is not clear in visual effect due to various possible reasons such as target motion, out-of-focus blur, lens distortion, low image resolution, etc.
[0050] For example, the first image can be a blurred image actually captured or scanned in a real shooting environment due to factors such as target motion, out-of-focus blur, lens distortion, or low image resolution. Alternatively, the first image can also be a simulated blurred image generated by blurring a clear image using image generation techniques, such as Gaussian blur or motion blur. Or, the first image can also be a simulated blurred image obtained by setting shooting parameters to simulate the blur conditions corresponding to the aforementioned target motion, out-of-focus blur, lens distortion, and low image resolution.
[0051] In some embodiments, the second image and the first image typically appear in pairs, usually different images of the same target captured under different conditions. The second image may refer to an image with higher resolution, richer detail, sharper edges, and no obvious blurring or defocusing compared to the first image. It should be understood that, compared to the first image, the second image is more accurately able to reflect image features that indicate image clarity, such as the shape, texture, and color of the captured object.
[0052] For example, the second image can be a high-resolution image obtained by capturing the target in a real shooting environment. Alternatively, the second image can also be an image obtained after processing to improve image clarity, such as sharpening and denoising.
[0053] For example, the collection target can be a human body, a face, a vehicle, an animal, a plant, a document, or other possible target object. Alternatively, the collection target can also be, for example, a supermarket, a station, a road, or other possible target scene.
[0054] In some embodiments, the image pair set includes one or more image pairs from different acquisition scenarios. That is, the image processing device can acquire multiple image pairs from different acquisition scenarios to form the aforementioned image pair set. It should be understood that in the process of forming the image pair set, it is necessary to ensure that each image pair in the image pair set contains a blurred image (i.e., the first image) obtained from image acquisition of the same acquisition object, and a relatively clear image (i.e., the second image).
[0055] For example, the scene being captured is associated with at least one of the following: lens type, lens focal length information, lighting environment, degree of shake, shooting angle, depth of field range, and object motion state.
[0056] Here, lens type refers to the type of optical component used in the image acquisition device to acquire images, typically categorized based on factors such as focal length, aperture, and angle of view. Examples include prime lenses, zoom lenses, wide-angle lenses, and fisheye lenses. That is, the image set can include image pairs acquired based on multiple lens types. In some embodiments, the first and second images in the same image pair used the same lens type during image acquisition. This improves the accuracy of the image processing model trained on this image set for recognizing images of different lens types, enhancing the image processing effect on images of different lens types. Using this image processing model for deblurring further improves the deblurring effect on images of different lens types.
[0057] Lens focal length information refers to the distance from the center point of the lens to the imaging plane. In other words, the image set can include image pairs acquired based on various lens focal lengths. This can improve the accuracy of image processing models trained on this image set for recognizing images with different lens focal lengths, and enhance the image processing performance for images with different lens focal lengths.
[0058] Lighting environment refers to the lighting conditions of the image acquisition environment, including light intensity, direction, and color. Images acquired under different lighting conditions may differ in exposure, color, and sharpness. In other words, an image set can include image pairs acquired under various lighting environments. This can improve the accuracy of image processing models trained on this image set for image recognition under different lighting environments, thus enhancing the image processing performance of images from diverse lighting conditions.
[0059] Camera shake refers to the movement of a camera during shooting due to various reasons (such as unsteady handhold, wind, etc.). Significant camera shake can lead to blurry photos and reduced image quality. Therefore, an image set can include image pairs acquired based on various levels of camera shake. This can improve the accuracy of image processing models trained on such image sets in recognizing images with different levels of shake, thus enhancing the image processing performance for images with varying degrees of shake.
[0060] Shooting angle can include multiple aspects such as shooting height, shooting direction, and shooting distance, which can affect the perspective and composition of an image. For example, shooting heights can include level shots, overhead shots, and low-angle shots; shooting directions can include frontal, side, and oblique shots; and images obtained from different shooting distances can be categorized as close-ups, mid-range shots, and distant shots. In other words, an image set can include image pairs acquired from multiple shooting angles. This can improve the accuracy of image processing models trained on this image set in recognizing images acquired from different shooting angles, thus enhancing the image processing performance of images acquired from various shooting angles.
[0061] Depth of field (DOF) refers to the range of distances, including the front and back of the subject, that allow for a sharp image to be captured from the front of the camera lens or other imager. A larger DOF range results in a wider area of sharp image; a smaller DOF range results in a narrower area of sharp image. In other words, an image set can include image pairs acquired based on multiple DOF ranges. This can improve the accuracy of image processing models trained on such image sets for recognizing images with different DOF ranges, thus enhancing the image processing performance for images with varying depths of field.
[0062] The motion state of an object refers to its static or dynamic state during image acquisition. In other words, an image set can include image pairs acquired under various object motion states. This improves the accuracy of image processing models trained on this image set for recognizing images under different object motion states, and enhances the image processing performance for images under different object motion states.
[0063] It should be noted that the image processing device can generate an image pair set for training the image processing model by acquiring a large number of image pairs under different focal lengths, lighting conditions, degrees of camera shake, shooting angles, depth of field ranges, and dynamic and static image types. This enhances the diversity of training samples, making the image processing method based on the model widely applicable to various complex shooting environments, such as real-time image processing, mobile device shooting, and real-time video processing. Furthermore, the trained image processing model maintains good image processing performance under different types of images and shooting conditions. Using this image processing model for deblurring can improve the deblurring effect on images with different lens types. Thus, the image processing model can adapt to more types of blurring situations and effectively address different types of blurring challenges, such as motion blur, depth of field blur, and lens distortion blur. This improves the universality, practicality, and robustness of the deblurring method based on this image processing model.
[0064] S102. Perform image alignment processing on each image pair in the image pair set to obtain the aligned image in each image pair.
[0065] Here, in each image pair, the alignment of the image with the first image is higher than the alignment of the second image with the first image.
[0066] Here, matching degree refers to the degree of matching or similarity between the second image and the first image in an image pair. Matching degree can be determined based on the similarity between the image viewpoint of the second image and the image viewpoint of the first image in the image pair, the similarity between the lighting environment when the second image was taken and the lighting environment when the first image was taken, the similarity between the shooting angle of the second image and the shooting angle of the first image, or other possible factors.
[0067] In this disclosure, the matching degree between the second image and the first image in an image pair may also be referred to as similarity, consistency, or other terms that express the same or similar characteristics, and this disclosure does not limit this terminology.
[0068] It should be noted that the first and second images in the acquired image pairs may differ in terms of viewing angle, lighting environment, and shooting angle, resulting in a low matching degree between the first and second images in the acquired image pairs. In such cases, training an image processing model based on these acquired image pairs will also result in a low matching degree between the output image and the input image, thus affecting the deblurring effect of the model. Therefore, image consistency processing can be performed on each image pair in the image pair set to improve the matching degree between the second and first images in each object pair, thereby improving the deblurring effect of the image processing model.
[0069] In some embodiments, the image processing apparatus aligns the second image to the first image for each image pair in the image pair set, based on image features of the first image in the image pair, image features of the second image in the image pair, and optical flow estimation techniques, to obtain an aligned image.
[0070] Here, the alignment accuracy between the aligned image and the first image is higher than that between the second image and the first image, and the matching degree between the aligned image and the first image is higher than that between the second image and the first image.
[0071] In other words, the image processing device can perform image alignment processing on each image pair in the image pair set, thereby improving the matching degree between the images in each image pair.
[0072] In some embodiments, image features include at least one of field of view, color, and brightness.
[0073] It should be noted that the image processing device can perform image alignment processing on the image pair based on the field of view, color, brightness, or other possible image features of the images in the image pair. This makes the two images in the image pair more consistent or similar, that is, improves the matching degree between the two images in the image pair. As a result, the matching degree between the second image and the first image in the processed image pair is greater than the matching degree between the second image and the first image in the unprocessed image pair, thereby improving the deblurring effect of the image processing model.
[0074] Optical flow refers to the pattern of pixel movement on the surface of an object in an image sequence. In computer vision and image processing, optical flow estimation is a technique used to estimate the displacement between pixels between adjacent image frames. The goal of optical flow estimation is to find corresponding pixels between consecutive image frames and calculate the displacement or velocity between them.
[0075] Here, the alignment accuracy between the aligned image and the first image is higher than that between the second image and the first image, and the matching degree between the aligned image and the first image is higher than that between the second image and the first image.
[0076] Here, alignment refers to the processing steps to make two images visually match or blend better. Aligning the second image with the first image means that image processing makes the features or content in the second image consistent or nearly consistent with the corresponding features or content in the first image in terms of position, thereby obtaining the aligned image mentioned above.
[0077] In this disclosure, alignment may also be referred to as image registration or other terms that express the same or similar meanings, and this disclosure does not limit it.
[0078] In some embodiments, the image processing apparatus may first estimate a high-precision optical flow field from the first image to the second image based on image features of the first image and the second image in the image pair, as well as optical flow estimation techniques. Then, based on the high-precision optical flow field from the first image to the second image, the aligned image is determined.
[0079] It should be noted that the image processing device can calculate the optical flow field between the first and second images, thereby enabling the second image to be accurately aligned to the first image, achieving higher-precision image registration (alignment) between the two images, such as pixel-level registration. This provides reliable input data for subsequent image fusion steps.
[0080] In some embodiments, the specific description of the image consistency processing procedure for any image pair can be referred to the embodiment shown in Figure 3 below, and will not be repeated here.
[0081] S103. Perform image fusion processing on the aligned image in each image pair and the first image to obtain the processed image pair set.
[0082] Here, the processed image pair set includes at least one processed image pair, and the matching degree between the second image and the first image in the processed image pair is greater than the matching degree between the second image and the first image in the unprocessed image pair.
[0083] In some embodiments, step S103 can be specifically implemented as the following steps S11-S14:
[0084] S11. Separate the low-frequency brightness information of the first image and the aligned image.
[0085] For example, the image processing device can separate (or extract) the low-frequency brightness information of the first image and the aligned image by applying a mean filter (such as average filtering or Gaussian filtering) multiple times.
[0086] Here, the low-frequency brightness information can also be referred to as the initial background image, which is a relatively smooth background image. The aforementioned mean filter is a smoothing filter that can reduce high-frequency details in the image, thereby separating out a smoother background image.
[0087] In some embodiments, the process of separating low-frequency luminance information can be achieved by the following formula (1): L light =B1ur n (Img) Formula (1)
[0088] Here, Img is the input image, and Blur is the input image. n (Img) represents applying the mean filter operation n times consecutively to the image Img, where n is a positive integer. L lightThis represents the extracted low-frequency brightness information (or initial background image). The input image can be either the first image or an aligned image.
[0089] S12. Generate a uniform background image based on the first image after removing low-frequency brightness information.
[0090] For example, the image processing device can calculate the mean value of each color channel (blue b, green g, red r) of the first image after removing low-frequency brightness information, and fill these mean values into a new image to form a uniform background image.
[0091] In some embodiments, the process of generating a uniform background image can be achieved by the following formula (2): back_img(x,y)=[rand b ,rand g ,rand r ] Formula (2)
[0092] Here, rand c This represents the mean value of each color channel c (blue b, green g, red r) of the image. `back_img(x,y)` refers to a generated random, uniform background image where all pixel values are set to the mean value of each channel.
[0093] S13. The first image after removing low-frequency brightness information and the aligned image after removing low-frequency brightness information are merged with the background image respectively to obtain the merged first image and the merged aligned image.
[0094] Here, the merged first image is the processed first image from step S202 above.
[0095] It should be understood that the merged first image obtained here retains some details of the original first image and has a relatively uniform background, which can be understood as a first image with uniform lighting. Similarly, the merged aligned image obtained here retains some details of the original aligned image and has a relatively uniform background, which can be understood as an aligned image with uniform lighting.
[0096] In some embodiments, the process of generating the merged image can be achieved by the following formula (3): new_img = img - L light +back_img formula (3)
[0097] Here, new_img is the merged image, and img-L light This represents the image after removing low-frequency brightness information; back_img is a uniform background image.
[0098] It should be understood that the merged first image and the merged aligned image can be obtained respectively using the method shown in the above formula (3).
[0099] S14. The brightness information of the merged aligned image and the chromaticity information of the merged first image are fused to obtain the processed second image.
[0100] For example, the image processing device can convert the merged first image to the YUV color space, and separate the luminance (Y) and chrominance (U, V) channels to obtain the converted first image. Here, the converted first image is represented as: YUV blurred =ConvertToYUV(blurred), ConvertToYUV(·) means converting the image to the YUV color space.
[0101] An image processing device can convert the merged aligned image to the YUV color space, separating the luminance (Y) and chrominance (U, V) channels to obtain the converted aligned image. Here, the converted aligned image is represented as: YUV sharp =ConvertToYUV(sharp).
[0102] Furthermore, the image processing device can fuse the luminance channel of the merged aligned image and the chroma channel of the merged first image to obtain a fused image. The fused image is then converted back to the blue-green-red (BGR) color space to obtain the processed second image. This achieves the fusion of luminance information from the merged aligned image and chroma information from the merged first image. Thus, the processed second image can achieve luminance and color matching between the first image and the aligned image.
[0103] Here, the fused image can be represented as M. YUV =[Y sharp U blurred V blurred ], Y sharp This represents the luminance channel of the merged aligned image in the YUV color space, U blurred V blurred This represents the chroma channel of the first image after merging in the YUV color space.
[0104] The processed second image can be represented as Image final =ConvertToBGR(M YUV ConvertToBGR(·) means converting the image to the BGR color space.
[0105] Here, the YUV color space is mainly used for television signal transmission and video compression. Y represents luminance, and U and V represent chrominance. The separation of luminance and color in the YUV space conforms to human visual perception, and color matching can be performed more naturally based on the YUV space. The BGR color space is mainly used for computer graphics and image processing. B, G, and R represent blue, green, and red, respectively.
[0106] It should be noted that, based on the image fusion process described above, brightness and color can be processed separately, extracting and removing the effects of illumination, thus avoiding detail loss or unnatural color transitions caused by global operations. Here, by performing fine image fusion in the brightness (Y channel) space on the aligned image, this separation of brightness and color processing can effectively reduce brightness distortion in blurred areas, making the final image more natural and richer in detail. This allows for better preservation of image details and edge information, reducing detail loss and artifact generation.
[0107] In some embodiments, the resulting set of processed image pairs can be used to train an image processing model to perform various image processing operations on the input image. It should be understood that the trained image processing apparatus can also achieve higher image processing performance because the image pairs used for training have a higher degree of matching, and the image pairs in the processed image pair set can be aligned with high precision and maintain realism and naturalness.
[0108] For example, the obtained set of processed image pairs can be used to train an image processing model for deblurring. The resulting deblurred image has higher clarity than the original image, improving the overall clarity, readability, and overall quality of the deblurred image, thereby enhancing the effect of image deblurring.
[0109] Based on the technical solution provided in this disclosure, image pairs can be processed to achieve high-precision alignment of the clear and blurred images within the image pair, followed by image fusion. The resulting processed image maintains realism and naturalness, achieving high-quality image output. Furthermore, the matching degree between the clear image (second image) and the blurred image (first image) in the image pair is improved. Therefore, using the processed image pairs to train an image processing model can enhance the model's image processing performance. Taking image processing for deblurring as an example, the deblurred image obtained by training with the image pair set obtained using this technical solution has higher clarity than the original image, improving the overall clarity, readability, and overall quality of the deblurred image, thereby enhancing the deblurring effect.
[0110] In some embodiments, as shown in FIG3, this disclosure also provides another image processing method, which can realize the image alignment processing of each image pair in step S102 above, specifically including:
[0111] S201. Perform feature extraction on the first image and the second image respectively to obtain multiple image feature layers of the first image and the second image.
[0112] Here, different image feature layers correspond to different image resolutions.
[0113] For example, the aforementioned multiple image feature layers can also be referred to as multi-layer feature maps. A convolutional neural network can be used to extract features from the first image (denoted as I1) to obtain a multi-layer feature map of the first image. Similarly, a convolutional neural network can be used to extract features from the second image (denoted as I2) to obtain a multi-layer feature map of the second image. Here, the image resolution of the multi-layer feature maps decreases progressively, forming a pyramid structure.
[0114] Here, a convolutional neural network (CNN) is a deep learning model used to process data with a grid-like topological structure (such as image data). Its core lies in the convolution operation, which aims to efficiently identify local features in the data and handle the translation invariance of the data. A pyramid structure is a hierarchical arrangement of topics or information; by transforming images at different scales, it generates a series of images with progressively decreasing resolution, thereby extracting and analyzing image features at different scales.
[0115] In some embodiments, the image processing device can cascade a pyramid structure convolutional neural network, that is, combine a convolutional neural network with a pyramid structure, and achieve accurate localization of key points in the image by constructing a multi-scale feature pyramid, thereby obtaining the above-mentioned multiple image feature layers (or multi-layer feature maps).
[0116] Thus, through multi-level feature extraction, rich image information can be provided for the subsequent optical flow estimation process.
[0117] S202. Based on the image feature layer with the smallest image resolution and optical flow estimation technology, the initial optical flow field from the first image to the second image is obtained.
[0118] Here, step S202 can also be called the initial optical flow estimation step.
[0119] For example, the image processing device can initialize the optical flow map from the bottom layer of the pyramid structure (i.e., the feature map with the lowest image resolution or the image feature layer with the smallest image resolution), that is, obtain the initial optical flow field from the first image to the second image.
[0120] S203. The initial optical flow field is interpolated and upsampled to obtain the upsampled optical flow field. The image resolution of the upsampled optical flow field is greater than that of the initial optical flow field.
[0121] Here, step S203 can also be called the upsampling step.
[0122] For example, if the image resolution of the input image (first image or second image) is H×W, the image resolution of the initial optical flow field obtained in step S202 can be a low image resolution of H / 8×W / 8. In this case, the image processing device can perform interpolation upsampling processing on this initial optical flow field to obtain an upsampled optical flow field upsampled to H / 4×W / 4. It can be seen that the image resolution H / 4×W / 4 of the upsampled optical flow field is greater than that of the initial optical flow field H / 8×W / 8.
[0123] For example, if the image resolution of the input initial optical flow field is H / 4×W / 4, the image processing device can perform interpolation upsampling processing on the initial optical flow field to obtain an upsampled optical flow field upsampled to H / 2×W / 2.
[0124] For example, if the image resolution of the initial optical flow field is H / 2×W / 2, the image processing device can perform interpolation upsampling on the initial optical flow field to obtain an upsampled optical flow field upsampled to H×W. At this time, the image resolution of the obtained upsampled optical flow field is consistent with the image resolution of the input image.
[0125] It should be noted that after each optical flow update iteration, the image processing device can perform interpolation upsampling to gradually improve the optical flow field of the current resolution to a higher image resolution until it matches the image resolution of the input image.
[0126] S204. The image feature layer corresponding to the upsampled optical flow field is fused with the upsampled optical flow field to obtain the fused optical flow field.
[0127] Here, step S204 can also be called the feature fusion step.
[0128] For example, after the interpolation upsampling process is completed in step S203, the image processing device can fuse the obtained upsampled optical flow field with the image feature layer (or feature map) corresponding to the image resolution of the sampled optical flow field to obtain the fused optical flow field.
[0129] In some embodiments, the image processing apparatus can optimize the fused optical flow field through a convolution fusion operation to obtain a more accurate optical flow field.
[0130] For example, the convolution operation can be shown as in equation (4): F high =Conv(F low +U(F prev )) Formula (4)
[0131] Here, Fhigh F represents the optical flow characteristics after upsampling. low For the image feature layer corresponding to the image resolution of this sampled optical flow field, U(F) prev ) represents the interpolated optical flow field, and Conv represents the convolution fusion operation.
[0132] S205. Use the fused optical flow field as the new initial optical flow field.
[0133] Here, step S205 can also be called the iterative optimization step.
[0134] S206. Repeat the upsampling, feature fusion and iterative optimization steps (i.e., repeat steps S203-S205) until the image resolution of the new initial optical flow field obtained in the iterative optimization step reaches the image resolution of the first image, and determine the new initial optical flow field as the high-precision optical flow field from the first image to the second image.
[0135] It should be noted that in each upsampling layer, the image processing device can continuously iterate and update the optical flow field obtained from the optical flow estimation through the operations in the upsampling and feature fusion steps described above, until a full-resolution optical flow map, i.e., the high-precision optical flow field mentioned above, is generated. The final output high-precision optical flow field has the same image resolution H×W as the input image. In this way, and through layer-by-layer feature fusion, high-precision and detail-rich optical flow estimation is achieved.
[0136] In some embodiments, steps S203-S206 above can also be referred to as a recursive update process, which generates a high-precision optical flow field with the same resolution as the input image by using a method of layer-by-layer interpolation upsampling and feature fusion, thereby further improving the accuracy of optical flow estimation.
[0137] For example, the image processing device can input the initial optical flow field obtained in step S202 into the recursive update module to gradually refine the initial optical flow field. This progressively improves the image resolution and accuracy of the optical flow map.
[0138] For example, the above recursive update module may employ a gated recurrent unit (GRU) or other possible recursive network models.
[0139] It should be understood that the input recursive update module can be used to progressively refine the optical flow map from low image resolution to high image resolution. The initial input is the optical flow map at the lowest image resolution, which is the initial optical flow field obtained in step S22. For each layer in the pyramid, the input recursive update module can receive the feature map of the current layer and the hidden state output by the previous layer's input recursive update module as new input. The optical flow map obtained based on this new input can be improved to a higher resolution through interpolation or other upsampling methods. This process continues iterating until the top layer of the pyramid is reached, i.e., the image resolution of the obtained optical flow map reaches the image resolution of the initial input image (the first image or the second image).
[0140] Then, the image processing device performs reverse mapping of each pixel in the second image to the first image based on the high-precision optical flow field from the first image to the second image, so as to map each pixel in the second image to the corresponding pixel position in the first image, thereby obtaining an aligned image.
[0141] In this way, the resulting aligned image achieves precise pixel-level alignment between the first and second images, providing reliable input data for subsequent image fusion, deblurring, and other processing, thus achieving higher processing accuracy. Furthermore, this high-precision alignment method can handle complex motion scenes and geometric distortions, providing higher-quality alignment results for image fusion.
[0142] In some embodiments, an effective region mask, or mask image, can also be constructed to mark pixels that remain in the effective region during the aforementioned inverse mapping transformation.
[0143] Based on the embodiment shown in Figure 3, through high-precision image alignment and image fusion processing, pixel-level precise alignment of images can be achieved based on the optical flow estimation technique of layer-by-layer upsampling. This effectively handles complex motion scenes and geometric distortions, significantly improving the accuracy of image alignment, thereby improving the overall image clarity while effectively maintaining the fine structure and clarity of the image.
[0144] In some embodiments, as shown in FIG4, this disclosure also provides another image processing method, which can implement the image processing process of obtaining each aligned image in step S102 above, specifically including:
[0145] S301. Based on the field of view information of the first image and the field of view information of the second image, the second image is cropped to obtain the cropped second image.
[0146] Here, the field of view of the cropped second image is the same as that of the first image.
[0147] For example, the image processing device can perform a field-of-view consistency check on the first image and the second image in the acquired image pair. If the fields of view are inconsistent, preliminary alignment can be performed by scale-invariant feature transform feature matching (SIFT), random sample consensus (RANSAC) algorithm or other methods to automatically crop out the part of the second image that is inconsistent with the field of view of the first image, so as to obtain the cropped second image.
[0148] For example, the image processing device can sequentially perform field-of-view consistency checks on all image pairs in the acquired image pair set, and perform field-of-view consistency processing on the image pairs by executing step S301 if the fields of view are inconsistent. In this way, it can be ensured that the field of view of the cropped second image is consistent with the field of view of the first image, providing a high-quality initial alignment for subsequent fine processing.
[0149] S302. Based on the color and brightness information of the first image, perform color and brightness migration processing on the cropped second image to obtain the migrated second image.
[0150] For example, for a color first image and a color second image, the image processing device may perform step S302, or for a black and white first image and a black and white second image, the image processing device may perform step S303.
[0151] For example, the image processing device can convert the first image and the cropped second image to the Lab color space, or CIELab color space, respectively. Here, L represents lightness, a represents the saturation of the color on the red-green axis, and b represents the saturation of the color on the blue-yellow axis.
[0152] Furthermore, based on the Lab color space, the brightness and color of the cropped second image are adjusted according to the color characteristics of the first image. After adjustment, it is converted back to the BGR color space to obtain the second image after migration processing. This second image after migration processing is the image with the same color and brightness as the first image.
[0153] This migration process solves the problem of poor image fusion results caused by color differences in existing technologies. Furthermore, consistent color and brightness improve the performance of the image processing module in deblurring, resulting in a more natural and realistic final image.
[0154] S303. Based on the brightness information of the first image, perform brightness migration processing on the cropped second image to obtain the migrated second image.
[0155] For example, the image processing device can convert both the first image and the cropped second image to the Lab color space. Here, since the image is black and white, the brightness of the cropped second image can be adjusted only in the luminance channel L. After adjustment, it is converted back to the BGR color space to obtain the second image after migration processing. This second image after migration processing is an image with the same brightness as the first image.
[0156] Thus, for black and white images, only luminance information needs to be transferred, making the processing method simpler and more convenient.
[0157] S304. Based on the image features of the first image and the second image in the image pair, as well as optical flow estimation techniques, the second image after migration processing is aligned to the first image to obtain an aligned image.
[0158] Here, the matching degree between the aligned image and the first image is higher than the matching degree between the second image and the first image.
[0159] Here, the specific implementation process of step S304 can be referred to the relevant description in step S102 above, and will not be repeated here.
[0160] Based on the embodiment shown in Figure 4, by processing for consistent field of view and consistent color and brightness, a high-quality initial alignment can be provided for subsequent fine processing, and the problem of poor image fusion effect caused by color differences in the prior art can be solved. Furthermore, consistent color and brightness can improve the performance of the image processing module in deblurring, making the final generated image more visually natural and realistic.
[0161] In some embodiments, this disclosure also provides another image processing method, which can be combined with any of the above embodiments, as shown in FIG5. The method includes: S104, training an initial model based on the processed image pair set to obtain an image processing model.
[0162] Here, the image processing model is used for image deblurring. That is, the image processing model can deblur the input image and output an image with higher clarity than the input image.
[0163] Here, image deblurring can be used to refer to the process of processing a blurred image using a specific algorithm or technique to restore its original sharpness and image details. In this disclosure, image deblurring can also be referred to as image sharpening processing, image restoration processing, image blur removal processing, or other terms with the same or similar meanings, and this disclosure does not limit it.
[0164] In some embodiments, the initial model can be implemented using a convolutional neural network (CNN) such as a U-shaped convolutional neural network (U-Ne), a deblur model based on a generative adversarial network (DeblurGAN), or other suitable deep learning models.
[0165] For example, the image processing device can employ an end-to-end training approach, directly processing the input data (the processed first image) through the entire network without requiring manual intervention or feature extraction in intermediate stages to obtain the output data (the deblurred output image). This simplifies the model structure, improves training efficiency, and enhances model performance.
[0166] Thus, through the training process, the initial model's network can learn how to utilize the structural information of the clear image (the processed second image) to effectively repair and restore details and edge information in the blurred image (the processed first image), thereby achieving image deblurring. It should be understood that a fully trained model can automatically learn features to remove various types of blur without relying on prior knowledge or manual adjustments, quickly generating high-quality deblurred images, thereby significantly improving image processing speed and efficiency.
[0167] In some embodiments, the process of training the initial model based on the processed image pair set includes the following steps S21-S25:
[0168] S21. The image processing device acquires the initial model and the processed image pair set.
[0169] S22. The image processing device inputs the first processed image from the target image pair in the processed image pair set into the initial model to be trained, and obtains the first output image of the initial model. Here, the target image pair is any image pair in the processed image pair set.
[0170] S23. The image processing device compares the first output image output by the initial model with the processed second image after the target image is matched, and determines the image loss value.
[0171] S24. The image processing device adjusts the model parameters of the initial model based on the aforementioned loss value.
[0172] S25. The image processing device takes another image pair in the processed image pair set as the new target image pair and repeats steps S21-S25 until convergence.
[0173] Here, the other image pair can be the target image mentioned above, or it can be any image pair in the processed image pair set other than the target image pair mentioned above. This disclosure does not limit this.
[0174] Here, the image processing device may determine whether the model has converged based on the loss value output each time during the initial model training process, or based on the number of training iterations. This disclosure does not limit this to any particular method.
[0175] For example, when the loss value is less than a loss value threshold, the image processing device can determine that the initial model has converged.
[0176] For example, when the number of training iterations exceeds a threshold, the image processing device can determine that the initial model has converged.
[0177] In this way, the image processing device can identify the converged initial model as the trained image processing model.
[0178] In some embodiments, each image pair in the processed object pair set may also carry a mask or mask image associated with the first image.
[0179] A mask is typically a binary image (or grayscale image, where each pixel value represents the validity of that pixel) with the same resolution as the input image. Masks can help ignore or specifically handle inaccurate or invalid pixels in subsequent image processing.
[0180] Based on the embodiment shown in Figure 5, the image pair can be processed to improve the matching degree between the clear image (second image) and the blurred image (first image) in the image pair. In this way, the model used for deblurring can be trained using the processed image pair, so that the matching degree between the clear image obtained by the model after deblurring is higher than that between the image before processing, thereby improving the clarity, readability and overall quality of the deblurred image, and thus improving the effect of image deblurring.
[0181] Furthermore, this solution can optimize the alignment of training data and the realism and naturalness of color fusion in deep learning networks, significantly improving the detail, clarity, and overall quality of the deblurred images. Even under hardware limitations, it can achieve high-quality image output, meeting users' high standards for mobile photography. This enables superior image effects in fields such as smartphones, medical imaging, and video surveillance, enhancing the overall user experience.
[0182] In some embodiments, in conjunction with the above embodiments, the image processing method provided in this disclosure can also be described by the flowchart shown in FIG6, as shown in FIG6, including:
[0183] S31, Image Acquisition.
[0184] For example, the image processing device can acquire a set of image pairs, each image pair including a first image (blurred image) and a second image (sharp image), wherein the image sharpness of the second image is greater than that of the first image.
[0185] S32, Preliminary image alignment.
[0186] For example, the image processing device can perform cropping processing on the second image based on the field of view information of the first image and the field of view information of the second image to obtain the cropped second image, thereby achieving preliminary image alignment.
[0187] S33, Color consistency processing.
[0188] For example, the image processing apparatus may perform color consistency processing on the cropped second image based on the first image and a migration processing method to obtain the migration-processed second image.
[0189] S34, pixel-level image alignment.
[0190] For example, an image processing device can achieve pixel-level image alignment based on layer-by-layer upsampling optical flow estimation techniques to obtain an aligned image.
[0191] S35, Image Fusion.
[0192] For example, the image processing apparatus can perform image fusion processing on the aligned image and the first image to obtain a processed first image and a processed second image.
[0193] S36, Defuzzification Training.
[0194] For example, the image processing apparatus can train an initial model based on a set of processed image pairs to obtain an image processing model.
[0195] The foregoing primarily describes the solutions provided in this disclosure from the perspective of interaction between various devices or apparatuses. It is understood that each device or apparatus, in order to achieve the aforementioned functions, includes corresponding hardware structures and / or software modules for executing those functions. Those skilled in the art should readily recognize that, based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0196] Figure 7 shows a schematic diagram of the composition of an electronic device provided in an embodiment of this disclosure. This electronic device can be the image processing device described above, capable of executing the image processing method provided in this disclosure. As shown in Figure 7, the electronic device 700 includes an acquisition module 701 and a processing module 702. In some embodiments, the electronic device 700 may further include a training module 703.
[0197] Here, the acquisition module 701 is used to acquire an image pair set, which includes at least one image pair, and each image pair includes a first image and a second image, wherein the image sharpness of the second image is greater than that of the first image.
[0198] The processing module 702 is used to perform image alignment processing on each image pair in the image pair set to obtain the aligned image in each image pair.
[0199] The processing module 702 is further configured to perform image fusion processing on the aligned image and the first image in each image pair to obtain a processed image pair set; here, the processed image pair set includes at least one processed image pair, and the matching degree between the second image and the first image in the processed image pair is greater than the matching degree between the second image and the first image in the unprocessed image pair.
[0200] In some embodiments, the processing module 702 is specifically used to align the second image to the first image for each image pair in the image pair set, based on the image features of the first image, the image features of the second image, and optical flow estimation techniques, to obtain an aligned image; here, the alignment accuracy between the aligned image and the first image is higher than the alignment accuracy between the second image and the first image, and the matching degree between the aligned image and the first image is higher than the matching degree between the second image and the first image.
[0201] In some embodiments, image features include at least one of field of view, color, and brightness.
[0202] In some embodiments, the processing module 702 is specifically used to estimate a high-precision optical flow field from the first image to the second image based on the image features of the first image, the image features of the second image, and optical flow estimation techniques in the image pair; and to perform reverse mapping of each pixel in the second image to the first image based on the high-precision optical flow field from the first image to the corresponding pixel position in the first image, so as to obtain an aligned image.
[0203] In some embodiments, the processing module 702 is specifically used for:
[0204] Feature extraction is performed on the first image and the second image respectively to obtain multiple image feature layers for the first image and the second image; here, different image feature layers correspond to different image resolutions;
[0205] Initial optical flow estimation step: Based on the image feature layer with the smallest image resolution and optical flow estimation techniques, the initial optical flow field from the first image to the second image is obtained;
[0206] Upsampling step: The initial optical flow field is interpolated and upsampled to obtain the upsampled optical flow field. The image resolution of the upsampled optical flow field is greater than that of the initial optical flow field.
[0207] Feature fusion step: The image feature layer corresponding to the image resolution based on the upsampled optical flow field is fused with the upsampled optical flow field to obtain the fused optical flow field;
[0208] Iterative optimization steps: Use the fused optical flow field as the new initial optical flow field;
[0209] Repeat the upsampling step, feature fusion step, and iterative optimization step until the image resolution of the new initial optical flow field reaches the image resolution of the first image, and then determine the new initial optical flow field as the high-precision optical flow field from the first image to the second image.
[0210] In some embodiments, the processing module 702 is specifically used for:
[0211] For each image pair, the low-frequency brightness information of the first image and the aligned image is separated;
[0212] A uniform background image is generated based on the first image after removing low-frequency brightness information;
[0213] The first image after removing low-frequency brightness information and the aligned image after removing low-frequency brightness information are merged with the background image to obtain the merged first image and the merged aligned image. The merged first image is then determined as the processed first image.
[0214] The chromaticity information of the merged first image and the luminance information of the merged aligned image are fused to obtain the processed second image;
[0215] Based on the processed first image and the processed second image, a processed image pair is obtained.
[0216] In some embodiments, the processing module 702 is specifically used to: crop the second image based on the field of view information of the first image and the field of view information of the second image to obtain a cropped second image, wherein the field of view of the cropped second image is consistent with the field of view of the first image; perform color and brightness transfer processing on the cropped second image based on the color and brightness information of the first image, or perform brightness transfer processing on the cropped second image based on the brightness information of the first image to obtain a transferred second image; and align the transferred second image to the first image based on the image features of the first image, the image features of the second image, and optical flow estimation techniques in the image pair to obtain an aligned image, wherein the matching degree between the aligned image and the first image is higher than the matching degree between the second image and the first image.
[0217] In some embodiments, the image pair set includes image pairs from different acquisition scenarios, where the acquisition scenario is associated with at least one of the following:
[0218] Lens type, lens focal length, lighting conditions, camera shake, shooting angle, depth of field, and object motion.
[0219] In some embodiments, the training module 703 is used to train the initial model based on the processed image pair set to obtain an image processing model; here, the image processing model is used for image deblurring.
[0220] For a more detailed description of the acquisition module 701, processing module 702, and training module 703, as well as a more detailed description of each technical feature therein and a description of the beneficial effects, please refer to the corresponding method embodiment section above, which will not be repeated here.
[0221] It should be noted that the modules in Figure 7 can also be called units; for example, the acquisition module can be called an acquisition unit. Furthermore, in the embodiment shown in Figure 7, the names of the modules may not be those shown in the figure; for example, the acquisition module can also be called a receiving module.
[0222] If the various units or modules in Figure 7 are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this disclosure, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this disclosure. Storage media for storing computer software products include: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0223] When the functions of the integrated modules described above are implemented in hardware, this disclosure provides a schematic diagram of the structure of an electronic device, which may include the electronic device 700 described above. As shown in FIG8, the electronic device 800 includes: a processor 802, a communication interface 803, and a bus 804. In some embodiments, the electronic device 800 may further include a memory 801.
[0224] Processor 802 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 802 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 802 may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a digital signal processor (DSP), and a microprocessor.
[0225] The communication interface 803 is used to connect to other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0226] The memory 801 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0227] In some embodiments, the memory 801 may exist independently of the processor 802. The memory 801 may be connected to the processor 802 via a bus 804 and is used to store instructions or program code. When the processor 802 calls and executes the instructions or program code stored in the memory 801, it can implement the methods provided in the embodiments of this disclosure.
[0228] In other embodiments, memory 801 may also be integrated with processor 802.
[0229] Bus 804 can be an extended industry standard architecture (EISA) bus, etc. Bus 804 can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in Figure 8, but this does not mean that there is only one bus or one type of bus.
[0230] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the equipment or device can be divided into different functional modules to complete all or part of the functions described above.
[0231] This disclosure also provides a computer-readable storage medium, which includes a non-transitory computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by computer instructions instructing related hardware. The program can be stored in the above-mentioned computer-readable storage medium, and when executed, it can include the processes of the above-described method embodiments. The computer-readable storage medium can be any of the foregoing embodiments or memory. The above-mentioned computer-readable storage medium can also be an external storage device of the above-mentioned device or apparatus, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the above-mentioned device or apparatus. Further, the above-mentioned computer-readable storage medium can also include both internal storage units of the above-mentioned device or apparatus and external storage devices. The above-mentioned computer-readable storage medium is used to store the above-mentioned computer program and other programs and data required by the above-mentioned device or apparatus. The above-mentioned computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0232] This disclosure also provides a computer program product comprising a computer program that, when run on a computer, causes the computer to perform any of the methods provided in the above embodiments.
[0233] Although this disclosure has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the appended claims in carrying out the claimed disclosure. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0234] Although this disclosure has been described in conjunction with specific features and embodiments, it will be apparent that various modifications and combinations can be made therein without departing from the spirit and scope of this disclosure. Accordingly, this specification and drawings are merely exemplary illustrations of the disclosure as defined by the appended claims and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this disclosure. It is obvious that those skilled in the art can make various alterations and modifications to this disclosure without departing from its spirit and scope. Thus, this disclosure is also intended to include any such modifications and modifications that fall within the scope of the claims of this disclosure and their equivalents.
[0235] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions within the technical scope disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. An image processing method, wherein, The method comprises: acquiring a set of image pairs, the set of image pairs comprising at least one image pair, each image pair comprising a first image and a second image, the second image having a higher image definition than the first image; respectively performing image alignment processing on each image pair in the set of image pairs to obtain an aligned image in each image pair; performing image fusion processing on the aligned image in each image pair and the first image to obtain a set of processed image pairs; wherein the set of processed image pairs comprises at least one processed image pair, and the matching degree between the second image and the first image in the processed image pair is higher than the matching degree between the second image and the first image in the image pair before processing.
2. The method of claim 1, wherein, The respective image alignment processing on each image pair in the set of image pairs to obtain an aligned image in each image pair comprises: for each image pair in the set of image pairs, aligning the second image to the first image based on the image features of the first image, the image features of the second image, and an optical flow estimation technique to obtain the aligned image; wherein the alignment accuracy of the aligned image to the first image is higher than the alignment accuracy of the second image to the first image, and the matching degree of the aligned image to the first image is higher than the matching degree of the second image to the first image.
3. The method of claim 2, wherein, The image features comprise at least one of a field of view, color, and brightness.
4. The method of claim 2 or 3, wherein, The aligning the second image to the first image based on the image features of the first image, the image features of the second image, and the optical flow estimation technique to obtain the aligned image comprises: estimating a high-precision optical flow field from the first image to the second image based on the image features of the first image, the image features of the second image, and the optical flow estimation technique; performing reverse mapping of each pixel point in the second image to the first image to map each pixel point in the second image to a corresponding pixel point position in the first image to obtain the aligned image according to the high-precision optical flow field from the first image to the second image.
5. The method of claim 4, wherein, The estimating a high-precision optical flow field from the first image to the second image based on the image features of the first image, the image features of the second image, and the optical flow estimation technique comprises: performing feature extraction on the first image and the second image respectively to obtain a plurality of image feature layers of the first image and the second image; wherein different image feature layers correspond to different image resolutions; an initial optical flow estimation step: obtaining an initial optical flow field from the first image to the second image based on the image feature layer with the smallest image resolution and the optical flow estimation technique; an upsampling step: performing interpolation upsampling processing on the initial optical flow field to obtain an upsampled optical flow field, the image resolution of the upsampled optical flow field being higher than that of the initial optical flow field; a feature fusion step: performing fusion processing on the upsampled optical flow field based on the image feature layer corresponding to the image resolution of the upsampled optical flow field to obtain a fused optical flow field; an iterative optimization step: taking the fused optical flow field as a new initial optical flow field; The up-sampling step, the feature fusion step, and the iterative optimization step are repeatedly performed until the image resolution of the new initial optical flow field reaches the image resolution of the first image, and the new initial optical flow field is determined as the high-precision optical flow field from the first image to the second image.
6. The method of any one of claims 1-5, wherein, The image fusion processing of the aligned image and the first image in each image pair obtains a set of processed image pairs, including: For each image pair, the low-frequency luminance information of the first image and the aligned image is separated out; A uniform background image is generated according to the first image after removing the low-frequency luminance information; The first image after removing the low-frequency luminance information and the aligned image after removing the low-frequency luminance information are respectively merged with the background image to obtain a merged first image and a merged aligned image, and the merged first image is determined as the processed first image; The chroma information of the merged first image and the luminance information of the merged aligned image are fused to obtain the processed second image; The processed first image and the processed second image are used to obtain the set of processed image pairs.
7. The method of any one of claims 2-4, wherein, The second image is aligned to the first image based on the image features of the first image, the image features of the second image, and the optical flow estimation technology to obtain an aligned image, including: The second image is cropped based on the field of view information of the first image and the field of view information of the second image to obtain a cropped second image, and the field of view of the cropped second image is consistent with the field of view of the first image; The cropped second image is subjected to color and luminance migration processing based on the color and luminance information of the first image, or the cropped second image is subjected to luminance migration processing based on the luminance information of the first image to obtain a migrated second image; The migrated second image is aligned to the first image based on the image features of the first image, the image features of the second image, and the optical flow estimation technology to obtain an aligned image, and the matching degree of the aligned image with the first image is higher than the matching degree of the second image with the first image.
8. The method of any one of claims 1-7, wherein, The set of image pairs includes image pairs in different acquisition scenarios, and the acquisition scenario is associated with at least one of the following: lens type, lens focal length information, lighting environment, shaking degree, shooting angle, depth of field range, object motion state.
9. The method of any one of claims 1-8, wherein, The method further includes: training an initial model based on the set of processed image pairs to obtain an image processing model, wherein the image processing model is used for image deblurring processing. 10.An electronic device, wherein including: a memory and a processor; the memory and the processor are coupled; the memory is used to store instructions executable by the processor; the processor executes the instructions to perform the method of any one of claims 1-9.
11. A computer program product, wherein, The computer program product contains a computer program that, when executed on a computer, causes the computer to perform the method of any one of claims 1-9. The computer program product contains