Super-resolution image generation method and device, storage medium and electronic equipment

By extracting and stitching the depth of field information of the image, combined with super-resolution model processing, the problem of poor combination of depth of field information and perspective relationship in the prior art is solved, and the natural and real effects of high-resolution images are achieved.

CN120070183APending Publication Date: 2025-05-30BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196326.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing super-resolution technology fails to effectively combine depth of field information and perspective relationships, resulting in image distortion, blurred foreground objects or excessively clear backgrounds, and failure of composition relationships.

Method used

Extract the depth of field information of the original image, generate the depth of field image, and splice the original image and the depth of field image into four channels of data, and pass it as input to the super-resolution model, and use the depth of field information for processing to restore high-resolution images.

Benefits of technology

Effectively retain the depth of field information of the image, improve the spatial sense and perspective consistency of the image, and the generated high-resolution images are more natural and realistic, avoiding the problems of blurred foregrounds or excessively clear backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070183A_ABST
    Figure CN120070183A_ABST
Patent Text Reader

Abstract

The invention relates to a super-resolution image generation method and device, a storage medium and electronic equipment. The method comprises the steps that depth-of-field information of an original image is extracted, a depth-of-field image of the original image is obtained, and the original image is a low-resolution image; splicing the original image and the depth-of-field image to form four-channel data, and obtaining a depth-of-field fused image; and inputting the depth-of-field fused image to a target super-resolution model to enable the target super-resolution model to perform super-resolution processing on the depth-of-field fused image to obtain a target image, the target image being a high-resolution image, and the depth-of-field information of the original image being present in the target image. The technical problem of image distortion caused by the fact that the super-resolution technology cannot effectively combine the depth-of-field information and the perspective relation is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, storage medium, and electronic device for super-resolution image generation. Background Art

[0002] In the field of image super-resolution (SR) technology, the goal is to restore a high-resolution image from a low-resolution image, thereby enhancing the details and clarity of the image. Standard super-resolution methods typically rely on deep learning models, aiming to generate a higher-resolution image from a low-resolution image through techniques such as convolutional neural networks. However, although these methods have made significant progress in improving image quality and restoring details, there are still some unsolved problems, especially when dealing with the depth-of-field effect and perspective relationship between the foreground and the background. First, the foreground changes from in-focus to out-of-focus, that is, the foreground objects that were clear in the low-resolution image become blurred after super-resolution processing. Second, the background changes from out-of-focus to in-focus, and the originally blurred background area becomes abnormally clear after super-resolution, resulting in an unnatural visual effect of the image. Finally, the composition relationship fails, that is, the overall perspective relationship and depth-of-field effect in the image are lost after super-resolution, weakening the sense of space and depth of the super-resolution image. The root cause of these problems is that standard super-resolution algorithms often only focus on enhancing details and improving clarity, while ignoring the depth-of-field information and perspective relationship in the image. Depth of field refers to the relationship between the distances of different objects in the image in three-dimensional space and their relative clarity. If the super-resolution technology fails to effectively consider this point, it is easy to cause inconsistent clarity between the foreground objects and the background in the image. In addition, there are certain limitations in the design of the loss function for super-resolution. The current mainstream loss functions, such as mean squared error or mean absolute error, mainly focus on pixel-level differences and ignore the geometric information and semantic content in the image. This makes it impossible for the super-resolution network to maintain the depth information and perspective structure of the image while restoring image details, resulting in the loss of depth-of-field consistency and perspective relationship. Although both super-resolution technology and depth-of-field estimation technology have made significant progress in the field of computer vision, the combination of the two is still relatively rare. Most research focuses on improving the clarity and resolution of images, and there is still insufficient exploration of how to preserve and restore depth-of-field information during the super-resolution process. Summary of the Invention

[0003] This application provides a method, apparatus, storage medium, and electronic device for super-resolution image generation to solve the technical problem that the super-resolution technology fails to effectively combine depth-of-field information and perspective relationship, resulting in image distortion.

[0004] In a first aspect, the present application provides a method for generating a super-resolution image, including: extracting the depth-of-field information of an original image to obtain the depth-of-field image of the original image, where the original image is a low-resolution image; splicing the original image and the depth-of-field image to form four-channel data, obtaining a depth-of-field fusion image; inputting the depth-of-field fusion image into a target super-resolution model, so that the target super-resolution model performs super-resolution processing on the depth-of-field fusion image to obtain a target image, where the target image is a high-resolution image, and the depth-of-field information of the original image exists in the target image.

[0005] In a second aspect, the present application provides an apparatus for generating a super-resolution image, including: an extraction module for extracting the depth-of-field information of an original image to obtain the depth-of-field image of the original image, where the original image is a low-resolution image; a first splicing module for splicing the original image and the depth-of-field image to form four-channel data, obtaining a depth-of-field fusion image; a first processing module for inputting the depth-of-field fusion image into a target super-resolution model, so that the target super-resolution model performs super-resolution processing on the depth-of-field fusion image to obtain a target image, where the target image is a high-resolution image, and the depth-of-field information of the original image exists in the target image.

[0006] As an optional example, the apparatus further includes: an acquisition module for acquiring a training image set and an initial super-resolution model before inputting the depth-of-field fusion image into the target super-resolution model, where the training image set includes a first training image, a second training image, and a first depth-of-field training image of the first training image, where the first training image is a low-resolution image, and the second training image is the high-resolution image corresponding to the first training image; a second splicing module for splicing the first training image and the first depth-of-field training image to form four-channel data, obtaining a depth-of-field fusion training image; a second processing module for inputting the depth-of-field fusion training image into the initial super-resolution model, so that the initial super-resolution model performs super-resolution processing on the depth-of-field fusion training image to obtain a third training image, where the third training image is a high-resolution image; a calculation module for calculating a comprehensive loss value of the initial super-resolution model according to the first training image, the second training image, and the third training image; an adjustment module for adjusting the model parameters of the initial super-resolution model according to the comprehensive loss value, and recalculating the comprehensive loss value until the recalculated comprehensive loss value is less than a target threshold to obtain the target super-resolution model.

[0007] As an alternative example, the above calculation module includes: an extraction unit configured to extract the depth-of-field information of the above third training image to obtain a second depth-of-field training image of the above third training image; a first calculation unit configured to define a depth-of-field consistency loss function and calculate, according to the above depth-of-field consistency loss function, the difference between the above first depth-of-field training image and the above second depth-of-field training image to obtain a depth-of-field consistency loss value of the above initial super-resolution model; a second calculation unit configured to define a super-resolution loss function and calculate, according to the above super-resolution loss function, the difference between the above third training image and the above second training image to obtain a super-resolution loss value of the above initial super-resolution model; and a third calculation unit configured to calculate, according to the above depth-of-field consistency loss value and the above super-resolution loss value, a comprehensive loss value.

[0008] As an alternative example, the above third calculation unit includes: a configuration subunit configured to respectively configure a first weight and a second weight for the above depth-of-field consistency loss value and the above super-resolution loss value; and a calculation subunit configured to perform weighted summation on the above depth-of-field consistency loss value and the above super-resolution loss value according to the above first weight and the above second weight to obtain the above comprehensive loss value.

[0009] In a third aspect, the present application provides a storage medium storing a computer program, where the computer program, when run by a processor, executes the above method for generating a super-resolution image.

[0010] In a fourth aspect, the present application further provides an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the above method for generating a super-resolution image through the computer program.

[0011] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:

[0012] This application extracts the depth-of-field information of the original image to obtain the depth-of-field image of the original image, where the original image is a low-resolution image; stitches the original image and the depth-of-field image to form four-channel data, obtaining a depth-of-field fusion image; inputs the depth-of-field fusion image into the target super-resolution model, enabling the target super-resolution model to perform super-resolution processing on the depth-of-field fusion image to obtain a target image, where the target image is a high-resolution image and the depth-of-field information of the original image exists in the target image. In the above method, by extracting the depth-of-field information of the low-resolution original image, generating a depth-of-field image, then stitching the original image and the depth-of-field image into four-channel data and passing it as input to the super-resolution model, and finally the model uses the depth-of-field information for processing to restore the high-resolution image. During the super-resolution process, the depth-of-field information helps the model correctly handle the depth relationship between the foreground and the background, avoiding the foreground being blurred or the background being overly clear, thereby achieving the purpose of effectively retaining the depth-of-field information of the image, improving the sense of space and perspective consistency of the image, and generating a more natural and realistic high-resolution image, and further solving the technical problem that the super-resolution technology fails to effectively combine the depth-of-field information and the perspective relationship, resulting in image distortion. Description of the Drawings

[0013] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0015] One or more embodiments are illustrated by way of example in the pictures in the corresponding accompanying drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0016] Figure 1 is a flowchart of an optional method for generating a super-resolution image according to an embodiment of the present application;

[0017] Figure 2 is a schematic diagram showing the image depth of an optional method for generating a super-resolution image according to an embodiment of the present application;

[0018] Figure 3 is a schematic structural diagram of an optional device for generating a super-resolution image according to an embodiment of the present application;

[0019] Figure 4 It is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0021] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0022] According to the first aspect of the embodiments of the present application, a method for generating a super-resolution image is provided. Optionally, as Figure 1 shown, the above method includes:

[0023] S102, extracting the depth-of-field information of the original image to obtain a depth-of-field image of the original image, where the original image is a low-resolution image;

[0024] S104, splicing the original image and the depth-of-field image to form four-channel data, and obtaining a depth-of-field fusion image;

[0025] S106, inputting the depth-of-field fusion image into a target super-resolution model so that the target super-resolution model performs super-resolution processing on the depth-of-field fusion image to obtain a target image, where the target image is a high-resolution image and the depth-of-field information of the original image exists in the target image.

[0026] Optionally, in this embodiment, the original image is a low-resolution image. First, the depth of field information needs to be extracted from the low-resolution image. The depth of field information reflects the distance and clarity difference of objects in the image. Usually, the depth of field information can be obtained by using existing depth estimation algorithms, such as using deep learning models (such as convolutional neural networks) or other computer vision techniques (such as stereo vision, structured light, etc.) to estimate the depth value of each pixel in the scene. Finally, through these algorithms, a depth of field image can be obtained, which represents the depth information of each pixel in the image. After obtaining the depth of field image, the original image is stitched with the depth of field image to generate a four-channel data. Specifically, the four-channel data is composed of the RGB three channels of the original image plus the depth channel of the depth of field image. In this way, the super-resolution model can not only see the detailed information of the image but also obtain the depth information of each pixel, so that the model can retain the depth of field effect in the original image during the super-resolution process. The generated depth of field fusion image is used as the input and passed to the target super-resolution model for image processing. The target super-resolution model is a mature deep learning-based super-resolution network (such as EDSR or RCAN) that has been trained extensively. Since the input of this model is four-channel data containing additional depth of field information, the model can make full use of this information when generating high-resolution images to ensure that the depth relationship between the foreground and background of the image is correctly retained, avoiding problems such as the foreground becoming blurred or the background being overly clear. After being processed by the super-resolution model, the finally output target image is a high-resolution image. Due to the input depth of field information, the target image can better retain the depth relationship and perspective structure of the original image, and its resolution and details are enhanced.

[0027] Optionally, in this embodiment, by combining the depth of field information with super-resolution image processing, the depth relationship and perspective effect of the original image can be better retained, improving the naturalness and realism of image processing.

[0028] As an optional example, before inputting the depth of field fusion image into the target super-resolution model, the above method further includes:

[0029] Obtain a training image set and an initial super-resolution model, where the training image set includes a first training image, a second training image, and a first depth of field training image of the first training image. The first training image is a low-resolution image, and the second training image is the high-resolution image corresponding to the first training image;

[0030] Stitch the first training image with the first depth of field training image to form four-channel data, obtaining a depth of field fusion training image;

[0031] Inputting the depth of field fusion training image into the initial super-resolution model so that the initial super-resolution model performs super-resolution processing on the depth of field fusion training image to obtain a third training image, wherein the third training image is a high-resolution image;

[0032] Calculate a comprehensive loss value of an initial super-resolution model according to the first training image, the second training image, and the third training image;

[0033] The model parameters of the initial super-resolution model are adjusted according to the comprehensive loss value, and the comprehensive loss value is recalculated until the recalculated comprehensive loss value is less than the target threshold, thereby obtaining the target super-resolution model.

[0034] Optionally, in this embodiment, the optimized target super-resolution model is obtained by training and optimizing the super-resolution model, so that the target super-resolution model can generate a high-resolution image while retaining the depth of field and perspective relationship in the image. Specifically, a low-resolution image first training image and a corresponding high-resolution image second training image are collected, and an existing depth estimation algorithm is used to obtain a low-resolution depth map first depth of field training image of the first training image, such as Figure 2 The schematic diagram of the image depth is shown in the figure, with the first training image of the low-resolution image on the left and the first training image of the low-resolution depth of field image on the right. The obtained low-resolution depth of field image is saved to improve the subsequent training speed. The first training image and the first depth of field training image are spliced ​​to form a four-channel data, that is, the depth channel containing the RGB channel of the low-resolution image and the depth channel of the depth of field information. This four-channel data is the depth of field fusion training image. The depth of field fusion training image is input into the initial super-resolution model, so that the initial super-resolution model performs super-resolution processing on the image to generate a third training image, that is, a high-resolution image. Based on the first training image, the second training image and the third training image, the comprehensive loss value of the super-resolution model is calculated. The comprehensive loss value measures the difference between the high-resolution image third training image output by the super-resolution model and the real high-resolution image second training image. According to the comprehensive loss value, the parameters of the initial model are adjusted using the gradient descent method, and the comprehensive loss value is recalculated. This process is repeated until the comprehensive loss value is lower than the target threshold, and finally the optimized target super-resolution model is obtained. The target super-resolution model, which has been continuously trained and optimized, can combine depth information, better understand the depth relationship between the foreground and background in the image during super-resolution processing, and maintain the spatial sense and perspective effect of the image. It can also better adapt to image super-resolution tasks in different scenes and depth ranges, and has higher generalization ability and robustness.

[0035] Optionally, in this embodiment, by combining the depth-of-field information and the training process of the super-resolution model, the super-resolution model can better understand the depth and perspective relationship of the image, so as to retain the true depth of field and perspective relationship of the original image when generating a high-resolution image.

[0036] Obtain a training image set and an initial super-resolution model. The training image set includes: the first training image: a low-resolution image; the second training image: the corresponding high-resolution image; the first depth-of-field training image: the depth-of-field image corresponding to the first training image, representing the depth information of the objects in the image.

[0037] As an optional example, according to the first training image, the second training image, and the third training image, the calculated comprehensive loss value of the initial super-resolution model includes:

[0038] Extract the depth-of-field information of the third training image to obtain the second depth-of-field training image of the third training image;

[0039] Define a depth-of-field consistency loss function, and according to the depth-of-field consistency loss function, calculate the difference between the first depth-of-field training image and the second depth-of-field training image to obtain the depth-of-field consistency loss value of the initial super-resolution model;

[0040] Define a super-resolution loss function, and according to the super-resolution loss function, calculate the difference between the third training image and the second training image to obtain the super-resolution loss value of the initial super-resolution model;

[0041] Calculate the comprehensive loss value according to the depth-of-field consistency loss value and the super-resolution loss value.

[0042] Optionally, in this embodiment, a depth-of-field consistency loss function is designed to measure the difference between the input depth-of-field information and the output depth-of-field information, ensuring the consistency of the depth-of-field relationship during the super-resolution process. Combine the super-resolution loss and the depth-of-field consistency loss to form a comprehensive loss for optimizing the training of the super-resolution model, so that while improving the image resolution, it retains the true depth of field and perspective relationship of the original image. Specifically, when calculating the comprehensive loss value of the initial super-resolution model, extract the depth-of-field information from the third training image to obtain the second depth-of-field training image, representing the depth information of the high-resolution image after super-resolution. Define a depth-of-field consistency loss function, which is used to measure the consistency of the depth-of-field information, that is, to ensure that the depth-of-field structure of the super-resolution image is consistent with the original image depth. Use the depth-of-field consistency loss function to calculate the difference between the first depth-of-field training image (the depth-of-field information of the low-resolution image) and the second depth-of-field training image (the depth-of-field information of the image after super-resolution). This difference value is the depth-of-field consistency loss value, which is used to optimize the model to maintain the correct depth-of-field information when restoring the high-resolution image. The depth-of-field consistency loss value is shown in formula (1):

[0043]

[0044] In formula (1), L depth is the depth of field consistency loss value, D LR is the first depth of field training image, D SR is the second depth of field training image.

[0045] Define the super-resolution loss function, which is used to measure the quality of the super-resolution image, that is, the difference between the generated high-resolution image and the target high-resolution image. Use the super-resolution loss function to calculate the difference between the third training image (the high-resolution image generated by the initial super-resolution model) and the second training image (the real high-resolution image). This difference value is the super-resolution loss value, which is used to optimize the model to generate higher-quality super-resolution images. The super-resolution loss value is shown in formula (2):

[0046]

[0047] In formula (2), L Sr is the super-resolution loss value, I SR is the third training image, I HR is the second training image.

[0048] The comprehensive loss value is the weighted sum of the depth of field consistency loss value and the super-resolution loss value. By balancing these two losses, the comprehensive loss value can take into account both the super-resolution effect and the depth of field consistency of the image, and finally be used to guide the optimization of the super-resolution model.

[0049] Optionally, in this embodiment, the depth of field consistency loss value ensures that while the super-resolution image restores details, it correctly maintains the depth of field information in the image, avoiding the distortion or unnaturalness of the depth relationship in the image. The super-resolution loss value ensures that the generated high-resolution image can be as close as possible to the real high-resolution target image, improving the details and quality of the image. By combining these two loss values, the optimized model can enhance the image resolution while ensuring the consistency of the spatial structure and depth sense of the image, thereby generating more natural and real high-resolution images.

[0050] As an optional example, according to the depth of field consistency loss value and the super-resolution loss value, the calculated comprehensive loss value includes:

[0051] Configure the first weight and the second weight for the depth of field consistency loss value and the super-resolution loss value respectively;

[0052] According to the first weight and the second weight, perform weighted summation on the depth of field consistency loss value and the super-resolution loss value to obtain the comprehensive loss value.

[0053] Optionally, in this embodiment, the first weight is used to weight the depth-of-field consistency loss value. The depth-of-field consistency loss value measures whether the depth-of-field information in the image is correctly maintained during the super-resolution process. Therefore, this weight determines the importance of the depth-of-field information in the final optimization. The second weight is used to weight the super-resolution loss value. The super-resolution loss value measures the difference between the image during the resolution improvement process and the true high-resolution image. Therefore, this weight determines the importance of the image resolution improvement effect in the final optimization. According to the first weight and the second weight, the depth-of-field consistency loss value and the super-resolution loss value are weighted and summed, and the two loss values are combined according to the specified weight ratio to obtain a comprehensive loss value. By adjusting the first weight and the second weight, the model can balance the relationship between the depth-of-field consistency and the super-resolution performance as needed. If the model needs to pay more attention to retaining the depth-of-field information (for example, the depth-of-field is crucial for the naturalness and authenticity of the image), the first weight can be increased so that the depth-of-field consistency loss value dominates in the comprehensive loss. If the model needs to pay more attention to improving the image resolution and clarity (for example, the quality of image detail restoration is the most critical), the second weight can be increased so that the super-resolution loss value dominates in the comprehensive loss. During the training process, an optimal balance point can be found by continuously adjusting these two weights, so as to ensure that the depth-of-field information is fully retained while the resolution of the super-resolution image is improved.

[0054] Optionally, in this embodiment, by weighting the comprehensive loss value, the optimized super-resolution model can not only improve the clarity and resolution of the image, but also ensure that the depth-of-field information is correctly retained, thereby avoiding image distortion. According to different application scenarios or requirements, the weights of the loss function are adjusted so that the super-resolution model can flexibly optimize the quality and structural consistency of the image. By reasonably configuring the weights, the model can better process various types of images and enhance the adaptability and robustness of the model.

[0055] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0056] According to another aspect of the embodiments of the present application, there is also provided an apparatus for generating a super-resolution image, as Figure 3 shown, including:

[0057] An extraction module 302, configured to extract the depth-of-field information of the original image to obtain a depth-of-field image of the original image, where the original image is a low-resolution image;

[0058] A first splicing module 304, configured to splice the original image and the depth-of-field image to form four-channel data, thereby obtaining a depth-of-field fusion image;

[0059] A first processing module 306, configured to input the depth-of-field fusion image into a target super-resolution model, so that the target super-resolution model performs super-resolution processing on the depth-of-field fusion image to obtain a target image, where the target image is a high-resolution image and the depth-of-field information of the original image exists in the target image.

[0060] It should be noted that the extraction module 302 in this embodiment can be used to execute step S102 in the embodiment of the present application, the first splicing module 304 in this embodiment can be used to execute step S104 in the embodiment of the present application, and the first processing module 306 in this embodiment can be used to execute step S106 in the embodiment of the present application.

[0061] As an optional example, the above device further includes:

[0062] An acquisition module, configured to acquire a training image set and an initial super-resolution model before inputting the depth-of-field fusion image into the target super-resolution model, where the training image set includes a first training image, a second training image, and a first depth-of-field training image of the first training image, where the first training image is a low-resolution image and the second training image is a high-resolution image corresponding to the first training image;

[0063] A second splicing module, configured to splice the first training image and the first depth-of-field training image to form four-channel data, thereby obtaining a depth-of-field fusion training image;

[0064] A second processing module, configured to input the depth-of-field fusion training image into the initial super-resolution model, so that the initial super-resolution model performs super-resolution processing on the depth-of-field fusion training image to obtain a third training image, where the third training image is a high-resolution image;

[0065] A calculation module, configured to calculate a comprehensive loss value of the initial super-resolution model according to the first training image, the second training image, and the third training image;

[0066] An adjustment module, configured to adjust the model parameters of the initial super-resolution model according to the comprehensive loss value, and recalculate the comprehensive loss value until the recalculated comprehensive loss value is less than a target threshold, thereby obtaining a target super-resolution model.

[0067] As an optional example, the calculation module includes:

[0068] An extraction unit for extracting the depth-of-field information of the third training image to obtain the second depth-of-field training image of the third training image;

[0069] A first calculation unit for defining a depth-of-field consistency loss function and calculating the difference between the first depth-of-field training image and the second depth-of-field training image according to the depth-of-field consistency loss function to obtain the depth-of-field consistency loss value of the initial super-resolution model;

[0070] A second calculation unit for defining a super-resolution loss function and calculating the difference between the third training image and the second training image according to the super-resolution loss function to obtain the super-resolution loss value of the initial super-resolution model;

[0071] A third calculation unit for calculating a comprehensive loss value according to the depth-of-field consistency loss value and the super-resolution loss value.

[0072] As an optional example, the third calculation unit includes:

[0073] A configuration subunit for respectively configuring a first weight and a second weight for the depth-of-field consistency loss value and the super-resolution loss value;

[0074] A calculation subunit for performing weighted summation on the depth-of-field consistency loss value and the super-resolution loss value according to the first weight and the second weight to obtain a comprehensive loss value.

[0075] For other examples of this embodiment, please refer to the above examples and will not be elaborated here.

[0076] Figure 4 It is a schematic diagram of an optional electronic device according to an embodiment of the present application, as Figure 4 shown, including a processor 402, a communication interface 404, a memory 406, and a communication bus 408. Among them, the processor 402, the communication interface 404, and the memory 406 complete mutual communication through the communication bus 408. Among them,

[0077] The memory 406 is used for storing a computer program;

[0078] The processor 402, when executing the computer program stored on the memory 406, implements the following steps:

[0079] Extract the depth-of-field information of the original image to obtain the depth-of-field image of the original image, where the original image is a low-resolution image;

[0080] Stitch the original image and the depth-of-field image to form four-channel data to obtain a depth-of-field fusion image;

[0081] Input the depth-of-field fused image into the target super-resolution model, so that the target super-resolution model performs super-resolution processing on the depth-of-field fused image to obtain a target image, where the target image is a high-resolution image and the depth-of-field information of the original image exists in the target image.

[0082] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above electronic device and other devices.

[0083] The memory may include a RAM, and may also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0084] As an example, the above memory 406 may but is not limited to include the extraction module 302, the first splicing module 304, and the first processing module 306 in the above device for generating a super-resolution image. In addition, it may also include but is not limited to other module units in the above device for generating a super-resolution image, which will not be elaborated in this example.

[0085] The above processor may be a general-purpose processor, which may include but is not limited to: a CPU (Central Processing Unit), an NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processing), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0086] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and will not be elaborated herein.

[0087] Those of ordinary skill in the art can understand that Figure 4The structure shown is only schematic. The device for implementing the above method for super-resolution image generation may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a personal digital assistant, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 4 It does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 4 in, or have a different configuration from that shown Figure 4 .

[0088] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a ROM, a RAM, a magnetic disk, or an optical disc, etc.

[0089] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program, when run by a processor, executes the steps in the above method for super-resolution image generation.

[0090] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.

[0091] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0092] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0093] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0094] In several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0095] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0096] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0097] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for generating a super-resolution image, characterized in that: include: Extracting depth of field information of an original image to obtain a depth of field image of the original image, wherein the original image is a low-resolution image; The original image and the depth of field image are spliced ​​to form four-channel data to obtain a depth of field fused image; The depth of field fusion image is input into a target super-resolution model so that the target super-resolution model performs super-resolution processing on the depth of field fusion image to obtain a target image, wherein the target image is a high-resolution image and the target image contains the depth of field information of the original image.

2. The method according to claim 1, characterized in that Before inputting the depth of field fusion image into the target super-resolution model, the method further includes: Acquire a training image set and an initial super-resolution model, wherein the training image set includes a first training image, a second training image, and a first depth of field training image of the first training image, wherein the first training image is a low-resolution image, and the second training image is a high-resolution image corresponding to the first training image; splicing the first training image with the first depth of field training image to form four-channel data, thereby obtaining a depth of field fusion training image; Inputting the depth of field fusion training image into the initial super-resolution model, so that the initial super-resolution model performs super-resolution processing on the depth of field fusion training image to obtain a third training image, wherein the third training image is a high-resolution image; Calculating a comprehensive loss value of the initial super-resolution model according to the first training image, the second training image, and the third training image; The model parameters of the initial super-resolution model are adjusted according to the comprehensive loss value, and the comprehensive loss value is recalculated until the recalculated comprehensive loss value is less than a target threshold, thereby obtaining the target super-resolution model.

3. The method according to claim 2, characterized in that The step of calculating the comprehensive loss value of the initial super-resolution model according to the first training image, the second training image, and the third training image includes: Extracting depth of field information of the third training image to obtain a second depth of field training image of the third training image; Defining a depth consistency loss function, and calculating the difference between the first depth training image and the second depth training image according to the depth consistency loss function, to obtain a depth consistency loss value of the initial super-resolution model; defining a super-resolution loss function, and calculating the difference between the third training image and the second training image according to the super-resolution loss function to obtain a super-resolution loss value of the initial super-resolution model; The comprehensive loss value is calculated based on the depth of field consistency loss value and the super-resolution loss value.

4. The method according to claim 3, characterized in that The calculating the comprehensive loss value according to the depth consistency loss value and the super-resolution loss value comprises: Assign a first weight and a second weight to the depth consistency loss value and the super-resolution loss value respectively; According to the first weight and the second weight, a weighted sum is performed on the depth consistency loss value and the super-resolution loss value to obtain the comprehensive loss value.

5. A device for generating a super-resolution image, characterized in that: include: An extraction module, used to extract the depth of field information of an original image to obtain a depth of field image of the original image, wherein the original image is a low-resolution image; A first stitching module, used for stitching the original image and the depth of field image to form four-channel data and obtain a depth of field fused image; The first processing module is used to input the depth of field fusion image into a target super-resolution model so that the target super-resolution model performs super-resolution processing on the depth of field fusion image to obtain a target image, wherein the target image is a high-resolution image and the target image contains the depth of field information of the original image.

6. The device according to claim 5, characterized in that The device also includes: an acquisition module, configured to acquire a training image set and an initial super-resolution model before inputting the depth-of-field fused image into a target super-resolution model, wherein the training image set includes a first training image, a second training image, and a first depth-of-field training image of the first training image, wherein the first training image is a low-resolution image, and the second training image is a high-resolution image corresponding to the first training image; A second stitching module is used to stitch the first training image with the first depth of field training image to form four-channel data to obtain a depth of field fusion training image; a second processing module, configured to input the depth of field fusion training image into the initial super-resolution model, so that the initial super-resolution model performs super-resolution processing on the depth of field fusion training image to obtain a third training image, wherein the third training image is a high-resolution image; A calculation module, used for calculating a comprehensive loss value of the initial super-resolution model according to the first training image, the second training image and the third training image; An adjustment module is used to adjust the model parameters of the initial super-resolution model according to the comprehensive loss value, and recalculate the comprehensive loss value until the recalculated comprehensive loss value is less than a target threshold, thereby obtaining the target super-resolution model.

7. The device according to claim 6, characterized in that The calculation module comprises: an extraction unit, configured to extract the depth information of the third training image to obtain a second depth training image of the third training image; A first calculation unit is used to define a depth consistency loss function, and calculate the difference between the first depth training image and the second depth training image according to the depth consistency loss function to obtain a depth consistency loss value of the initial super-resolution model; a second calculation unit, configured to define a super-resolution loss function, and calculate a difference between the third training image and the second training image according to the super-resolution loss function to obtain a super-resolution loss value of the initial super-resolution model; The third calculation unit is used to calculate the comprehensive loss value according to the depth of field consistency loss value and the super-resolution loss value.

8. The device according to claim 7, characterized in that The third computing unit comprises: A configuration subunit, configured to configure a first weight and a second weight for the depth consistency loss value and the super-resolution loss value respectively; A calculation subunit is used to perform weighted summation of the depth consistency loss value and the super-resolution loss value according to the first weight and the second weight to obtain the comprehensive loss value.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 4 through the computer program.