Image restoration method, model training method, electronic equipment and storage medium
Patent Information
- Application Number
- CN202380010920.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively repair the problems of missing, defilement and coverage in digital images, especially during image acquisition, storage and transmission.
By evaluating image quality to be repaired, image quality evaluation information is obtained, and repair operations are performed based on this information, and images are trained and repaired using deep learning models.
Improves the accuracy and quality of image repair, and allows more efficient processing of missing and dirty areas in the image.
Smart Images

Figure CN120112938A_ABST
Abstract
Description
Image restoration method, model training method, electronic device and storage medium Technical Field
[0001] The present disclosure relates to artificial intelligence technology, in particular to the fields of computer vision technology, deep learning technology, and image restoration technology. More specifically, it relates to an image restoration method, a deep learning model training method, an electronic device, and a storage medium. Background Art
[0002] With the development of computer technology and the emergence of digital devices such as digital cameras and mobile phones, the application of digital images has become increasingly widespread. Image restoration can refer to the process of repairing digital images that have been lost, damaged, or covered during the acquisition, storage, and transmission of digital images, or during post-processing.
[0003] Summary of the Invention
[0004] In view of this, the present disclosure provides an image restoration method, a deep learning model training method, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to one aspect of the present disclosure, there is provided an image restoration method, comprising: performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information, wherein the image quality assessment information includes image quality assessment values corresponding to at least one image quality assessment item, and the image quality assessment values represent the degree of interference of the image quality assessment items on the first image to be restored; and performing a restoration operation on the first image to be restored according to the image quality assessment information to obtain a first restored image.
[0006] According to another aspect of the present disclosure, a training method for a deep learning model is provided, comprising: obtaining a first repaired sample image based on a sample image to be repaired; determining quantized sample information corresponding to the above-mentioned sample image to be repaired, wherein the above-mentioned sample image to be repaired includes at least one sample image area to be repaired, and the above-mentioned quantized sample information includes a quantized value corresponding to each of the above-mentioned at least one sample image area to be repaired, and the quantized value corresponding to the above-mentioned sample image area to be repaired represents the importance of the above-mentioned sample image area to be repaired; and, using the above-mentioned first repaired sample image, a predetermined sample image corresponding to the above-mentioned sample image to be repaired, and the quantized sample information corresponding to the above-mentioned sample image to be repaired, to train a deep learning model to obtain a first image repair model.
[0007] According to another aspect of the present disclosure, an image restoration method is provided, comprising: obtaining a second image to be restored; and inputting the second image to be restored into a first image restoration model to obtain a second restored image; wherein the first image restoration model is trained using a deep learning model training method.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more instructions, wherein when the one or more instructions are executed by the one or more processors, the one or more processors implement the method described in the present disclosure.
[0009] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which executable instructions are stored. When the executable instructions are executed by a processor, the processor implements the method described in the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided. The computer program product includes computer-executable instructions. When the computer-executable instructions are executed, they are used to implement the method described in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0012] FIG1 schematically illustrates a system architecture to which an image restoration method and a deep learning model training method according to an embodiment of the present disclosure can be applied;
[0013] FIG2 schematically shows a flow chart of an image restoration method according to an embodiment of the present disclosure;
[0014] FIG3 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to an embodiment of the present disclosure;
[0015] FIG4 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0016] FIG5 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0017] FIG6 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0018] FIG7 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0019] FIG8 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0020] FIG9 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0021] FIG10 schematically shows an example of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure;
[0022] FIG11 schematically shows an example of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to an embodiment of the present disclosure;
[0023] FIG12A schematically shows an example of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to an embodiment of the present disclosure;
[0024] FIG12B schematically illustrates an example of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to another embodiment of the present disclosure;
[0025] FIG12C schematically illustrates an example of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to another embodiment of the present disclosure;
[0026] FIG13 schematically shows an example diagram of an image restoration process according to an embodiment of the present disclosure;
[0027] FIG14 schematically shows a flow chart of a method for training a deep learning model according to an embodiment of the present disclosure;
[0028] FIG15 schematically illustrates an example of a process of training a first deep learning model using a first restoration sample image and a first predetermined sample image corresponding to a first sample image to be restored to obtain a first image restoration model according to an embodiment of the present disclosure;
[0029] FIG16A schematically shows an example schematic diagram of performing a first quantization operation on a first edge sample image to obtain first quantized sample information corresponding to a first sample image to be repaired according to an embodiment of the present disclosure;
[0030] FIG16B schematically shows an example schematic diagram of performing a first quantization operation on a first edge sample image to obtain first quantized sample information corresponding to a first sample image to be repaired according to another embodiment of the present disclosure;
[0031] FIG16C schematically shows an example diagram of first quantized sample information according to an embodiment of the present disclosure;
[0032] FIG17A schematically shows an example of a process of determining first quantized sample information corresponding to a first sample image to be restored according to another embodiment of the present disclosure;
[0033] FIG17B schematically shows an example of a process of determining first quantized sample information corresponding to a first sample image to be restored according to another embodiment of the present disclosure;
[0034] FIG18 schematically shows an example of a process of performing a restoration operation on a first sample image to be restored according to first sample image quality assessment information to obtain a first restored sample image according to an embodiment of the present disclosure;
[0035] FIG19 schematically shows an example of a process of obtaining a first reference sample image according to an embodiment of the present disclosure;
[0036] FIG20A schematically shows an example diagram of a deep learning model training process according to an embodiment of the present disclosure;
[0037] FIG20B schematically shows an example diagram of a deep learning model training process according to another embodiment of the present disclosure;
[0038] FIG20C schematically illustrates an example diagram of a deep learning model training process according to another embodiment of the present disclosure;
[0039] FIG21 schematically shows an example diagram of an image quality assessment model training process according to an embodiment of the present disclosure;
[0040] FIG22 schematically shows a flow chart of an image restoration method according to another embodiment of the present disclosure;
[0041] FIG23 schematically shows a flow chart of a method for training a deep learning model according to another embodiment of the present disclosure;
[0042] FIG24 schematically shows a flow chart of an image restoration method according to another embodiment of the present disclosure;
[0043] FIG25 schematically shows a block diagram of an image restoration apparatus according to an embodiment of the present disclosure;
[0044] FIG26 schematically shows a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure;
[0045] FIG27 schematically shows a block diagram of an image restoration apparatus according to another embodiment of the present disclosure;
[0046] FIG28 schematically shows a block diagram of a training apparatus for a deep learning model according to another embodiment of the present disclosure;
[0047] FIG29 schematically shows a block diagram of an image restoration apparatus according to another embodiment of the present disclosure; and
[0048] Figure 30 schematically shows a block diagram of an electronic device suitable for implementing an image restoration method and a deep learning model training method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0049] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0050] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0051] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0052] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0053] In the embodiments of the present disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of the data involved (for example, including but not limited to user personal information) all comply with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken with respect to user personal information to prevent unauthorized access to user personal information data and to maintain the security of user personal information, network security, and national security.
[0054] In the embodiments of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0055] For example, after obtaining the first image to be restored, your information can be desensitized using methods including de-identification or anonymization to protect the security of your information.
[0056] To at least partially address the technical problems existing in the related art, the present disclosure provides an image restoration solution. For example, an image quality assessment operation is performed on a first image to be restored to obtain image quality assessment information. The image quality assessment information includes an image quality assessment value corresponding to each of at least one image quality assessment items, each of which represents the degree to which the image quality assessment item interferes with the first image to be restored. A restoration operation is performed on the first image to be restored based on the image quality assessment information to obtain a first restored image.
[0057] Figure 1 schematically illustrates a system architecture to which an image restoration method and a deep learning model training method according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 is merely an example of a system architecture to which the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, and does not imply that the present disclosure cannot be applied to other devices, systems, environments, or scenarios.
[0058] As shown in FIG1 , the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0059] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0060] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0061] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received user requests and other data, and feed back processing results (e.g., web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0062] It should be noted that the training method of the deep learning model provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the training device of the deep learning model provided in the embodiment of the present disclosure can generally be set in the server 105. The training method of the deep learning model provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the training device of the deep learning model provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0063] Alternatively, the training method of the deep learning model provided in the embodiment of the present disclosure may also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or may also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the training apparatus of the deep learning model provided in the embodiment of the present disclosure may also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0064] It should be noted that the image restoration method provided in the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or can also be executed by a terminal device other than the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the image restoration apparatus provided in the embodiments of the present disclosure can also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or can be provided in a terminal device other than the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0065] Alternatively, the image restoration method provided in the embodiment of the present disclosure may also be executed by the server 105. Accordingly, the image restoration apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The image restoration method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the image restoration apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0066] It should be understood that the number of first terminal devices, second terminal devices, third terminal devices, networks and servers in Figure 1 is merely illustrative and any number of first terminal devices, second terminal devices, third terminal devices, networks and servers may be provided as required.
[0067] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0068] FIG2 schematically shows a flow chart of an image restoration method according to an embodiment of the present disclosure.
[0069] As shown in FIG. 2 , the image restoration method 200 may include S210 to S220 .
[0070] In operation S210, an image quality assessment operation is performed on the first image to be repaired to obtain image quality assessment information.
[0071] In operation S220, a restoration operation is performed on the first image to be restored according to the image quality evaluation information to obtain a first restored image.
[0072] According to an embodiment of the present disclosure, the image quality assessment information includes an image quality assessment value corresponding to each of at least one image quality assessment items. The image quality assessment value represents the degree of interference of the image quality assessment item on the first image to be restored.
[0073] According to an embodiment of the present disclosure, the first image to be repaired may refer to an image that requires an image repair operation. The first image to be repaired may include at least one of the following: a dynamic image to be repaired and a static image to be repaired. The dynamic image to be repaired may include at least one of the following: a dynamic text image to be repaired and a dynamic non-text image to be repaired. The dynamic text image to be repaired may include at least one of the following: a dynamic document text image to be repaired and a dynamic scene text image to be repaired. The dynamic non-text image to be repaired may refer to a video frame image. The static image to be repaired may include at least one of the following: a static text image to be repaired and a static non-text image to be repaired. The static text image to be repaired may include at least one of the following: a static document text image to be repaired and a static scene text image to be repaired.
[0074] According to embodiments of the present disclosure, a document text image may refer to a text image with a neat layout, controlled lighting, and a relatively simple background. A scene text image may refer to a text image with a relatively complex background, diverse text forms, and uncontrolled lighting. Text forms may include at least one of the following: text color, size, font, orientation, and irregular layout. Irregular layout may include at least one of bending, tilting, wrinkling, deformation, and incompleteness.
[0075] According to an embodiment of the present disclosure, after obtaining the first image to be repaired, an image quality assessment may be performed on the first image to be repaired. Image Quality Assessment (IQA) may refer to a process for characterizing the quality of an image by analyzing and studying the characteristics of the image. For example, an image quality assessment method may include at least one of the following: a subjective image quality assessment method and an objective image quality assessment method. For example, a subjective image quality assessment method may refer to a process in which a person is used as an observer to subjectively evaluate an image so as to truly reflect the person's visual perception. An objective image quality assessment method may refer to a process in which a certain mathematical model is used to reflect the subjective perception of the human eye and to obtain a result based on digital calculation.
[0076] According to an embodiment of the present disclosure, a subjective image quality evaluation method may include at least one of the following: an absolute image quality evaluation method and a relative image quality evaluation method. For example, an absolute image quality evaluation method may refer to a process in which an observer evaluates the absolute quality of an image according to certain specific evaluation performances based on his or her own knowledge and understanding. The observer may use a double stimulus continuous quality grading method (DSCQS) to evaluate the first image to be repaired with reference to the original image to obtain an image quality evaluation value. For example, the first image to be repaired and the original image may be played alternately for a certain period of time to the observer according to certain rules, and then a certain time interval may be left after the playback for the observer to score, and finally all the given scores may be averaged as the evaluation value of the sequence, i.e., the image quality evaluation value of the first image to be repaired.
[0077] According to an embodiment of the present disclosure, a relative image quality assessment method may refer to a process in which an observer compares a batch of first images to be restored, thereby determining the order of excellence for each first image to be restored, and providing a corresponding image quality assessment value. The observer may employ a single stimulus continuous quality evaluation (SSCQE) method on the first images to be restored to obtain the image quality assessment value. For example, the batch of first images to be restored may be played in a certain sequence, and the observer may provide an image quality assessment value corresponding to each first image to be restored while viewing the batch of first images to be restored.
[0078] According to an embodiment of the present disclosure, an objective image quality assessment method may include at least one of the following: a full-reference (FR) image quality assessment method, a reduced-reference (RR) image quality assessment method, and a no-reference (NR) image quality assessment method. For example, the full-reference image quality assessment method may refer to a process of comparing the difference between an ideal image and a first image to be restored, analyzing the degree of distortion of the first image to be restored, and thereby obtaining an image quality assessment value of the first image to be restored, while selecting an ideal image as a reference image. The full-reference image quality assessment method may include at least one of the following: an image quality assessment method based on pixel statistics, an image quality assessment method based on information theory, and an image quality assessment method based on structural information.
[0079] According to an embodiment of the present disclosure, a pixel statistics-based image quality assessment method can measure the quality of the first image to be restored from a statistical perspective by calculating the difference between the grayscale values of pixels corresponding to the first image to be restored and the reference image. For example, the pixel statistics-based image quality assessment method can include at least one of the following: an image quality assessment method based on Peak-Signal to Noise Ratio (PSNR) and an image quality assessment method based on Mean Square Error (MSE).
[0080] According to an embodiment of the present disclosure, an image quality assessment method based on information theory can measure the quality of the first image to be restored by calculating the mutual information between the first image to be restored and a reference image. For example, the image quality assessment method based on information theory can include at least one of the following: an image quality assessment method based on the Information Fidelity Criterion (IFC) or an image quality assessment method based on Visual Information Fidelity (VIF).
[0081] According to an embodiment of the present disclosure, a method for image quality assessment based on structural information can construct a structural similarity between the first image to be repaired and the reference image by correlating the pixels of the first image to be repaired and the reference image, thereby measuring the quality of the first image to be repaired based on the structural similarity. For example, the method for image quality assessment based on structural information can refer to an image quality assessment method based on structural similarity (SSIM). In this case, the larger the image quality assessment value (i.e., the SSIM value), the better the quality of the first image to be repaired.
[0082] According to an embodiment of the present disclosure, a partial reference image quality assessment method may refer to a process of extracting partial feature information from an ideal image and a first image to be restored, and then performing a comparative analysis of the partial feature information of the first image to be restored using the partial feature information of the ideal image as a reference to obtain an image quality assessment value. For example, the partial reference image quality assessment method may include at least one of the following: an image quality assessment method based on original image features, a quality assessment method based on a digital watermark image, and an image quality assessment method based on a Wavelet domain statistical model.
[0083] According to embodiments of the present disclosure, a non-reference image quality assessment method may refer to a process of making an assumption about the characteristics of an ideal image, establishing a corresponding mathematical analysis model based on the assumption, and calculating the characteristic representation of the first image to be restored in the mathematical analysis model to obtain an image quality assessment value. For example, the non-reference image quality assessment method may include at least one of the following: a mean-based image quality assessment method, a standard deviation-based image quality assessment method, and an average gradient-based image quality assessment method.
[0084] According to an embodiment of the present disclosure, the image quality evaluation item is used to quantitatively evaluate the first image to be restored. For example, the image quality evaluation item may include at least one of the following: image blur (i.e., Blur), scaling (i.e., Resize), noise level (i.e., Noise), and compression (i.e., JPEG).
[0085] Alternatively, the image quality assessment item may further include at least one of the following: a full reference assessment item, a partial reference assessment item, and a no reference assessment item. For example, the full reference assessment item may be used to compare the first image to be restored with all reference images, the partial reference assessment item may be used to compare the first image to be restored with some reference images, and the no reference assessment item may be used to directly assess the first image to be evaluated.
[0086] For example, the full-reference evaluation item may include at least one of the following: Mean Absolute Error (MAE), Mean Squared Error (MSE), Universal Quality Index (UQI), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), Multi-scale Structure Similarity Index Measure (MS-SSIM), Learned Perceptual Image Patch Similarity (LPIPS), and Fréchet Inception Distance (FID). The semi-reference evaluation item may include at least one of the following: Border Pixel Error (BPE). The no-reference evaluation item may include at least one of the following: Inception Score (IS) and Modified Inception Score (MIS).
[0087] According to an embodiment of the present disclosure, after obtaining a first image to be repaired, the above-mentioned image quality assessment method can be used for each image quality assessment item in at least one image quality assessment item, and an image quality assessment operation can be performed on the first image to be repaired based on the image quality assessment item to obtain an image quality assessment value corresponding to the image quality assessment item. The image quality assessment value can be used to characterize the degree of interference of the image quality assessment item on the first image to be repaired. After obtaining the image quality assessment value corresponding to each image quality assessment item, image quality assessment information can be determined based on the image quality assessment value corresponding to each image quality assessment item. For example, the image quality assessment value corresponding to each at least one image quality assessment item may include: an image quality assessment value corresponding to each image quality assessment item in the at least one image quality assessment item. Alternatively, an image quality assessment value corresponding to each image quality assessment item in the at least one image quality assessment item. Alternatively, an image quality assessment value corresponding to some of the image quality assessment items in the at least one image quality assessment item.
[0088] According to embodiments of the present disclosure, after obtaining image quality assessment information, image restoration can be performed on the first image to be restored based on the image quality assessment information. For example, image restoration can refer to the process of predicting the content of a missing region based on the content of a known region of the first image to be restored using an appropriate algorithm to obtain a first restored image. For example, image restoration methods can include at least one of the following: a diffusion-based image restoration method, a block-based image restoration method, and a deep learning-based image restoration method.
[0089] According to an embodiment of the present disclosure, a diffusion-based image restoration method may be used to propagate known information at the edge of the area to be restored to the area to be restored through a propagation mechanism for small-sized deletions such as text, scratches, and noise in the first image to be restored, thereby completing the image restoration process by completing the content of the missing area. The diffusion-based image restoration method may include at least one of the following: a variational image restoration method based on a geometric image model and an image restoration method based on partial differentials. For example, a variational image restoration method based on a geometric image model may refer to a process of approximating the image restoration problem as a variational problem of finding the extreme value of a functional by establishing a mathematical model, that is, by utilizing functionals and total variation theory to establish a data model based on the known information of the first image to be restored.
[0090] According to an embodiment of the present disclosure, a partial differential-based image restoration method may refer to a process of smoothly propagating pixels in a known region of a first image to be restored to a missing region using partial differential equations in mathematics or physics to restore the first image to be restored. For example, the partial differential-based image restoration method may include at least one of the following: an image restoration method based on a BSCB (Bertalmio-Sapiro-CasellesBallester) model, an image restoration method based on a total variation (TV) model, and an image restoration method based on a curvature driven diffusion (CDD) model.
[0091] According to an embodiment of the present disclosure, a block-based image restoration method may refer to a process of repairing the first image to be restored by calculating and searching for an image block with the highest similarity between a missing region and a known region in the first image to be restored, and copying and pasting the image block into the missing region. For example, the block-based image restoration method may include at least one of the following: a non-parametric texture synthesis image restoration method based on a Markov random field, an image restoration method based on a PatchMatch algorithm, and an image restoration method based on a Criminisi algorithm. For example, the image restoration method based on the PatchMatch algorithm may iteratively update the missing region by continuously searching for image blocks with similar textures to the missing region from the valid region in the first image to be restored, and filling the missing region with the image blocks. The image restoration method based on the Criminisi algorithm may perform priority calculation on the missing region in the first image to be restored, then propagate the structure and texture information of the known region to the missing region, and then update the confidence of each region, and repeat the above steps until the image restoration is completed.
[0092] According to embodiments of the present disclosure, a deep learning-based image restoration method may refer to a process of extracting deep features of a first image to be restored through a stacked network and restoring the first image to be restored based on the deep features. For example, the deep learning-based image restoration method may include at least one of the following: a unit image restoration method and a multivariate image restoration method.
[0093] According to an embodiment of the present disclosure, a unit image restoration method may refer to a process of restoring a single first image to be restored to obtain a single first restored image. For example, the unit image restoration method may include at least one of the following: an image restoration method based on an encoder and decoder (i.e., Encoder-Decoder) model, an image restoration method based on a U-Net model, an image restoration method based on a Generative Adversarial Network (GAN), and an image restoration method based on a Transformer (i.e., Transformer) model.
[0094] According to an embodiment of the present disclosure, a multivariate image restoration method may refer to a process of restoring a single first image to be restored to obtain multiple first restored images. For example, the multivariate image restoration method may include at least one of the following: an image restoration method based on a variational autoencoder (VAE) model, an image restoration method based on a convolutional variational autoencoder (CVAE) model, and an image restoration method based on a bidirectional transformer (i.e., bidirectional transformer) model.
[0095] According to the embodiments of the present disclosure, because the image quality assessment value included in the image quality assessment information can represent the degree of interference of the image quality assessment item on the first image to be repaired, the image quality assessment information can provide guidance for the restoration of the first image to be repaired. Furthermore, by repairing the first image to be repaired based on the image quality assessment information, the accuracy of the image restoration can be improved, thereby improving the quality of the first repaired image.
[0096] According to an embodiment of the present disclosure, performing a restoration operation on the first image to be restored according to image quality assessment information to obtain the first restored image may include the following operations.
[0097] An image quality assessment vector is obtained according to the image quality assessment information, and a restoration operation is performed on the first image to be restored according to the image quality assessment vector to obtain a first restored image.
[0098] According to an embodiment of the present disclosure, after obtaining image quality assessment information, a first number of image quality assessment values corresponding to each of at least one image quality assessment items in the image quality assessment information can be mapped to a second number of first image quality assessment values. Based on this, an image quality assessment vector can be determined based on the second number of first image quality assessment values. After obtaining the image quality assessment vector, the above-described image restoration method can be used to restore the first image to be restored based on the image quality assessment vector, thereby obtaining a first restored image.
[0099] According to the embodiment of the present disclosure, the second number can be set according to actual business needs and is not limited here. For example, the second number can be set to 64. On this basis, the first number of image quality evaluation values is w 11 、w 12 、…、w 1n For example, if the first number is n1, then n1 image quality assessment values can be mapped to 64 first image quality assessment values s1, s2, ..., s64 .
[0100] According to the embodiments of the present disclosure, since the image quality assessment vector is obtained based on the image quality assessment information, by repairing the first image to be repaired based on the image quality assessment vector, the quality of the first repaired image can be improved and the repair effect can be enhanced.
[0101] According to an embodiment of the present disclosure, obtaining an image quality assessment vector according to image quality assessment information may include the following operations.
[0102] Performing a dimension transformation operation on the image quality assessment information to obtain an image quality assessment vector, wherein the number of dimensions of the image quality assessment vector is determined according to the first image to be repaired.
[0103] According to an embodiment of the present disclosure, the number of dimensions of an image evaluation vector can be determined based on the first image to be restored. This number of dimensions is determined as a second number. Based on this second number, the first number of image quality assessment values is input into a fourth deep learning model to obtain a second number of first image quality assessment values. Based on this, the second number of first quality assessment values can be input into a fifth deep learning model to obtain an image quality assessment vector.
[0104] According to the embodiments of the present disclosure, the fourth deep learning model and the fifth deep learning model can be set according to actual business needs and are not limited here. For example, the fourth deep learning model can be a fully connected (FC) layer. The fifth deep learning model can be a split model.
[0105] According to an embodiment of the present disclosure, since the number of dimensions of the image quality assessment vector is determined based on the first image to be repaired, by performing dimensionality transformation on the image quality assessment information based on the number of dimensions, an image quality assessment vector used as a guide for image repair can be obtained, which is beneficial to improving the accuracy of the image repair method.
[0106] FIG3 schematically shows an example of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to an embodiment of the present disclosure.
[0107] As shown in FIG3 , in step 300 , an image quality assessment operation can be performed on a first image to be restored 303 to obtain image quality assessment information 301. A dimension transformation operation is performed on image quality assessment information 301 to obtain an image quality assessment vector 302. A restoration operation is performed on the first image to be restored 303 based on image quality assessment vector 302 to obtain a first restored image 304.
[0108] According to an embodiment of the present disclosure, when 1<e≤E, performing a restoration operation on the first image to be restored according to the image quality evaluation vector to obtain the first restored image may include the following operations.
[0109] Based on the (e-1)th first intermediate feature map, obtain the (e)th second intermediate feature map, where the first fused feature map is obtained based on the (e-1)th first intermediate feature map and the image quality assessment vector, and the (e)th first intermediate feature map is obtained based on the first image to be restored. A first fusion operation is performed on the (e)th second intermediate feature map and the image quality assessment vector to obtain the (e)th first intermediate feature map. Based on the (e)th first intermediate feature map, obtain the first restored image. E is an integer greater than 1, and e is an integer greater than or equal to 1 and less than or equal to E.
[0110] According to an embodiment of the present disclosure, a feature extraction operation can be performed on the first image to be repaired to obtain a first first intermediate feature map. For example, the first image to be repaired can be input into the sixth deep learning model to obtain a first first intermediate feature map. The sixth deep learning model can be configured according to actual business needs and is not limited here. For example, the sixth deep learning model can be a convolutional neural network (CNN). In this case, the first image to be repaired can be converted into a series of first intermediate feature maps through layer-by-layer convolution operations and pooling operations. Each first intermediate feature map can correspond to the output of a layer of the convolutional neural network, and is used to capture the abstract features of the first image to be repaired at different levels.
[0111] According to the embodiments of the present disclosure, by processing the (e-1)th first intermediate feature map to obtain the (e)th second intermediate feature map, and fusing the (e)th second intermediate feature map with the image quality assessment vector to obtain the (e)th first intermediate feature map, the image quality assessment vector and each first intermediate feature map can be effectively utilized. On this basis, the first restored image is obtained based on the (e)th first intermediate feature map, which is conducive to improving the quality and accuracy of the first restored image.
[0112] According to an embodiment of the present disclosure, performing a first fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map may include the following operations.
[0113] Perform a first channel fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map.
[0114] According to an embodiment of the present disclosure, after obtaining the e-th second intermediate feature map and the image quality assessment vector, the first channel fusion method can be used to perform a first channel fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map.
[0115] According to an embodiment of the present disclosure, the first channel fusion operation can be used to integrate and fuse information between different feature channels. The first channel fusion method can be configured according to actual business needs and is not limited here. For example, the first channel fusion method can include at least one of the following: channel concatenation, channel attention, channel-wise convolution, and channel-wise pooling.
[0116] According to the embodiments of the present disclosure, since the e-th first intermediate feature map is obtained by performing a channel fusion operation on the e-th second intermediate feature map and the image quality assessment vector, it is beneficial to better utilize the information in the image quality assessment vector while retaining the useful information in the e-th second intermediate feature map, which is beneficial to guiding the subsequent image restoration process, thereby improving the quality of the restored image.
[0117] According to an embodiment of the present disclosure, the image quality assessment vector includes image quality assessment dimension information corresponding to F dimensions, the number of channels of the e-th second intermediate feature map includes F, and the number of channels of the first intermediate feature map includes F, where F is an integer greater than or equal to 1.
[0118] According to an embodiment of the present disclosure, performing a first channel fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map may include the following operations.
[0119] Multiply the f-th channel of the e-th second intermediate feature map by the f-th image quality assessment dimension information to obtain the f-th channel of the e-th first intermediate feature map, where the f-th image quality assessment dimension information represents the image quality assessment dimension information corresponding to the f-th dimension.
[0120] According to an embodiment of the present disclosure, for the fth channel among the F channels of the eth second intermediate feature map, the fth image quality assessment dimension information corresponding to the fth dimension can be determined. On this basis, the fth channel of the eth second intermediate feature map can be multiplied by the fth image quality assessment dimension information to obtain the fth channel of the eth first intermediate feature map.
[0121] According to an embodiment of the present disclosure, the e-1th first intermediate feature map can be input into the seventh deep learning model to obtain the eth second intermediate feature map. The seventh deep learning model can be configured according to actual business needs and is not limited here. For example, the seventh deep learning model may include at least one of the following: Residual Network (ResNet), Wider Residual Network (Wider-ResNet), Dilated Residual Network (Dilated ResNet), Residual Network eXtension (ResNeXt), Squeeze-and-Excitation Network (SENet), Residual Network Based on Nested-ResNet (ResNeSt) and Selective Kernel Convolution (SKNet) Network.
[0122] According to an embodiment of the present disclosure, the first image to be repaired can be input into the fifth deep learning model to obtain image quality assessment dimension information corresponding to each of the F dimensions. The fifth deep learning model can be configured according to actual business needs and is not limited here. For example, the fifth deep learning model may include a fully connected layer model and a Split model. The fully connected layer model and the Split model can split the mapped image quality assessment dimension information corresponding to each of the F dimensions one by one and input them into the backbone network, and multiply them with the e-th second intermediate feature map output by the fourth deep learning model at the channel layer.
[0123] Among them, GMF (Or ′, s ′) represents the fth channel of the first intermediate feature map of the eth, Or ′ i′ Characterize the fth channel of the eth second intermediate feature map, s′ i′ represents the f-th image quality assessment dimension information, and Π() represents the multiplication operation.
[0124] According to an embodiment of the present disclosure, the fth channel of the eth second intermediate feature map is multiplied by the fth image quality assessment dimension information to obtain the fth channel of the eth first intermediate feature map. The multiplication operation can enhance the relevant features of the fth channel in the eth second intermediate feature map. Since the fth image quality assessment dimension information represents the image quality assessment dimension information corresponding to the fth dimension, the fth image quality assessment dimension information can be better reflected in the fth channel of the eth first intermediate feature map, thereby improving the accuracy and effectiveness of the channel.
[0125] According to an embodiment of the present disclosure, obtaining the e-th second intermediate feature map according to the e-1-th first intermediate feature map may include the following operations.
[0126] According to the e-1th first intermediate feature map, the eth third intermediate feature map is obtained. According to the e-1th first intermediate feature map and the eth third intermediate feature map, the eth second intermediate feature map is obtained.
[0127] According to an embodiment of the present disclosure, taking ResBlock, a basic module in a residual network, as an example, ResBlock can include two convolutional layers and a skip connection (i.e., Skip Connection). Skip connection can add input data directly to the output of the convolutional layer, so that the output contains the input information, preventing the gradient from decaying or exploding layer by layer during the back propagation process. ResBlock optimizes network parameters by learning the residual (i.e., the difference between input and output), thereby better approximating the objective function.
[0128] According to an embodiment of the present disclosure, ResBlock may include Basic Block and Bottleneck Block. For example, a residual network with a number of network layers less than a predetermined threshold may include Basic Block. Alternatively, a residual network with a number of network layers greater than or equal to a predetermined threshold may include Bottleneck Block.
[0129] According to an embodiment of the present disclosure, the network structure of the residual network may include an input part, an intermediate convolution part, and an output part. By changing the number of blocks and parameters of the intermediate convolution part, different network structures of the residual network can be obtained.
[0130] According to the embodiments of the present disclosure, since the e-th second intermediate feature map is obtained based on the e-1-th first intermediate feature map and the e-th third intermediate feature map, and the e-th third intermediate feature map is obtained based on the e-1-th first intermediate feature map, through layer-by-layer processing, the feature map of the previous layer and the feature map of the current layer can be better utilized, thereby transferring information between each layer, which is conducive to improving the accuracy of image restoration.
[0131] According to an embodiment of the present disclosure, obtaining the e-th third intermediate feature map according to the e-1-th first intermediate feature map may include the following operations.
[0132] Perform a first conversion operation on the e-1th first intermediate feature map to obtain an eth fourth intermediate feature map. Perform a first channel grouping operation on the eth fourth intermediate feature map to obtain multiple eth fifth intermediate feature maps. Perform a second conversion operation on the multiple eth fifth intermediate feature maps to obtain multiple eth sixth intermediate feature maps. Perform a second fusion operation on the multiple eth sixth intermediate feature maps to obtain an eth third intermediate feature map.
[0133] According to an embodiment of the present disclosure, the first conversion operation can be used to extract, convert and transmit key feature information from the e-1th first intermediate feature map. The second conversion operation can be used to extract, convert and transmit key feature information from the eth fifth intermediate feature map. The specific first conversion operation and the second conversion operation can be configured according to actual business needs and are not limited here. For example, the first conversion method and the second conversion method may include at least one of the following: convolution (i.e., Convolution), pooling (i.e., Pooling), upsampling (i.e., Upsampling), downsampling (i.e., Downsampling), transposed convolution (i.e., Transpose Convolution) and normalization (i.e., Normalization).
[0134] According to an embodiment of the present disclosure, the first channel grouping operation can be used to group the e-th fourth intermediate feature map channel. By dividing the e-th fourth intermediate feature map channel into multiple groups, the diversity and generalization ability of the features can be improved. The specific first channel grouping method can be configured according to actual business needs and is not limited here. For example, the first channel grouping method may include at least one of the following: Depthwise Grouped Convolution, Pointwise Grouped Convolution, and Grouped Convolution.
[0135] According to an embodiment of the present disclosure, the second fusion operation can be used to merge or combine multiple e-th sixth intermediate feature maps into an e-th third intermediate feature map, which can integrate information of different scales, different levels, and different representations. The specific second fusion method can be configured according to actual business needs and is not limited here. For example, the second fusion method may include at least one of the following: direct addition, weighted addition, element-by-element multiplication, splicing, attention mechanism, and skip connection.
[0136] According to an embodiment of the present disclosure, after obtaining the e-1th first intermediate feature map, the first conversion method described above can be used to perform a first conversion operation on the e-1th first intermediate feature map to obtain an e-th fourth intermediate feature map. The first channel grouping method described above can be used to perform a first channel grouping operation on the e-th fourth intermediate feature map to obtain multiple e-th fifth intermediate feature maps. The second conversion method described above can be used to perform a second conversion operation on the multiple e-th fifth intermediate feature maps to obtain multiple e-th sixth intermediate feature maps. Based on this, the second fusion method described above can be used to perform a second fusion operation on the multiple e-th sixth intermediate feature maps to obtain an e-th third intermediate feature map.
[0137] According to an embodiment of the present disclosure, by performing a first conversion operation on the e-1th first intermediate feature map to obtain the e-4th intermediate feature map, the representation of the feature map can be changed, thereby extracting features with greater discriminative power. By performing a first channel grouping operation on the e-4th intermediate feature map to obtain multiple e-5th intermediate feature maps, the features can be refined and richer and more diverse features can be extracted on different channels. By performing a second conversion operation on multiple e-5th intermediate feature maps to obtain multiple e-6th intermediate feature maps, features can be further extracted and processed, further improving the overall expressive power of the features. On this basis, by performing a second fusion operation on multiple e-6th intermediate feature maps to obtain the e-3rd intermediate feature map, a more comprehensive and comprehensive e-3rd intermediate feature map can be obtained.
[0138] According to an embodiment of the present disclosure, obtaining the e-th third intermediate feature map according to the e-1-th first intermediate feature map may include the following operations.
[0139] Perform a second channel grouping operation on the e-1th first intermediate feature map to obtain a plurality of e-th seventh intermediate feature maps. Perform a third conversion operation on the plurality of e-th seventh intermediate feature maps to obtain a plurality of e-th eighth intermediate feature maps. Perform a third fusion operation on the plurality of e-th eighth intermediate feature maps to obtain a third intermediate feature map.
[0140] According to an embodiment of the present disclosure, the second channel grouping method, the third conversion method and the third fusion method can refer to the above-mentioned relevant contents about the first channel grouping method, the first conversion method and the second fusion method, which will not be repeated here.
[0141] According to an embodiment of the present disclosure, after obtaining the e-1th first intermediate feature map, the second channel grouping method described above can be used to perform a second channel grouping operation on the e-1th first intermediate feature map to obtain multiple e-th seventh intermediate feature maps. The third conversion method described above can be used to perform a third conversion operation on the multiple e-th seventh intermediate feature maps to obtain multiple e-th eighth intermediate feature maps. Based on this, the third fusion method described above can be used to perform a third fusion operation on the multiple e-th eighth intermediate feature maps to obtain an e-th third intermediate feature map.
[0142] According to an embodiment of the present disclosure, by performing a second channel grouping operation on the e-1th first intermediate feature map, a plurality of e-7th intermediate feature maps are obtained, which can refine the features and extract richer and more diverse features across different channels. By performing a third conversion operation on the plurality of e-7th intermediate feature maps, a plurality of e-8th intermediate feature maps are obtained, which can change the representation of the feature maps and thus extract more discriminative features. On this basis, by performing a third fusion operation on the plurality of e-8th intermediate feature maps, a more comprehensive and comprehensive e-3rd intermediate feature map can be obtained, which is beneficial for improving the effect of subsequent image restoration.
[0143] According to an embodiment of the present disclosure, obtaining the e-th second intermediate feature map according to the e-1th first intermediate feature map and the e-th third intermediate feature map may include the following operations.
[0144] The e-th third intermediate feature map is processed based on a first attention strategy to obtain an e-th ninth intermediate feature map, where the first attention strategy includes one of the following: a first channel attention strategy, a first spatial attention strategy, a first hybrid attention strategy, and a first self-attention strategy. The e-th second intermediate feature map is obtained based on the e-1th first intermediate feature map and the e-th ninth intermediate feature map.
[0145] According to an embodiment of the present disclosure, the first attention strategy may include at least one of the following: a first channel attention (i.e., Channel Attention) strategy, a first spatial attention (i.e., Spatial Attention) strategy, a first mixed attention (i.e., Mixture Attention) strategy, a first self-attention (i.e., Self-Attention) strategy, a first temporal attention (i.e., Temporal Attention) strategy and a first branch attention (i.e., Branch Attention) strategy.
[0146] According to embodiments of the present disclosure, the first-channel attention strategy adaptively adjusts the importance of different channels by learning dynamic weights. This strategy can focus on the content of useful information in the image, i.e., the "what." For example, the first-channel attention strategy can compress the original feature map into a global vector through global average pooling or global maximum pooling operations, then use a fully connected layer or convolutional layer to generate channel attention weights, and finally apply the channel attention weights to the original feature map.
[0147] According to an embodiment of the present disclosure, the first spatial attention strategy focuses on the correlation between different locations and strengthens or suppresses features at different locations by learning location-aware weights, so as to better utilize the local information of the input data in the spatial dimension. The first spatial attention strategy can focus on the locations in the image that have application information, that is, "where". For example, the first spatial attention strategy generates a position weight map through a convolution operation, and then multiplies or adds the weight map with the original feature map.
[0148] According to an embodiment of the present disclosure, a first hybrid attention strategy simultaneously focuses on the content and location of useful information in an image. For example, the first hybrid attention strategy can separately determine the channel attention features and spatial attention features of the input content, and then fuse the channel attention features and spatial attention features to obtain a hybrid feature that simultaneously focuses on the content and location of useful information in the image.
[0149] According to an embodiment of the present disclosure, the first self-attention strategy can capture the long-distance dependencies between different elements and is not limited by the length of the sequence, and is used to model the relationship between sequence or set data. For example, the first self-attention strategy can calculate the query matrix (Query Matrix), key matrix (Key Matrix) and value matrix (Value Matrix) based on the input sequence, and obtain the correlation score matrix of each query and key by performing a dot product operation on the query matrix and the key matrix, and perform a Softmax function normalization on the correlation score matrix to obtain an attention weight matrix. On this basis, the attention weight matrix can be multiplied by the value matrix to obtain a weighted value matrix, thereby obtaining a contextual representation of each element. Thus, each element can take into account the information of other elements at the same time, thereby better modeling the relationship between elements in the sequence.
[0150] According to embodiments of the present disclosure, a first-time attention strategy is used to process time series data, which can enhance the ability to model relationships between different time steps. For example, the first-time attention strategy can use a recurrent neural network or an attention mechanism to perform a weighted sum or concatenation of inputs at different time steps in the sequence.
[0151] According to an embodiment of the present disclosure, the first branch attention strategy can model features of different types or scales by introducing additional branches. Each branch focuses on different feature extraction and combines the outputs of these branches through an adaptive weight mechanism.
[0152] According to an embodiment of the present disclosure, after obtaining the e-th third intermediate feature map, the first attention strategy described above can be used to process the e-th third intermediate feature map to obtain the e-th ninth intermediate feature map. On this basis, the e-th second intermediate feature map can be obtained based on the e-1th first intermediate feature map and the e-th ninth intermediate feature map.
[0153] According to the embodiments of the present disclosure, since the e-th ninth intermediate feature map is obtained by processing the e-th third intermediate feature map based on the first attention strategy, by introducing the attention mechanism when processing the feature map, it is possible to better focus on important features. On this basis, by obtaining the e-th second intermediate feature map based on the e-1th first intermediate feature map and the e-th ninth intermediate feature map, it is possible to fully utilize the features of each pass.
[0154] According to an embodiment of the present disclosure, the first attention strategy includes a first channel attention strategy.
[0155] According to an embodiment of the present disclosure, processing the e-th third intermediate feature map based on the first attention strategy to obtain the e-th ninth intermediate feature map may include the following operations.
[0156] Perform a first compression operation on the e-th third intermediate feature map to obtain the e-th tenth intermediate feature map. Perform a first activation operation on the e-th tenth intermediate feature map to obtain the e-th first channel attention feature map. Perform a fourth fusion operation on the e-th third intermediate feature map and the e-th first channel attention feature map to obtain the e-th ninth intermediate feature map.
[0157] According to an embodiment of the present disclosure, the first compression operation can be used to reduce the dimension or size of the third intermediate feature map, thereby reducing the model parameters and the amount of calculation. The specific first compression method can be configured according to actual business needs and is not limited here. For example, the first compression method may include at least one of the following: maximum pooling, average pooling, convolutional kernel dimensionality reduction (i.e., Convolutional Kernel Dimension Reduction), compression algorithm (i.e., Compression Algorithms) and network pruning (i.e., Network Pruning).
[0158] According to an embodiment of the present disclosure, the first activation operation can be used to apply a nonlinear function to the e-th third intermediate feature map to enhance the expressive power and nonlinear modeling capability of the model. The specific first activation method can be configured according to actual business needs and is not limited here. For example, the first activation method may include at least one of the following: Rectified Linear Unit (ReLU), Leaky Rectified Linear Unit (Leaky ReLU), Parametric Rectified Linear Unit (PReLU), Sigmoid function, Hyperbolic Tangent (Tanh) and Softmax function.
[0159] According to an embodiment of the present disclosure, the fourth fusion operation can refer to the above-mentioned relevant content about the second fusion operation, which will not be repeated here.
[0160] According to an embodiment of the present disclosure, after obtaining the e-th third intermediate feature map, the above-mentioned first compression method can be used to perform a first compression operation on the e-th third intermediate feature map to obtain the e-th tenth intermediate feature map, so as to reduce the dimension of the e-th third intermediate feature map to a global vector. Using the above-mentioned first activation method, a first activation operation is performed on the e-th tenth intermediate feature map to obtain the e-th first channel attention feature map, so as to convert the global vector into a channel weight vector. On this basis, the above-mentioned second fusion method can be used to perform a fourth fusion operation on the e-th third intermediate feature map and the e-th first channel attention feature map to obtain the final attention-enhanced e-th ninth intermediate feature map.
[0161] According to an embodiment of the present disclosure, by performing a first compression operation on the e-th third intermediate feature map to obtain the e-th tenth intermediate feature map, the size and dimension of the e-th third intermediate feature map can be reduced for subsequent processing. By performing a first activation operation on the e-th tenth intermediate feature map to obtain the e-th first channel attention feature map, a nonlinear transformation can be introduced, so that the effective feature map channel weights are large, and the invalid or ineffective feature map channel weights are small, thereby enhancing the expressive power of the feature map. On this basis, by performing a fourth fusion operation on the e-th third intermediate feature map and the e-th first channel attention feature map to obtain the e-th ninth intermediate feature map, the e-th third intermediate feature map and the e-th first channel attention feature map can be combined, thereby reducing the dimension of the feature map while retaining key features, which is beneficial to improving processing efficiency and the efficiency of subsequent image restoration.
[0162] According to an embodiment of the present disclosure, the first attention strategy includes a first channel attention strategy.
[0163] According to an embodiment of the present disclosure, processing the e-th third intermediate feature map based on the first attention strategy to obtain the e-th ninth intermediate feature map may include the following operations.
[0164] A first convolution operation of multiple first parallel levels is performed on the e-th third intermediate feature map to obtain an e-th eleventh intermediate feature map corresponding to each of the multiple first parallel levels. A fifth fusion operation is performed on the multiple e-th eleventh intermediate feature maps to obtain an e-th twelfth intermediate feature map. A second compression operation is performed on the e-th twelfth intermediate feature map to obtain an e-th thirteenth intermediate feature map. A second activation operation is performed on the e-th thirteenth intermediate feature map to obtain an e-th second channel attention feature map corresponding to each of the multiple first parallel levels. A sixth fusion operation is performed on the e-th eleventh intermediate feature map and the e-th second channel attention feature map corresponding to each of the multiple first parallel levels to obtain an e-th fourteenth intermediate feature map corresponding to each of the multiple first parallel levels. A seventh fusion operation is performed on the multiple e-th fourteenth intermediate feature maps to obtain an e-th ninth intermediate feature map.
[0165] According to an embodiment of the present disclosure, the first convolution operation can be used to extract a multi-level feature representation of the e-th third intermediate feature map. The specific first convolution method can be configured according to actual business needs and is not limited here. For example, the first convolution method may include at least one of the following: multi-scale (Multi-Scale) convolution and residual (Residual) connection. Multi-scale convolution can capture feature information of different scales by performing convolution operations at different scales at the same time, for example, it can be achieved by using convolution kernels of different sizes or by sampling the input multiple times. The residual connection can directly pass the original feature information to the next layer by adding the input feature map to the output feature map element by element. For example, a residual block (Residual Block) can be used. The residual block includes two convolution layers and a jump connection.
[0166] According to an embodiment of the present disclosure, after obtaining the e-th third intermediate feature map, the above-mentioned first convolution method can be used to input the e-th third intermediate feature map into multiple convolution branches, so as to perform multiple first parallel level first convolution operations on the e-th third intermediate feature map, and obtain the e-th eleventh intermediate feature map corresponding to each of the multiple first parallel levels.
[0167] According to an embodiment of the present disclosure, the description of the second compression operation, the second activation operation, the fifth fusion operation, the sixth fusion operation and the seventh fusion operation can refer to the above-mentioned relevant content about the first compression operation, the first activation operation and the second fusion operation, which will not be repeated here.
[0168] According to an embodiment of the present disclosure, after obtaining multiple e-th eleventh intermediate feature maps, the above-mentioned second fusion method can be used to perform a fifth fusion operation on the multiple e-th eleventh intermediate feature maps output by multiple convolution branches to obtain an e-th twelfth intermediate feature map. Using the above-mentioned first compression method, a second compression operation is performed on the e-th twelfth intermediate feature map to obtain an e-th thirteenth intermediate feature map. The e-th thirteenth intermediate feature map can be used to characterize the global information of the multiple e-th eleventh intermediate feature maps. Using the above-mentioned first activation method, a second activation operation is performed on the e-th thirteenth intermediate feature map to obtain an e-th second channel attention feature map corresponding to each of the multiple first parallel levels. The e-th second channel attention feature map can be used to characterize the weight vectors corresponding to each of the different channel dimensions.
[0169] According to an embodiment of the present disclosure, after obtaining the e-th second channel attention feature map corresponding to each of the multiple first parallel levels, the above-mentioned second fusion method can be used to perform a sixth fusion operation on the e-th eleventh intermediate feature map corresponding to each of the multiple first parallel levels obtained by the first convolution operation and the e-th second channel attention feature map corresponding to each of the multiple first parallel levels obtained by the second activation operation, thereby obtaining the e-th fourteenth intermediate feature map corresponding to each of the multiple first parallel levels. On this basis, the above-mentioned second fusion method can be used to perform a seventh fusion operation on the multiple e-th fourteenth intermediate feature maps to obtain the e-th ninth intermediate feature map.
[0170] According to an embodiment of the present disclosure, by performing a first convolution operation of multiple first parallel levels on the e-th third intermediate feature map, an e-th eleventh intermediate feature map corresponding to each of the multiple first parallel levels is obtained, which can extract richer feature expressions, thereby enhancing the feature representation capability. By performing a fifth fusion operation on multiple e-th eleventh intermediate feature maps, an e-th twelfth intermediate feature map is obtained, which can integrate the information of multiple feature maps, which is conducive to utilizing the feature representation capabilities of different levels to extract more global and detailed feature information. By performing a second compression operation on the e-th twelfth intermediate feature map, an e-th thirteenth intermediate feature map is obtained, which can reduce the resolution or number of channels of the feature map, thereby reducing the amount of calculation and improving processing efficiency. By performing a second activation operation on the e-th thirteenth intermediate feature map, an e-th second channel attention feature map corresponding to each of the multiple first parallel levels is obtained, introducing a nonlinear mapping, which can strengthen the feature information of interest and suppress irrelevant features, thereby enhancing the discriminative ability of the feature. By performing a sixth fusion operation on the e-th eleventh intermediate feature map and the e-th second channel attention feature map corresponding to each of the multiple first parallel levels, a fourteenth intermediate feature map corresponding to each of the multiple first parallel levels is obtained. This allows the fusion of feature maps from multiple different sources, facilitating the extraction of a more global and richer feature representation. Furthermore, by performing a seventh fusion operation on the multiple fourteenth intermediate feature maps, a ninth intermediate feature map is obtained. This allows the fusion of feature maps from multiple different levels, combining feature information at different scales, and thus establishing a multi-scale feature representation.
[0171] According to an embodiment of the present disclosure, the first attention strategy includes a first spatial attention strategy.
[0172] According to an embodiment of the present disclosure, processing the e-th third intermediate feature map based on the first attention strategy to obtain the e-th ninth intermediate feature map may include the following operations.
[0173] Perform a maximum pooling operation on the e-th third intermediate feature map to obtain the e-th fifteenth intermediate feature map. Perform an average pooling operation on the e-th third intermediate feature map to obtain the e-th sixteenth intermediate feature map. Perform a second channel fusion operation on the e-th fifteenth intermediate feature map and the e-th sixteenth intermediate feature map to obtain the e-th seventeenth intermediate feature map. Based on the e-th seventeenth intermediate feature map, obtain the e-th ninth intermediate feature map.
[0174] According to an embodiment of the present disclosure, a spatial attention module can be used to process the e-th third intermediate feature map to obtain the e-th ninth intermediate feature map. The convolution block attention module may further include a channel attention module. The channel attention module may be used to learn weight information for each channel to improve the response of important channels. The spatial attention module may be used to learn weight information for different spatial positions in the feature map to improve the response of important spatial regions.
[0175] According to an embodiment of the present disclosure, a maximum pooling operation can be used to compress a feature map, thereby reducing the size of the feature map while retaining the main features. For example, the maximum pooling operation can divide the input feature map into non-overlapping regions and select the maximum value in each region as the output.
[0176] According to an embodiment of the present disclosure, an average pooling operation can be used to more smoothly reduce the dimension of a feature map. For example, the average pooling operation can divide the input feature map into non-overlapping regions and calculate the average value of each region as the output.
[0177] According to an embodiment of the present disclosure, after obtaining the e-th third intermediate feature map, the above-mentioned maximum pooling method can be used to perform a maximum pooling operation on the e-th third intermediate feature map to obtain an e-th fifteenth intermediate feature map. The above-mentioned average pooling method can be used to perform an average pooling operation on the e-th third intermediate feature map to obtain an e-th sixteenth intermediate feature map.
[0178] According to the embodiment of the present disclosure, the description of the second channel fusion operation can refer to the above-mentioned relevant content about the first channel fusion operation, which will not be repeated here. After obtaining the e-th sixteenth intermediate feature map, the above-mentioned first channel fusion method can be used to perform a second channel fusion operation on the e-th fifteenth intermediate feature map and the e-th sixteenth intermediate feature map to obtain the e-th seventeenth intermediate feature map that can be used to characterize the spatial weight. On this basis, the e-th seventeenth intermediate feature map can be element-wise multiplied with the original e-th third intermediate feature map to achieve adaptive weighting of different spatial positions, thereby obtaining the e-th ninth intermediate feature map that highlights the important spatial areas.
[0179] According to an embodiment of the present disclosure, by performing a maximum pooling operation on the e-th third intermediate feature map, the role of the most significant features in the feature map can be emphasized, which helps to extract the main feature information and reduce the influence of noise, thereby obtaining the e-th fifteenth intermediate feature map. By performing an average pooling operation on the e-th third intermediate feature map, the effect of a smooth feature map can be achieved, and the detail information in the feature map can be blurred to obtain the e-th sixteenth intermediate feature map. By performing a second channel fusion operation on the e-th fifteenth intermediate feature map and the e-th sixteenth intermediate feature map, the advantages of the two can be effectively combined to obtain the e-th seventeenth intermediate feature map, which is beneficial to improving the expressive power of the feature map. On this basis, by further processing the e-th seventeenth intermediate feature map, the e-th ninth intermediate feature map can be obtained, which is beneficial to improving the accuracy of subsequent image restoration.
[0180] According to an embodiment of the present disclosure, the first attention strategy includes a first spatial attention strategy.
[0181] According to an embodiment of the present disclosure, processing the e-th third intermediate feature map based on the first attention strategy to obtain the e-th ninth intermediate feature map may include the following operations.
[0182] Perform a second convolution operation on the e-th third intermediate feature map to obtain the e-th eighteenth intermediate feature map, the e-th nineteenth intermediate feature map, and the e-th twentieth intermediate feature map. Perform a deformation operation on the e-th eighteenth intermediate feature map, the e-th nineteenth intermediate feature map, and the e-th twentieth intermediate feature map to obtain the e-th twenty-first intermediate feature map corresponding to the e-th eighteenth intermediate feature map, the e-th twenty-second intermediate feature map corresponding to the e-th nineteenth intermediate feature map, and the e-th twenty-third intermediate feature map corresponding to the e-th twentieth intermediate feature map. Perform an eighth fusion operation on the e-th twenty-first intermediate feature map and the e-th twenty-second intermediate feature map to obtain the e-th twenty-fourth intermediate feature map. Based on the e-th twenty-fourth intermediate feature map, obtain the e-th twenty-fifth intermediate feature map. Perform an eighth fusion operation on the e-th twenty-third intermediate feature map and the e-th twenty-fifth intermediate feature map to obtain the e-th twenty-sixth intermediate feature map. Perform a ninth fusion operation on the e-th third intermediate feature map and the e-th twenty-sixth intermediate feature map to obtain the e-th ninth intermediate feature map.
[0183] According to embodiments of the present disclosure, the e-th third intermediate feature map can be processed to obtain an e-th ninth intermediate feature map. For example, a dual attention network can be used to process the e-th third intermediate feature map to obtain a series of feature maps. On this basis, after obtaining spatial position and channel relationship information of the series of feature maps, this relationship information is fused onto the e-th third intermediate feature map in an element-wise addition manner to obtain the e-th ninth intermediate feature map.
[0184] According to an embodiment of the present disclosure, the description of the second convolution operation, the eighth fusion operation and the ninth fusion operation can refer to the above-mentioned relevant content about the first convolution operation and the second fusion operation, which will not be repeated here. The deformation operation can be used to perform an adaptive geometric transformation on the feature map. The specific deformation method can be configured according to actual business needs and is not limited here. For example, the deformation method may include at least one of the following: bilinear interpolation, nearest neighbor interpolation, cubic spline interpolation, spatial transformation network (STN), convolutional neural network or deconvolutional neural network (DNN).
[0185] According to an embodiment of the present disclosure, after obtaining the e-th third intermediate feature map, the e-th third intermediate feature map can be subjected to a second convolution operation using the above-mentioned first convolution method to obtain an e-th eighteenth intermediate feature map, an e-th nineteenth intermediate feature map, and an e-th twentieth intermediate feature map. The e-th eighteenth intermediate feature map, the e-th nineteenth intermediate feature map, and the e-th twentieth intermediate feature map can be subjected to a deformation operation using the above-mentioned deformation method to obtain an e-th twenty-first intermediate feature map corresponding to the e-th eighteenth intermediate feature map, an e-th twenty-second intermediate feature map corresponding to the e-th nineteenth intermediate feature map, and an e-th twenty-third intermediate feature map corresponding to the e-th twentieth intermediate feature map.
[0186] According to an embodiment of the present disclosure, after obtaining the e-th twenty-first intermediate feature map, the e-th twenty-second intermediate feature map, and the e-th twenty-third intermediate feature map, the second fusion method described above can be used to perform an eighth fusion operation on the e-th twenty-first intermediate feature map and the e-th twenty-second intermediate feature map to obtain the e-th twenty-fourth intermediate feature map. Based on the e-th twenty-fourth intermediate feature map, the e-th twenty-fifth intermediate feature map is obtained. Using the second fusion method described above, the e-th twenty-third intermediate feature map and the e-th twenty-fifth intermediate feature map are performed an eighth fusion operation to obtain the e-th twenty-sixth intermediate feature map. Based on this, the second fusion method described above can be used to perform a ninth fusion operation on the e-th third intermediate feature map and the e-th twenty-sixth intermediate feature map to obtain the e-th ninth intermediate feature map.
[0187] According to an embodiment of the present disclosure, the first attention strategy includes a first hybrid attention strategy.
[0188] According to an embodiment of the present disclosure, processing the e-th third intermediate feature map based on the first attention strategy to obtain the e-th ninth intermediate feature map may include the following operations.
[0189] Perform the first channel attention operation on the e-th third intermediate feature map to obtain the e-th third channel attention feature map. Perform the first spatial attention operation on the e-th third intermediate feature map to obtain the e-th first spatial attention feature map. Perform the tenth fusion operation on the e-th third channel attention feature map and the e-th first spatial attention feature map to obtain the e-th fused attention feature map. Perform the eleventh fusion operation on the e-th third intermediate feature map and the e-th fused attention feature map to obtain the e-th ninth intermediate feature map.
[0190] According to an embodiment of the present disclosure, the e-th third intermediate feature map can be processed to obtain the e-th ninth intermediate feature map.
[0191] According to an embodiment of the present disclosure, the first channel attention operation can adjust the importance of each channel through the learned weights. The specific first channel attention method can be configured according to actual business needs and is not limited here. For example, the first channel attention method may include at least one of the following: global average pooling processing, fully connected layer processing, activation function processing, and channel weighted processing. The above-mentioned first channel attention method can be used to perform a first channel attention operation on the e-th third intermediate feature map to obtain the e-th third channel attention feature map.
[0192] According to an embodiment of the present disclosure, the first spatial attention operation can be used to adjust the importance of feature maps at different spatial positions. The specific first spatial attention method can be configured according to actual business needs and is not limited here. For example, the first spatial attention method may include at least one of the following: convolution processing, normalization processing, attention weight processing, and feature weighted processing. The above-mentioned first spatial attention method can be used to perform a first spatial attention operation on the e-th third intermediate feature map to obtain the e-th first spatial attention feature map.
[0193] According to an embodiment of the present disclosure, the description of the tenth fusion operation and the eleventh fusion operation can be found in the above-mentioned relevant content about the second fusion operation, and will not be repeated here. The above-mentioned second fusion method can be used to perform a fusion operation on the e-th third intermediate feature map, the e-th third channel attention feature map, and the e-th first spatial attention feature map to obtain the e-th ninth intermediate feature map.
[0194] According to an embodiment of the present disclosure, by performing a channel attention operation on the third intermediate feature map, the degree of attention to different channel features can be strengthened, and important channel information can be highlighted, thereby obtaining the eth third channel attention feature map. By performing a first spatial attention operation on the eth third intermediate feature map, the degree of attention to different spatial positions can be enhanced, so that the attention is more focused on the target area, thereby obtaining the eth first spatial attention feature map. By performing a tenth fusion operation on the eth third channel attention feature map and the eth first spatial attention feature map, the information of the two can be comprehensively utilized to obtain a richer and more useful feature representation, which helps to enhance the expressive power of the feature map, thereby obtaining the eth fused attention feature map that can better capture the key features of the target and related contextual information. On this basis, by performing an eleventh fusion operation on the eth third intermediate feature map and the eth fused attention feature map, feature information at different levels can be integrated to obtain the eth ninth intermediate feature map that more comprehensively captures feature expressions at multiple levels, thereby improving the representation power and semantic richness of the feature map.
[0195] According to an embodiment of the present disclosure, the first attention strategy includes a first hybrid attention strategy.
[0196] According to an embodiment of the present disclosure, processing the e-th third intermediate feature map based on the first attention strategy to obtain the e-th ninth intermediate feature map may include the following operations.
[0197] Perform the second channel attention operation on the e-th third intermediate feature map to obtain the e-th fourth channel attention feature map. Perform the twelfth fusion operation on the e-th third intermediate feature map and the e-th fourth channel attention feature map to obtain the e-th twenty-seventh intermediate feature map. Perform the second spatial attention operation on the e-th twenty-seventh intermediate feature map to obtain the e-th second spatial attention feature map. Perform the thirteenth fusion operation on the e-th twenty-seventh intermediate feature map and the e-th second spatial attention feature map to obtain the e-th ninth intermediate feature map.
[0198] According to an embodiment of the present disclosure, the convolutional block attention module can be used to process the e-th third intermediate feature map to obtain the e-th ninth intermediate feature map. The description of the second channel attention operation, the second spatial attention operation, the twelfth fusion operation, and the thirteenth fusion operation can be found in the above-mentioned related content about the first channel attention operation, the first spatial attention operation, and the second fusion operation, which will not be repeated here.
[0199] According to an embodiment of the present disclosure, after obtaining the e-th third intermediate feature map, the first channel attention method described above can be used to perform a second channel attention operation on the e-th third intermediate feature map to obtain an e-th fourth channel attention feature map that can be used to characterize the useful information in the image. On this basis, the second fusion method described above can be used to perform a twelfth fusion operation on the e-th third intermediate feature map and the e-th fourth channel attention feature map to obtain an e-th twenty-seventh intermediate feature map.
[0200] According to an embodiment of the present disclosure, after obtaining the e-th twenty-seventh intermediate feature map, the second spatial attention method described above can be used to perform a second spatial attention operation on the e-th twenty-seventh intermediate feature map to obtain an e-th second spatial attention feature map that can be used to characterize the location of useful information in the image. Based on this, the second fusion method described above can be used to perform a thirteenth fusion operation on the e-th twenty-seventh intermediate feature map and the e-th second spatial attention feature map to obtain an e-th ninth intermediate feature map that fuses channel features and spatial features.
[0201] According to the embodiments of the present disclosure, by performing a second channel attention operation on the e-th third intermediate feature map, the focus on different channel features can be strengthened and important channel information can be highlighted, thereby obtaining the e-th fourth channel attention feature map. By performing a twelfth fusion operation on the e-th third intermediate feature map and the e-th fourth channel attention feature map, the information of both can be comprehensively utilized to obtain a richer and more useful feature representation, which helps to enhance the expressive power of the feature map, thereby obtaining the e-th twenty-seventh intermediate feature map that can better capture the key features of the target and relevant contextual information. By performing a second spatial attention operation on the e-th twenty-seventh intermediate feature map, the focus on different spatial locations can be enhanced, so that attention is more focused on the target area, thereby obtaining the e-th second spatial attention feature map. On this basis, by performing a thirteenth fusion operation on the e-th twenty-seventh intermediate feature map and the e-th second spatial attention feature map, feature information at different levels can be integrated to obtain the e-th ninth intermediate feature map that more comprehensively captures feature expressions at multiple levels, thereby improving the expressive power and semantic richness of the feature map.
[0202] According to an embodiment of the present disclosure, obtaining a first repaired image according to the Eth first intermediate feature map may include the following operations.
[0203] The Eth first intermediate feature map is processed based on a second attention strategy to obtain an eth twenty-eighth intermediate feature map (i.e., the eth tenth intermediate feature map), where the second attention strategy includes one of the following: a second channel attention strategy, a second spatial attention strategy, a second hybrid attention strategy, and a second self-attention strategy. A first inpainted image is obtained based on the Eth first intermediate feature map and the eth twenty-eighth intermediate feature map.
[0204] According to an embodiment of the present disclosure, the second attention strategy may include at least one of the following: a second channel attention strategy, a second spatial attention strategy, a second hybrid attention strategy, a second self-attention strategy, a second temporal attention strategy, and a second branch attention strategy. The description of the second channel attention strategy, the second spatial attention strategy, the second hybrid attention strategy, the second self-attention strategy, the second temporal attention strategy, and the second branch attention strategy can be found in the above-mentioned relevant content about the first channel attention strategy, the first spatial attention strategy, the first hybrid attention strategy, the first self-attention strategy, the first temporal attention strategy, and the first branch attention strategy, which will not be repeated here.
[0205] According to an embodiment of the present disclosure, after obtaining the Eth first intermediate feature map, the second attention strategy described above can be used to process the Eth first intermediate feature map to obtain the Eth twenty-eighth intermediate feature map. On this basis, a first inpainted image can be obtained based on the Eth first intermediate feature map and the Eth twenty-eighth intermediate feature map.
[0206] According to the embodiments of the present disclosure, since the e-th twenty-eighth intermediate feature map is obtained by processing the e-th first intermediate feature map based on the second attention strategy, by introducing the attention mechanism when processing the feature map, it is possible to better focus on important features. On this basis, by obtaining the first restored image based on the e-th first intermediate feature map and the e-th twenty-eighth intermediate feature map, it is possible to fully utilize the features of each time, thereby improving the quality of the first restored image and the accuracy of image restoration.
[0207] According to an embodiment of the present disclosure, obtaining the e-th second intermediate feature map according to the e-1-th first intermediate feature map may include the following operations.
[0208] In the case of 2<g≤G, according to the twenty-ninth intermediate feature map of the e1th to the twenty-ninth intermediate feature map of the e1th g-1 The twenty-ninth intermediate feature map (i.e., the e g-1 eleventh intermediate feature map), get the e g The 29th intermediate feature map, wherein the 29th intermediate feature map of e2 (i.e. the 11th intermediate feature map of e2) is obtained based on the 29th intermediate feature map of e1 (i.e. the 11th intermediate feature map of e1), and the 29th intermediate feature map of e1 is the first intermediate feature map of e-1. G The twenty-ninth intermediate feature map (i.e., the e G eleventh intermediate feature map), and obtain the eth second intermediate feature map. G is an integer greater than 2, and g is an integer greater than or equal to 1 and less than or equal to G.
[0209] According to an embodiment of the present disclosure, the eighth deep learning model can be used to analyze the e1th to the twenty-ninth intermediate feature maps. g-1 The twenty-ninth intermediate feature map is processed to obtain the e-th second intermediate feature map. For example, the eighth deep learning model may include a densely connected convolutional network (DenseNet). For example, the densely connected convolutional network may include an input layer, a dense block (Dense Block), a transition layer (Transition Layer) and a global average pooling layer (Global Average Pooling).
[0210] According to an embodiment of the present disclosure, a dense block may include multiple dense layers, each dense layer includes multiple convolutional layers, and the input between different dense layers can be the concatenation of all previous layers, that is, each dense layer in the dense block receives the feature maps of all previous layers as input.
[0211] For example, the first intermediate feature map of the e-1th layer can be used as the 29th intermediate feature map of the e1th layer, and the 29th intermediate feature map of the e1th layer can be input into the first dense layer to obtain the 29th intermediate feature map of the e2th layer. The 29th intermediate feature map of the e1th layer and the 29th intermediate feature map of the e2th layer can be input into the second dense layer to obtain the 29th intermediate feature map of the e3th layer. Similarly, the 29th intermediate feature map of the e1th layer can be used as the 29th intermediate feature map of the e1th layer. g-1 The twenty-ninth intermediate feature map is input to the g-1th dense layer to obtain the e-th g Twenty-ninth intermediate feature map.
[0212] According to an embodiment of the present disclosure, a transition layer can be used to reduce the dimension and size of a feature map. For example, the transition layer can include a convolutional layer and a pooling layer. A global average pooling layer can be used to convert the feature map output by the transition layer into a fixed-length vector representation.
[0213] According to the embodiment of the present disclosure, due to the e g The twenty-ninth intermediate feature map is a graph based on the twenty-ninth intermediate feature map from the e1th to the e2nd. g-1 The 29th intermediate feature map is obtained by introducing dense connections, which can enhance feature reuse and gradient flow. G The twenty-ninth intermediate feature map can obtain the second intermediate feature map containing richer and more useful features.
[0214] According to an embodiment of the present disclosure, according to the e G The twenty-ninth intermediate feature map, obtaining the e-th second intermediate feature map, may include the following operations.
[0215] Processing the first G The 29th intermediate feature map is obtained, wherein the third attention strategy includes one of the following: a third channel attention strategy, a third spatial attention strategy, a third mixed attention strategy and a third self-attention strategy. G The twenty-ninth intermediate feature map and the e-th thirtieth intermediate feature map obtain the e-th second intermediate feature map.
[0216] According to an embodiment of the present disclosure, the third attention strategy may include at least one of the following: a third channel attention strategy, a third spatial attention strategy, a third mixed attention strategy, a third self-attention strategy, a third temporal attention strategy, and a third branch attention strategy. For the description of the third channel attention strategy, the third spatial attention strategy, the third mixed attention strategy, the third self-attention strategy, the third temporal attention strategy, and the third branch attention strategy, please refer to the above-mentioned relevant content on the first channel attention strategy, the first spatial attention strategy, the first mixed attention strategy, the first self-attention strategy, the first temporal attention strategy, and the first branch attention strategy, which will not be repeated here.
[0217] According to an embodiment of the present disclosure, after obtaining the e G After the 29th intermediate feature map, the third attention strategy mentioned above can be used to process the e G The 29th intermediate feature map is obtained by the 30th intermediate feature map. G The twenty-ninth intermediate feature map and the e-th thirtieth intermediate feature map obtain the e-th second intermediate feature map.
[0218] According to an embodiment of the present disclosure, obtaining the e-th second intermediate feature map according to the e-1-th first intermediate feature map may include the following operations.
[0219] Perform the third convolution operation of multiple second parallel levels (i.e., the first parallel level) on the e-1th first intermediate feature map to obtain an e-th thirty-first intermediate feature map (i.e., the e-th twelfth intermediate feature map) corresponding to each of the multiple second parallel levels. Obtain an e-th second intermediate feature map based on the multiple e-th thirty-first intermediate feature maps.
[0220] According to an embodiment of the present disclosure, for the description of the third convolution operation, please refer to the relevant content about the first convolution operation mentioned above, and no further details will be given here. After obtaining the e-1th first intermediate feature map, the above-mentioned third convolution method can be used to input the e-1th first intermediate feature map into multiple convolution branches, so as to perform the third convolution operation of multiple second parallel levels on the e-1th first intermediate feature map, and obtain the e-th thirty-first intermediate feature map corresponding to each of the multiple second parallel levels.
[0221] According to an embodiment of the present disclosure, by performing multiple second-parallel-level third convolution operations on the e-1th first intermediate feature map, multiple e-31st intermediate feature maps with different dimensions and representation capabilities can be obtained at different second-parallel levels. Through multi-level parallel operations, richer and more diverse feature representations can be extracted, which can better capture the details and semantic information of the target. On this basis, by integrating the features of multiple second-parallel levels based on multiple e-31st intermediate feature maps, the e-th second intermediate feature map with more comprehensive and rich representation capabilities can be obtained.
[0222] According to an embodiment of the present disclosure, obtaining an e-th second intermediate feature map based on multiple e-th thirty-first intermediate feature maps may include the following operations.
[0223] A fourteenth fusion operation is performed on the plurality of the e-th thirty-first intermediate feature maps (i.e., the e-th twelfth intermediate feature maps) to obtain an e-th thirty-second intermediate feature map (i.e., the e-th thirteenth intermediate feature map). The e-th thirty-second intermediate feature map is processed based on a fourth attention strategy to obtain an e-th thirty-third intermediate feature map (i.e., the e-th fourteenth intermediate feature map), wherein the fourth attention strategy includes one of the following: a fourth channel attention strategy, a fourth spatial attention strategy, a fourth hybrid attention strategy, and a fourth self-attention strategy. Based on the e-th thirty-second intermediate feature map and the e-th thirty-third intermediate feature map, an e-th second intermediate feature map is obtained.
[0224] According to an embodiment of the present disclosure, the fourth attention strategy may include at least one of the following: a fourth channel attention strategy, a fourth spatial attention strategy, a fourth hybrid attention strategy, a fourth self-attention strategy, a fourth temporal attention strategy, and a fourth branch attention strategy. For the description of the fourteenth fusion operation, the fourth channel attention strategy, the fourth spatial attention strategy, the fourth hybrid attention strategy, the fourth self-attention strategy, the fourth temporal attention strategy, and the fourth branch attention strategy, please refer to the above-mentioned relevant content about the second fusion operation, the first channel attention strategy, the first spatial attention strategy, the first hybrid attention strategy, the first self-attention strategy, the first temporal attention strategy, and the first branch attention strategy, which will not be repeated here.
[0225] According to an embodiment of the present disclosure, after obtaining multiple e-th 31st intermediate feature maps, the aforementioned second fusion method can be used to perform a fourteenth fusion operation on the multiple e-th 31st intermediate feature maps to obtain an e-th 32nd intermediate feature map. The aforementioned fourth attention strategy is used to process the e-th 32nd intermediate feature map to obtain an e-th 33rd intermediate feature map. Based on this, an e-th 2nd intermediate feature map can be obtained based on the e-th 32nd intermediate feature map and the e-th 33rd intermediate feature map.
[0226] According to embodiments of the present disclosure, by performing a fourteenth fusion operation on multiple e-th thirty-first intermediate feature maps, the feature information of the multiple e-th thirty-first intermediate feature maps can be comprehensively utilized to obtain an e-th thirty-second intermediate feature map. By processing the e-th thirty-second intermediate feature map based on the fourth attention strategy, specific features can be focused on in the channel or spatial dimension, resulting in an e-th thirty-third intermediate feature map with higher expressiveness. Furthermore, based on the e-th thirty-second intermediate feature map and the e-th thirty-third intermediate feature map, the expressiveness of the e-th second intermediate feature map can be further improved.
[0227] According to an embodiment of the present disclosure, performing a restoration operation on the first image to be restored according to the image quality evaluation vector to obtain the first restored image may include the following operations.
[0228] The target first image to be inpainted and the image quality assessment vector are processed at multiple first cascade levels to obtain thirty-fourth intermediate feature maps (i.e., fifteenth intermediate feature maps) corresponding to each of the multiple first cascade levels, wherein the target first image to be inpainted is obtained based on the first image to be inpainted. The target first image to be inpainted is processed at multiple second cascade levels to obtain a first inpainted image.
[0229] According to an embodiment of the present disclosure, the first cascade level processing can be used to extract features from the target first image to be repaired and the image quality assessment vector. For example, the multiple first cascade levels may include the 1st first cascade level, the 2nd first cascade level, ..., the x1th first cascade level. On this basis, performing multiple first cascade level processing on the target first image to be repaired and the image quality assessment vector to obtain the thirty-fourth intermediate feature map corresponding to each of the multiple first cascade levels may include: using the 1st first cascade level to extract features from the target first image to be repaired and the image quality assessment vector to obtain the 1st thirty-fourth intermediate feature map. Using the 2nd first cascade level to extract features from the 1st thirty-fourth intermediate feature map to obtain the 2nd thirty-fourth intermediate feature map. Similarly, using the x1th first cascade level to extract features from the x1-1th thirty-fourth intermediate feature map to obtain the x1th thirty-fourth intermediate feature map.
[0230] According to an embodiment of the present disclosure, the second cascade level processing can be used to perform feature fusion on the thirty-fourth intermediate feature maps corresponding to each of the multiple first cascade levels. For example, the multiple second cascade levels may include the 1st second cascade level, the 2nd second cascade level, ..., the x2th second cascade level. On this basis, performing multiple second cascade level processing on the thirty-fourth intermediate feature maps corresponding to each of the multiple first cascade levels to obtain the first repaired image may include: using the 1st second cascade level to perform feature fusion on the thirty-fourth intermediate feature maps corresponding to each of the multiple first cascade levels to obtain the 1st first processed feature map. Using the 2nd second cascade level to perform feature fusion on the 1st first processed feature map to obtain the 2nd first processed feature map. Similarly, using the x2th second cascade level to perform feature fusion on the x2-1th first processed feature map to obtain the first repaired image.
[0231] According to the embodiments of the present disclosure, by performing multiple first-cascade-level processing on the target first image to be inpainted and the image quality assessment vector, feature information at different levels and aspects can be extracted, thereby obtaining a thirty-fourth intermediate feature map corresponding to each of the multiple first-cascade-levels. Furthermore, by performing multiple second-cascade-level processing on the thirty-fourth intermediate feature map corresponding to each of the multiple first-cascade-levels, further features can be extracted and a more accurate inpainting result can be generated, thereby improving the quality of the first inpainted image and the accuracy of the image inpainting.
[0232] FIG4 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0233] As shown in Figure 4, in step 400, an image quality assessment operation is performed on a first image to be restored 403 to obtain image quality assessment information 401. Based on the first image to be restored 403, a target first image to be restored 404 is obtained.
[0234] Based on the image quality assessment information 401, an image quality assessment vector 402 is obtained. Multiple first cascade-level processing is performed on the target first image to be restored 404 and the image quality assessment vector 402, resulting in a thirty-fourth intermediate feature map 405 corresponding to each of the multiple first cascade-level processings. Multiple second cascade-level processing is performed on the thirty-fourth intermediate feature map 405 corresponding to each of the multiple first cascade-level processings, resulting in a first restored image 406.
[0235] According to an embodiment of the present disclosure, the target first image to be restored is obtained based on the first image to be restored, and may include one of the following.
[0236] The target first image to be repaired is the first image to be repaired. The target first image to be repaired is obtained based on the thirty-fifth intermediate feature map (i.e., the sixteenth intermediate feature map) corresponding to each of the multiple third parallel levels (i.e., the second parallel level), and the thirty-fifth intermediate feature map corresponding to each of the multiple third parallel levels is obtained by performing multiple third parallel level processing on the first image to be repaired. The target first image to be repaired is obtained based on the first intermediate repair image and the first structural image, the first intermediate repair image is obtained based on the first image to be repaired, and the first structural image is obtained based on the first intermediate repair image. The target first image to be repaired is obtained based on the first image to be repaired and the second intermediate repair image, and the second intermediate repair image is obtained by performing a structural repair operation on the first image to be repaired.
[0237] According to an embodiment of the present disclosure, after obtaining a first image to be inpainted, the first image to be inpainted can be input into multiple third-parallel hierarchical branches, so that the multiple third-parallel hierarchical branches perform multiple third-parallel hierarchical processing on the first image to be inpainted, thereby obtaining a thirty-fifth intermediate feature map corresponding to each of the multiple third-parallel hierarchical processes. Based on this, the target first image to be inpainted can be obtained based on the thirty-fifth intermediate feature map corresponding to each of the multiple third-parallel hierarchical processes.
[0238] According to an embodiment of the present disclosure, after obtaining the first image to be repaired, a first intermediate repaired image can be obtained based on the first image to be repaired. For example, a predetermined detection algorithm can be used to perform basic processing on the first image to be repaired to obtain a first intermediate repaired image for characterizing the image edge information of the non-missing area in the first image to be repaired. The predetermined detection algorithm can be configured according to actual business needs and is not limited here. For example, the predetermined detection algorithm can include at least one of the following: a Canny edge detection algorithm, an edge detection algorithm based on a Sobel operator, an edge detection algorithm based on a Roberts operator, an edge detection algorithm based on a Laplacian operator, and an edge detection algorithm based on a Prewitt operator.
[0239] According to an embodiment of the present disclosure, after obtaining the first intermediate repaired image, a first structural image can be obtained based on the first intermediate repaired image. For example, the first intermediate repaired image can be trained using a predetermined generative model to predict a first structural image for characterizing the image edge information of the missing area in the first image to be repaired. It can be configured according to actual business needs and is not limited here. For example, the predetermined generative model can be configured according to actual business needs and is not limited here. For example, the predetermined generative model can include at least one of the following: original generative adversarial network (i.e., GAN), conditional generative adversarial network (Conditional GAN, cGAN), deep convolutional generative adversarial network (Deep Convolutional GAN, DCGAN), Wasserstein generative adversarial network (Wasserstein GAN, WGAN), progressive generative adversarial network (Progressive GAN).
[0240] According to an embodiment of the present disclosure, after obtaining the first intermediate restoration image and the first structural image, the image edge information of the non-missing area in the first image to be restored and the image edge information of the missing area in the first image to be restored can be used as guidance to determine the target first image to be restored.
[0241] According to embodiments of the present disclosure, after obtaining a first image to be restored, a structural restoration operation can be performed on the first image to be restored to obtain a second intermediate restored image. For example, a coarse processing module of a predetermined restoration model can be used to perform a structural restoration operation on the first image to be restored to obtain a second intermediate restored image. The coarse processing module can fill in low-dimensional information or high-dimensional edge structures in missing image regions. The coarse processing module is used to map the spatial information of the first image to be restored from a low-resolution to a high-resolution space and extract high-level features for subsequent processing.
[0242] According to embodiments of the present disclosure, a target first image to be restored can be obtained based on a first image to be restored and a second intermediate restoration image. For example, the first image to be restored and the second intermediate restoration image can be processed using a fine processing module of a predetermined restoration model to obtain the target first image to be restored. The fine processing module can reconstruct the semantic information of the image based on the global structural information output by the coarse processing module. The fine processing module is configured to receive the feature map output by the coarse processing module and then utilize a deeper structure to enhance and reconstruct the feature map.
[0243] According to an embodiment of the present disclosure, by performing multiple third parallel level processing on the first image to be repaired, a variety of feature information can be extracted and integrated to obtain a thirty-fifth intermediate feature map corresponding to each of the multiple third parallel levels. By obtaining a first structural image based on the first intermediate repair image, and obtaining a target first image to be repaired based on the first intermediate repair image and the first structural image, the target first image to be repaired can be obtained by combining image structural information. By performing a structural repair operation on the first image to be repaired to obtain a second intermediate repair image, and obtaining a target first image to be repaired based on the first image to be repaired and the second intermediate repair image, it is possible to detect and fill in missing areas of the image, which is beneficial to improving the integrity of the target first image to be repaired, thereby improving the quality and accuracy of image repair.
[0244] According to an embodiment of the present disclosure, the second intermediate restoration image is obtained by performing a structural restoration operation on the first image to be restored, and may include the following operations.
[0245] The second intermediate restoration image is obtained by performing multiple fourth cascade level processing on the thirty-sixth intermediate feature maps corresponding to each of the multiple third cascade levels, and the thirty-sixth intermediate feature maps corresponding to each of the multiple third levels are obtained by performing multiple third cascade level processing on the first image to be restored.
[0246] According to an embodiment of the present disclosure, the third cascade level processing can be used to extract features from the first image to be repaired. For example, the multiple third cascade levels may include the 1st third cascade level, the 2nd third cascade level, ..., the y1th third cascade level. On this basis, performing multiple third cascade level processing on the first image to be repaired to obtain the thirty-sixth intermediate feature map corresponding to each of the multiple third levels may include: using the 1st third cascade level to extract features from the first image to be repaired to obtain the 1st thirty-sixth intermediate feature map. Using the 2nd third cascade level to extract features from the 1st thirty-sixth intermediate feature map to obtain the 2nd thirty-sixth intermediate feature map. Similarly, using the y1th third cascade level to extract features from the y1-1th thirty-sixth intermediate feature map to obtain the y1th thirty-sixth intermediate feature map.
[0247] According to an embodiment of the present disclosure, the fourth cascade level processing can be used to perform feature fusion on the thirty-sixth intermediate feature maps corresponding to each of the multiple third levels. For example, the multiple fourth cascade levels may include the first fourth cascade level, the second fourth cascade level, ..., the y2th fourth cascade level. On this basis, performing multiple fourth cascade level processing on the thirty-sixth intermediate feature maps corresponding to each of the multiple third levels to obtain the second intermediate repaired image may include: using the first fourth cascade level to perform feature fusion on the thirty-sixth intermediate feature maps corresponding to each of the multiple fourth cascade levels to obtain the first second processed feature map. Using the second fourth cascade level to perform feature fusion on the first second processed feature map to obtain the second second processed feature map. Similarly, using the y2th fourth cascade level to perform feature fusion on the y2-1th second processed feature map to obtain the second intermediate repaired image.
[0248] According to an embodiment of the present disclosure, the plurality of third parallel stages includes H stages, where H is an integer greater than 1.
[0249] According to an embodiment of the present disclosure, the target first image to be repaired is obtained based on the thirty-fifth intermediate feature maps corresponding to each of the multiple third parallel levels, which may include the following operations.
[0250] The target first image to be repaired is obtained based on the thirty-fifth intermediate feature map corresponding to each of the H third parallel levels. The thirty-fifth intermediate feature map is obtained by decoding the thirty-seventh intermediate feature map, and the thirty-seventh intermediate feature map is obtained by encoding the first image to be repaired.
[0251] According to an embodiment of the present disclosure, the encoding operation can convert the feature map into a more compact and representative representation. After obtaining the first image to be repaired, the first image to be repaired can be encoded using H first and third parallel layers to obtain a thirty-seventh intermediate feature map. The specific encoding method can be configured according to actual business needs and is not limited here. For example, the encoding method may include at least one of the following: global average pooling, maximum pooling, image pyramid, downsampling, convolution dimensionality reduction and sparse coding, etc. The thirty-seventh intermediate feature map is more inclined to express basic feature units such as points, lines, edge contours, etc., and contains more spatial information.
[0252] According to an embodiment of the present disclosure, the decoding operation can reconvert the encoded feature representation into the original feature map. After obtaining the thirty-seventh intermediate feature map, the H second and third parallel levels can be used to perform a decoding operation on the thirty-seventh intermediate feature map to obtain a thirty-fifth intermediate feature map corresponding to each of the H second and third parallel levels. The specific decoding method can be configured according to actual business needs and is not limited here. For example, the decoding method may include at least one of the following: inverse pooling, deconvolution, bilinear interpolation, and upsampling, etc. The thirty-fifth intermediate feature map is more inclined to express the semantic information of the image and contains more semantic information.
[0253] According to the embodiments of the present disclosure, by encoding the first image to be restored to obtain the thirty-seventh intermediate feature map, the first image to be restored can be converted into a higher-level and more abstract feature representation, thereby better capturing the semantic and structural information of the first image to be restored. By decoding the thirty-seventh intermediate feature map to obtain the thirty-fifth intermediate feature map, the abstract feature representation can be converted back into a representation of the first image to be restored, thereby restoring the detail and shape information in the first image to be restored. Furthermore, by fusing and integrating different feature information at multiple levels based on the thirty-fifth intermediate feature maps corresponding to each of the H third parallel levels, a more accurate target first image to be restored can be obtained.
[0254] According to an embodiment of the present disclosure, the plurality of first cascade levels includes I, where I is an integer greater than 1.
[0255] According to an embodiment of the present disclosure, performing multiple first cascade level processing on the target first image to be repaired and the image quality assessment vector to obtain a thirty-fourth intermediate feature map corresponding to each of the multiple first cascade levels may include the following operations.
[0256] When 1<i≤I, the (i-1)th (i-39)th intermediate feature map is obtained based on the (i-1)th (i-38)th intermediate feature map, where the (i-1)th (i-38)th intermediate feature map is obtained based on the target first image to be restored and the image quality assessment vector. A fifteenth fusion operation is performed on the (i-39)th intermediate feature map and the image quality assessment vector to obtain the (i-34)th intermediate feature map. i is an integer greater than or equal to 1 and less than or equal to I. For example, I=4.
[0257] According to an embodiment of the present disclosure, by performing downsampling processing on the target first image to be restored and the image quality assessment vector at multiple first cascade levels, a thirty-fourth intermediate feature map corresponding to each of the multiple first cascade levels can be obtained. The downsampling processing can include, for example, convolution and maximum pooling.
[0258] According to an embodiment of the present disclosure, for example, when i=1, the target first image to be restored and the image quality assessment vector can be concatenated to obtain the 1st to 38th intermediate feature map. Alternatively, when 1<i≤I, for the i-th first cascade level among the I first cascade levels, the i-th to 39th intermediate feature map can be obtained based on the i-1th to 38th intermediate feature map. For example, the i-1th to 38th intermediate feature map can be convolved to obtain the i-th to 39th intermediate feature map.
[0259] According to an embodiment of the present disclosure, the description of the fifteenth fusion operation can be found in the above-mentioned relevant content regarding the second fusion operation, and will not be repeated here. After obtaining the (i)th thirty-ninth intermediate feature map, the (i)th thirty-ninth intermediate feature map and the image quality assessment vector can be subjected to the fifteenth fusion operation using the fifteenth fusion method to obtain the (i)th thirty-fourth intermediate feature map. For example, the (i)th thirty-ninth intermediate feature map and the image quality assessment vector can be subjected to maximum pooling to obtain the (i)th thirty-fourth intermediate feature map.
[0260] According to an embodiment of the present disclosure, by obtaining the i-th thirty-ninth intermediate feature map based on the i-1 thirty-eighth intermediate feature map, and performing the fifteenth fusion operation on the i-th thirty-ninth intermediate feature map and the image quality assessment vector, the effect of image restoration or enhancement can be enhanced to obtain the i-th thirty-fourth intermediate feature map.
[0261] According to an embodiment of the present disclosure, the plurality of second cascade levels includes 1.
[0262] According to an embodiment of the present disclosure, performing multiple second cascade level processing on the thirty-fourth intermediate feature map corresponding to each of the multiple first cascade levels to obtain a first repaired image may include the following operations.
[0263] When 1≤i<I, an i-th thirty-fifth intermediate feature map is obtained based on the i+1th thirty-fifth intermediate feature map and the i-th thirty-fourth intermediate feature map, where the i-th thirty-fifth intermediate feature map is obtained based on the i-th thirty-fourth intermediate feature map. A first restored image is obtained based on the i-th thirty-fifth intermediate feature map.
[0264] According to an embodiment of the present disclosure, a first restored image may be obtained by performing upsampling processing of multiple second cascade levels on the thirty-fourth intermediate feature maps corresponding to each of the multiple first cascade levels. The upsampling processing may include, for example, deconvolution processing and feature fusion processing.
[0265] According to an embodiment of the present disclosure, for example, when i=I, the i-th-thirty-fifth intermediate feature map can be obtained based on the i-th-thirty-fourth intermediate feature map. Alternatively, when 1≤i<I, for the i-th second cascade level in the I second cascade levels, the i-th-thirty-fifth intermediate feature map can be obtained based on the i+1-th thirty-fifth intermediate feature map and the i-th thirty-fourth intermediate feature map. For example, the i+1-th thirty-fifth intermediate feature map and the i-th thirty-fourth intermediate feature map can be fused to obtain the i-th thirty-fifth intermediate feature map. After obtaining the i-th-thirty-fifth intermediate feature map, the first restored image can be obtained based on the i-th thirty-fifth intermediate feature map.
[0266] According to an embodiment of the present disclosure, obtaining the i-th thirty-ninth intermediate feature map according to the (i-1)th thirty-eighth intermediate feature map may include the following operations.
[0267] Based on the (i-1)th 38th intermediate feature map, obtain the (i)th 40th intermediate feature map. Process the (i)th 40th intermediate feature map based on the fifth self-attention strategy to obtain the (i)th self-attention feature map. Based on the (i)th self-attention feature map, obtain the (i)th 41st intermediate feature map. Based on the (i)th self-attention feature map and the (i)th 41st intermediate feature map, obtain the (i)th 39th intermediate feature map.
[0268] According to an embodiment of the present disclosure, after obtaining the (i-1)th 38th intermediate feature map, the (i-1)th 40th intermediate feature map can be obtained based on the (i-1)th 38th intermediate feature map. For example, a convolution process can be performed on the (i-1)th 38th intermediate feature map to obtain the (i-1)th 40th intermediate feature map.
[0269] According to an embodiment of the present disclosure, the fifth self-attention strategy can be used to learn the correlation between different locations in the image and capture the global context information of the image. The fifth self-attention strategy can include at least one of a split operation, a flattening operation, an addition operation, and a processing operation.
[0270] For example, after obtaining the i-th 40th intermediate feature map, a partitioning operation can be performed on the i-th 40th intermediate feature map to obtain multiple intermediate sub-feature map blocks (i.e., patches) of fixed size. A flattening operation is performed on the multiple intermediate sub-feature map blocks to obtain first intermediate sub-feature vectors corresponding to each of the multiple intermediate sub-feature map blocks, so as to arrange the pixels in the multiple intermediate sub-feature map blocks in a certain order.
[0271] On this basis, an addition operation can be performed on the multiple first intermediate sub-feature vectors. For example, the operation includes a predetermined identification addition operation and a position coding addition operation to obtain a second intermediate sub-feature vector corresponding to each of the multiple first intermediate sub-feature vectors. The predetermined identification addition operation may refer to adding an identification vector to the beginning of the first intermediate sub-feature vector corresponding to each of the multiple intermediate sub-feature blocks, so as to represent the classification information of the entire image. The position coding addition operation may refer to adding a position coding vector to the first intermediate sub-feature vector corresponding to each of the multiple intermediate sub-feature blocks, so as to provide position information of the image block in the original image, thereby perceiving the global context.
[0272] According to an embodiment of the present disclosure, a plurality of second intermediate sub-feature vectors may be processed to obtain an i-th self-attention feature map. For example, a plurality of second intermediate sub-feature vectors may be encoded to obtain an i-th self-attention feature map. After obtaining the i-th self-attention feature map, an i-th forty-first intermediate feature map may be obtained based on the i-th self-attention feature map. For example, the i-th self-attention feature map may be decoded to obtain an i-th forty-first intermediate feature map.
[0273] According to an embodiment of the present disclosure, after obtaining the i-th 41st intermediate feature map, the i-th 39th intermediate feature map can be obtained based on the i-th self-attention feature map and the i-th 41st intermediate feature map. For example, the i-th self-attention feature map and the i-th 41st intermediate feature map can be fused to obtain the i-th 39th intermediate feature map.
[0274] According to an embodiment of the present disclosure, processing the i-th fortieth intermediate feature map based on the fifth self-attention strategy to obtain the i-th self-attention feature map may include one of the following.
[0275] Process the i-th 40th intermediate feature map based on the multi-head self-attention strategy to obtain the i-th self-attention feature map. Process the i-th 40th intermediate feature map based on the window multi-head self-attention strategy to obtain the i-th multi-head self-attention feature map. Based on the i-th multi-head self-attention feature map, obtain the i-th 42nd intermediate feature map. Based on the i-th multi-head self-attention feature map and the i-th 42nd intermediate feature map, obtain the i-th 43rd intermediate feature map. Process the i-th 43rd intermediate feature map based on the shifted window multi-head self-attention strategy to obtain the i-th self-attention feature map.
[0276] According to embodiments of the present disclosure, a multi-head self-attention strategy can explicitly model the input sequence by learning the correlation between each position in the input feature sequence and other positions, thereby improving the expressiveness of the feature map. For example, the i-th self-attention feature map can be processed based on the multi-head self-attention strategy to obtain the i-th self-attention feature map.
[0277] According to an embodiment of the present disclosure, the windowed multi-head self-attention strategy refers to a windowed strategy introduced on the basis of multi-head self-attention processing, which can divide the input feature sequence into small blocks of fixed size and calculate the attention correlation between each small block.
[0278] For example, based on the window multi-head self-attention strategy, a window partitioning operation can be performed on the i-th 40th intermediate feature map to obtain non-overlapping small blocks, namely the i-th multi-head self-attention feature map. Intra-block attention calculation is performed on the i-th multi-head self-attention feature map to obtain the i-th 42nd intermediate feature map for capturing the attention correlation between positions within the window. After obtaining the i-th 42nd intermediate feature map, a cross-block attention calculation is performed on the i-th multi-head self-attention feature map and the i-th 42nd intermediate feature map to obtain the i-th 43rd intermediate feature map for capturing the attention correlation between the current window and its surrounding windows.
[0279] According to an embodiment of the present disclosure, a shifted window multi-head self-attention strategy may refer to a strategy that introduces a shift operation on the basis of a window multi-head self-attention strategy, and can shift each window on the input feature sequence so that the position at the window boundary can also perform attention calculation with adjacent windows. After obtaining the i-th 43rd intermediate feature map, the i-th 43rd intermediate feature map can be processed based on the shifted window multi-head self-attention strategy to obtain an i-th self-attention feature map that can better capture the correlation at the window boundary.
[0280] According to an embodiment of the present disclosure, by processing the i-th forty-second intermediate feature map based on a multi-head self-attention strategy, it is possible to highlight important information at different positions in the feature map and enhance the interaction between them, thereby obtaining the i-th self-attention feature map. By processing the i-th forty-second intermediate feature map based on a window multi-head self-attention strategy, a window mechanism is introduced, that is, only focusing on a specific local area, which can further extract relevant information within a specific window and enhance the expression of this information, thereby obtaining the i-th multi-head self-attention feature map. On this basis, by processing the i-th forty-third intermediate feature map based on a shifted window multi-head self-attention strategy, a shift operation is introduced within the window, which can better capture the relationship between local features and context, thereby improving the expressive power and accuracy of the i-th self-attention feature map.
[0281] According to an embodiment of the present disclosure, performing a restoration operation on the first image to be restored according to the image quality evaluation vector to obtain the first restored image may include the following operations.
[0282] The first image to be inpainted is processed through J fifth cascade levels (i.e., the third cascade level) to obtain a 44th intermediate feature map (i.e., the 17th intermediate feature map) corresponding to each of the J fifth cascade levels, where J is an integer greater than 1. A 45th intermediate feature map (i.e., the 18th intermediate feature map) is obtained based on the J-th 44th intermediate feature map and the image quality assessment vector. A first inpainted image is obtained based on the 45th intermediate feature map and the 44th intermediate feature maps corresponding to each of the J fifth cascade levels.
[0283] According to an embodiment of the present disclosure, the J fifth cascade levels may include J first fifth cascade levels and J second fifth cascade levels. After obtaining the first image to be restored, the J first fifth cascade levels may be used to perform encoding processing on the first image to be restored, thereby obtaining forty-fourth intermediate feature maps corresponding to each of the J fifth cascade levels.
[0284] According to an embodiment of the present disclosure, after obtaining the forty-fourth intermediate feature map corresponding to each of the J fifth cascade levels, self-attention processing can be performed on the J-th forty-fourth intermediate feature map and the image quality assessment vector, and the self-attention mechanism can be used to perform global modeling on the forty-fourth intermediate feature map corresponding to each of the J fifth cascade levels to obtain a forty-fifth intermediate feature map.
[0285] According to an embodiment of the present disclosure, after obtaining the 45th intermediate feature map, the J second to fifth cascade levels can be used to decode the 45th intermediate feature map and the 44th intermediate feature maps corresponding to the J fifth cascade levels, so as to map the globally modeled 45th intermediate feature map to the same size as the first image to be repaired, thereby obtaining a first repaired image.
[0286] According to an embodiment of the present disclosure, by processing the first image to be repaired using J fifth cascade levels, information between different levels can be fully utilized to obtain a 44th intermediate feature map corresponding to each of the J fifth cascade levels. By further optimizing the quality and expression of the feature map based on the J-th 44th intermediate feature map and the image quality assessment vector, a 45th intermediate feature map can be obtained. On this basis, by integrating the 45th intermediate feature map with the 44th intermediate feature map corresponding to each of the J fifth cascade levels, feature information from different levels can be fused, thereby helping to improve the image repair effect and the quality of the first repaired image.
[0287] FIG5 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0288] As shown in FIG5 , in step 500 , an image quality assessment operation can be performed on a first image to be restored 503 to obtain image quality assessment information 501. Based on image quality assessment information 501, an image quality assessment vector 502 is obtained. The first image to be restored 503 is processed through J fifth cascade levels to obtain a 44th intermediate feature map 504 corresponding to each of the J fifth cascade levels. Based on the Jth 44th intermediate feature map and the image quality assessment vector 504, a 45th intermediate feature map 505 is obtained. Based on the 45th intermediate feature map 505 and the 44th intermediate feature maps 504 corresponding to each of the J fifth cascade levels, a first restored image 506 is obtained.
[0289] According to an embodiment of the present disclosure, a forty-fifth intermediate feature map is obtained according to the forty-fourth intermediate feature map and the image quality assessment vector, including:
[0290] When 1<k≤K, a k-th forty-seventh intermediate feature map is obtained based on the k-1th forty-sixth intermediate feature map, where the k-th forty-sixth intermediate feature map is obtained based on the j-th forty-fourth intermediate feature map and the image quality assessment vector. A sixteenth fusion operation is performed on the k-th forty-seventh intermediate feature map and the image quality assessment vector to obtain a k-th forty-sixth intermediate feature map. A forty-fifth intermediate feature map is obtained based on the k-th forty-sixth intermediate feature map. K is an integer greater than 1, and k is an integer greater than or equal to 1 and less than or equal to K.
[0291] According to an embodiment of the present disclosure, when k=1, the jth 44th intermediate feature map and the image quality assessment vector may be concatenated to obtain the jth 46th intermediate feature map. When 1<k≤K, the kth 47th intermediate feature map may be obtained based on the k-1th 46th intermediate feature map. For example, the k-1th 46th intermediate feature map may be convolved to obtain the kth 47th intermediate feature map.
[0292] According to an embodiment of the present disclosure, the description of the sixteenth fusion operation can be found in the above-mentioned relevant content regarding the second fusion operation, and will not be repeated here. After obtaining the k-th forty-seventh intermediate feature map, the above-mentioned sixteenth fusion method can be used to perform the sixteenth fusion operation on the k-th forty-seventh intermediate feature map and the image quality assessment vector to obtain the k-th forty-sixth intermediate feature map.
[0293] According to an embodiment of the present disclosure, by performing the sixteenth fusion operation on the kth forty-seventh intermediate feature map and the image quality assessment vector, a more excellent kth forty-sixth intermediate feature map that combines the kth forty-seventh intermediate feature map and the image quality assessment vector can be obtained. On this basis, by using the kth forty-sixth intermediate feature map, the expressive power of the forty-fifth intermediate feature map can be improved, thereby facilitating improved effects of subsequent image restoration.
[0294] According to an embodiment of the present disclosure, obtaining a first repaired image according to the forty-fifth intermediate feature map and the forty-fourth intermediate feature maps corresponding to each of the J fifth cascade levels includes:
[0295] When 1≤j<J, a j-th 48th intermediate feature map is obtained based on the j+1th 48th intermediate feature map and the j-th 44th intermediate feature map, where the j-th 48th intermediate feature map is obtained based on the 45th intermediate feature map. A restored image is obtained based on the j-th 48th intermediate feature map. j is an integer greater than or equal to 1 and less than or equal to J.
[0296] According to an embodiment of the present disclosure, a first restored image may be obtained by upsampling the 45th intermediate feature map and the 44th intermediate feature maps corresponding to the J fifth cascade levels. The upsampling process may include, for example, deconvolution and feature fusion.
[0297] According to an embodiment of the present disclosure, for example, when j=J, the jth forty-eighth intermediate feature map can be determined based on the forty-fifth intermediate feature map. Alternatively, when 1≤j<J, for the jth fifth cascade level among the J fifth cascade levels, the jth forty-eighth intermediate feature map can be obtained based on the j+1th forty-eighth intermediate feature map and the jth forty-fourth intermediate feature map. For example, the j+1th forty-eighth intermediate feature map and the jth forty-fourth intermediate feature map can be fused to obtain the jth forty-eighth intermediate feature map. After obtaining the jth forty-eighth intermediate feature map, the first restored image can be obtained based on the first forty-eighth intermediate feature map.
[0298] According to an embodiment of the present disclosure, performing a restoration operation on a first image to be restored according to an image quality assessment vector to obtain a first restored image includes:
[0299] Based on the first image to be restored and the image quality assessment vector, forty-ninth intermediate feature maps (i.e., nineteenth intermediate feature maps) corresponding to L sixth cascade levels (i.e., fourth cascade levels) are obtained, where L is an integer greater than 1. Based on the Lth forty-ninth intermediate feature maps, a fiftieth intermediate feature map (i.e., twentieth intermediate feature map) is obtained. Based on the fiftieth intermediate feature map and the forty-ninth intermediate feature maps corresponding to the L sixth cascade levels, a first restored image is obtained.
[0300] According to an embodiment of the present disclosure, the L sixth cascade levels may include L first sixth cascade levels and L second sixth cascade levels. After obtaining the first image to be restored, the first image to be restored and the image quality assessment vector may be encoded using the L first sixth cascade levels to obtain forty-ninth intermediate feature maps corresponding to each of the L sixth cascade levels.
[0301] According to an embodiment of the present disclosure, after obtaining the forty-ninth intermediate feature map corresponding to each of the L sixth cascade levels, self-attention processing can be performed on the Lth forty-ninth intermediate feature map, and the self-attention mechanism can be used to perform global modeling on the Lth forty-ninth intermediate feature map to obtain a fiftieth intermediate feature map.
[0302] According to an embodiment of the present disclosure, after obtaining the fiftieth intermediate feature map, the L second-sixth cascade levels can be used to decode the fiftieth intermediate feature map and the forty-ninth intermediate feature maps corresponding to the L sixth cascade levels, so as to map the globally modeled fiftieth intermediate feature map to the same size as the first image to be repaired, thereby obtaining a first repaired image.
[0303] According to an embodiment of the present disclosure, by capturing important information at different levels of the first image to be restored and the image quality assessment vector based on the first image to be restored, the forty-ninth intermediate feature map corresponding to each of the L sixth cascade levels is obtained. By obtaining a fiftieth intermediate feature map based on the Lth forty-ninth intermediate feature map, and based on the fiftieth intermediate feature map and the forty-ninth intermediate feature map corresponding to each of the L sixth cascade levels, a first restored image can be obtained based on the correlation between the feature maps and the feature information extracted in the previous processing steps, thereby improving the quality of the first restored image and the image restoration effect.
[0304] FIG6 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0305] As shown in FIG6 , in step 600 , an image quality assessment operation can be performed on a first image to be restored 603 to obtain image quality assessment information 601. Based on image quality assessment information 601, an image quality assessment vector 602 is obtained. Based on the first image to be restored 603 and the image quality assessment vector 602, a 49th intermediate feature map 604 corresponding to each of the L sixth cascade levels is obtained. Based on the Lth 49th intermediate feature map 604, a 50th intermediate feature map 605 is obtained. Based on the 50th intermediate feature map 605 and the 49th intermediate feature map 604 corresponding to each of the L sixth cascade levels, a first restored image 606 is obtained.
[0306] According to an embodiment of the present disclosure, based on the first image to be repaired and the image quality assessment vector, a forty-ninth intermediate feature map corresponding to each of the L sixth cascade levels is obtained, including:
[0307] When 1<l≤L, a l-th fifty-first intermediate feature map is obtained based on the l-th forty-ninth intermediate feature map, where the l-th forty-ninth intermediate feature map is obtained based on the first image to be restored and the image quality assessment vector. A seventeenth fusion operation is performed on the l-th fifty-first intermediate feature map and the image quality assessment vector to obtain a l-th forty-ninth intermediate feature map. l is an integer greater than or equal to 1 and less than or equal to L.
[0308] According to an embodiment of the present disclosure, when l = 1, the first image to be restored and the image quality assessment vector may be concatenated to obtain the 1st to 49th intermediate feature map. When 1 < l ≤ L, the 1st to 51st intermediate feature map may be obtained based on the 1-1st to 49th intermediate feature map. For example, the 1-1st to 49th intermediate feature map may be convolved to obtain the 1st to 51st intermediate feature map.
[0309] According to an embodiment of the present disclosure, the description of the seventeenth fusion operation can be found in the above-mentioned related content regarding the second fusion operation, and will not be repeated here. After obtaining the first fifty-first intermediate feature map, the seventeenth fusion method can be used to perform the seventeenth fusion operation on the first fifty-first intermediate feature map and the image quality assessment vector to obtain the first forty-ninth intermediate feature map.
[0310] According to the embodiments of the present disclosure, by obtaining the (1-1)th forty-ninth intermediate feature map from the (1-1)th forty-ninth intermediate feature map, information in the (1-1)th forty-ninth intermediate feature map can be further extracted and optimized. Furthermore, by performing the seventeenth fusion operation on the (1-1)th forty-first intermediate feature map and the image quality assessment vector, useful information in the (1-1)th forty-first intermediate feature map can be captured to a certain extent, and the expressiveness of the features can be further improved, resulting in a more accurately expressed (1-1)th forty-ninth intermediate feature map.
[0311] According to an embodiment of the present disclosure, performing a restoration operation on a first image to be restored according to an image quality assessment vector to obtain a first restored image includes:
[0312] When 1<m≤M, the mth fifty-third intermediate feature map (mth twenty-second intermediate feature map) is obtained based on the (1st) fifty-second intermediate feature map (i.e., the (1st) twenty-first intermediate feature map) to the (m-1st) fifty-second intermediate feature map (m-1st twenty-first intermediate feature map), where the (1st) fifty-second intermediate feature map (i.e., the (1st) twenty-first intermediate feature map) is obtained based on the (1st) fifty-third intermediate feature map (i.e., the (1st) twenty-second intermediate feature map) and the image quality assessment vector, and the (1st) fifty-third intermediate feature map is obtained based on the first image to be restored. The (mth) fifty-third intermediate feature map and the image quality assessment vector are subjected to an eighteenth fusion operation to obtain the (mth) fifty-second intermediate feature map (i.e., the (mth) twenty-first intermediate feature map). A first restored image is obtained based on the (Mth) fifty-second intermediate feature map. M is an integer greater than 1, and m is an integer greater than or equal to 1 and less than or equal to M.
[0313] According to an embodiment of the present disclosure, when m=1, a convolution process can be performed on the first image to be restored to obtain a 1st 53rd intermediate feature map. The 1st 53rd intermediate feature map and the image quality assessment vector are concatenated to obtain a 1st 52nd intermediate feature map. When 1<m≤M, an mth 53rd intermediate feature map can be obtained based on the 1st 52nd intermediate feature map to the m-1th 52nd intermediate feature map.
[0314] For example, the 1st 52nd intermediate feature map can be used as the 1st 53rd intermediate feature map, and the 1st 53rd intermediate feature map can be input into the 1st processing layer to obtain the 2nd 53rd intermediate feature map. The 1st 53rd intermediate feature map and the 2nd 53rd intermediate feature map can be input into the 2nd processing layer to obtain the 3rd 53rd intermediate feature map. Similarly, the 1st 53rd intermediate feature map to the (m-1)th 53rd intermediate feature map can be input into the (m-1)th dense layer to obtain the (m)th 53rd intermediate feature map.
[0315] According to the embodiments of the present disclosure, the description of the eighteenth fusion operation can be found in the above-mentioned related content regarding the second fusion operation, and will not be repeated here. After obtaining the (m)th fifty-third intermediate feature map, the above-mentioned eighteenth fusion method can be used to perform the eighteenth fusion operation on the (m)th fifty-third intermediate feature map and the image quality assessment vector to obtain the (m)th fifty-second intermediate feature map. Based on this, the first restored image can be obtained based on the (M)th fifty-second intermediate feature map.
[0316] FIG7 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0317] As shown in FIG7 , in step 700 , an image quality assessment operation is performed on a first image to be restored 707 to obtain image quality assessment information 701. Based on image quality assessment information 701, an image quality assessment vector 702 is obtained. Based on the first image to be restored 707, a first fifty-third intermediate feature map 708 is obtained. Based on the first fifty-third intermediate feature map 708 and image quality assessment vector 702, a first fifty-second intermediate feature map 704 is obtained.
[0318] When 1<m≤M, an m-th fifty-third intermediate feature map 706 is obtained based on the (1st) fifty-second intermediate feature map 704 to the (m-1st) fifty-second intermediate feature map 705. An eighteenth fusion operation is performed on the (m)th fifty-third intermediate feature map 706 and the image quality assessment vector 702 to obtain an (m)th fifty-second intermediate feature map 709. A first restored image 710 is obtained based on the (M)th fifty-second intermediate feature map 709.
[0319] According to an embodiment of the present disclosure, performing a restoration operation on a first image to be restored according to an image quality assessment vector to obtain a first restored image includes:
[0320] At least one scaled sub-image to be restored is obtained based on the first image to be restored. A step-by-step restoration operation is performed on the at least one scaled sub-image to be restored based on the image quality assessment vector to obtain a first restored image.
[0321] According to embodiments of the present disclosure, scale may refer to image resolution. Each scale may have at least one sub-image to be restored corresponding to that scale. After obtaining a first image to be restored, a scale division operation may be performed on the first image to be restored to obtain multiple sub-images to be restored of at least one fixed size. Specific scale division methods may include at least one of the following: a scale division method based on a high-pass filter, a scale division method based on a Basel wavelet transform, and a scale division method based on a pyramid decomposition.
[0322] According to an embodiment of the present disclosure, a step-by-step restoration operation may refer to a process of dividing an image into multiple levels, performing partial filling and reconstruction at each level, and passing the results to the next level for further processing. A damaged image can be reconstructed by filling in missing areas. After obtaining a sub-image to be restored at at least one scale, a step-by-step restoration operation may be performed on the sub-image to be restored at at least one scale based on an image quality assessment vector to obtain a first restored image. For example, a step-by-step restoration method may include at least one of the following: a diffusion-based image restoration method, a block-based image restoration method, and a deep learning-based image restoration method.
[0323] According to an embodiment of the present disclosure, by obtaining a sub-image to be repaired of at least one scale based on a first image to be repaired, performing step-by-step repair operations on the sub-images to be repaired of different scales, and fusing the final results, it is possible to meet image quality requirements of different needs while ensuring the repair effect.
[0324] FIG8 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored based on image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0325] As shown in FIG8 , in step 800, an image quality assessment operation can be performed on a first image to be restored 803 to obtain image quality assessment information 801. Based on image quality assessment information 801, an image quality assessment vector 802 is obtained. Based on the first image to be restored 803, a sub-image to be restored of at least one scale 804 is obtained. Based on image quality assessment vector 802, a step-by-step restoration operation is performed on the sub-image to be restored of at least one scale 804 to obtain a first restored image 805.
[0326] According to an embodiment of the present disclosure, performing a step-by-step restoration operation on a sub-image to be restored of at least one scale according to an image quality assessment vector to obtain a first restored image includes:
[0327] When 1<n≤N, an nth first fused image is obtained based on the (n-1)th third intermediate restored image and the nth sub-image to be restored, where the first first fused image is obtained by performing a first restoration sub-operation on the first sub-image to be restored based on the image quality assessment vector. A second restoration sub-operation is performed on the nth first fused image based on the image quality assessment vector to obtain an nth third intermediate restored image. A first restored image is obtained based on the Nth third intermediate restored image. N is an integer greater than 1, and n is an integer greater than or equal to 1 and less than or equal to N.
[0328] According to an embodiment of the present disclosure, the first restoration sub-operation may refer to a restoration operation performed at the lowest level. The first restoration sub-operation may include a diffusion-based image restoration method. For example, when n=1, the first restoration sub-operation may be performed on the first sub-image to be restored based on the image quality assessment vector to obtain a first fused image.
[0329] According to an embodiment of the present disclosure, the second restoration sub-operation may refer to an operation that transfers the restoration structure at the lowest level, i.e., the first fused image, to a subsequent level for more accurate filling. The second restoration sub-operation may include a block-based image restoration method and a deep learning-based image restoration method. For example, when 1<n≤N, the nth sub-image to be restored may be fused with the n-1th third intermediate restoration image to obtain the nth first fused image.
[0330] According to embodiments of the present disclosure, after obtaining the nth third intermediate restoration image, all of them can be combined to obtain the final first restoration image. Based on this, subsequent processing such as color adjustment and smoothing can be performed on the Nth third intermediate restoration image to improve the naturalness and coherence of the first restoration image.
[0331] According to the embodiments of the present disclosure, in each iterative step, a higher-level nth first fused image is gradually generated by associating or transforming the n-1th third intermediate inpainted image with the nth sub-image to be inpainted. Based on this, a second inpainting sub-operation is performed on the nth first fused image in conjunction with the image quality assessment vector to obtain the nth third intermediate inpainted image. The final first inpainted image is then generated based on the Nth third intermediate inpainted image, thereby improving the accuracy and quality of the first inpainted image.
[0332] According to an embodiment of the present disclosure, performing a restoration operation on a first image to be restored according to an image quality assessment vector to obtain a first restored image includes:
[0333] A third restoration sub-operation (i.e., the first restoration sub-operation) is performed on the first image to be restored based on the image quality assessment vector, obtaining a fourth intermediate restoration image. A second fused image is obtained based on the first image to be restored and the fourth intermediate restoration image. A fourth restoration sub-operation (i.e., the second restoration sub-operation) is performed on the second fused image based on the image quality assessment vector, obtaining a fifth intermediate restoration image. A third fused image is obtained based on the first image to be restored and the fifth intermediate restoration image. A fifth restoration sub-operation (i.e., the third restoration sub-operation) is performed on the third fused image based on the image quality assessment vector, obtaining a first restoration image.
[0334] According to an embodiment of the present disclosure, after obtaining a damaged RGB three-channel color first image to be repaired, a third repair sub-operation can be performed on the first image to be repaired based on the image quality evaluation vector to obtain a fourth intermediate repair image of single-channel gradient information for characterizing the image structure of the damaged area.
[0335] According to embodiments of the present disclosure, after obtaining the fourth intermediate inpainted image, a second fused image can be obtained based on the damaged RGB three-channel color image to be inpainted and the fourth intermediate inpainted image representing the single-channel gradient information of the image structure of the damaged area. Based on this, a fourth inpainting sub-operation can be performed on the second fused image based on the image quality assessment vector to obtain a fifth intermediate inpainted image representing the single-channel grayscale image of the damaged area.
[0336] According to an embodiment of the present disclosure, after obtaining the fifth intermediate inpainted image, a third fused image can be obtained based on the damaged first inpainted image (RGB three-channel color) and the fifth intermediate inpainted image (a single-channel grayscale image representing the damaged area). The fifth inpainting sub-operation is performed on the third fused image based on the image quality assessment vector to obtain the first inpainted image (RGB three-channel color) of the damaged area.
[0337] According to the embodiments of the present disclosure, through multiple restoration sub-operations and fusion operations, the image quality assessment vector and information of the first image to be restored can be fully utilized to gradually generate the first restored image, which is conducive to improving the quality and integrity of the first restored image.
[0338] FIG9 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0339] As shown in FIG9 , in step 900 , an image quality assessment operation can be performed on a first image to be restored 903 to obtain image quality assessment information 901. Based on image quality assessment information 901, an image quality assessment vector 902 is obtained. A third restoration sub-operation is performed on the first image to be restored 903 based on image quality assessment vector 902 to obtain a fourth intermediate restoration image 904. Based on the first image to be restored 903 and the fourth intermediate restoration image 904, a second fused image 905 is obtained. Based on the image quality assessment vector 902, a fourth restoration sub-operation is performed on the second fused image 905 to obtain a fifth intermediate restoration image 906. Based on the first image to be restored 903 and the fifth intermediate restoration image 906, a third fused image 907 is obtained. Based on the image quality assessment vector 902, a fifth restoration sub-operation is performed on the third fused image 907 to obtain a first restoration image 908.
[0340] According to an embodiment of the present disclosure, performing a restoration operation on a first image to be restored according to an image quality evaluation vector to obtain a first restored image may include the following operations.
[0341] A sixth intermediate restoration image is obtained based on the first image to be restored. A second structural image is obtained based on the sixth intermediate restoration image. A seventh fused image is obtained based on the sixth intermediate restoration image and the second structural image. A sixth restoration sub-operation (i.e., the fourth restoration sub-operation) is performed on the seventh fused image based on the image quality assessment vector to obtain a first restoration image.
[0342] According to an embodiment of the present disclosure, after obtaining the first image to be restored, a rough sixth intermediate restoration image can be obtained based on the first image to be restored. Based on this, the rough sixth intermediate restoration image can be refined to obtain the final restoration result, i.e., the first restoration image.
[0343] According to an embodiment of the present disclosure, after obtaining the sixth intermediate restored image, a second structural image can be obtained based on the sixth intermediate restored image. The second structural image can be used to characterize the similarity between any damaged area block (i.e., foreground block) and the known area block (i.e., background block) in the first image to be restored. Based on the sixth intermediate restored image and the second structural image, a seventh fused image can be obtained, in which the output result is constrained by the second structural image. On this basis, the sixth restoration sub-operation can be performed on the seventh fused image based on the image quality assessment vector to obtain the first restored image.
[0344] FIG10 schematically shows an example schematic diagram of a process of performing a restoration operation on a first image to be restored according to image quality assessment information to obtain a first restored image according to another embodiment of the present disclosure.
[0345] As shown in FIG10 , in step 1000, an image quality assessment operation can be performed on a first image to be restored 1003 to obtain image quality assessment information 1001. Based on image quality assessment information 1001, an image quality assessment vector 1002 is obtained. Based on the first image to be restored 1003, a sixth intermediate restored image 1004 is obtained. Based on the sixth intermediate restored image 1004, a second structural image 1005 is obtained. Based on the sixth intermediate restored image 1004 and the second structural image 1005, a seventh fused image 1006 is obtained. Based on the image quality assessment vector 1002, a sixth restoration sub-operation is performed on the seventh fused image 1006 to obtain a first restored image 1007.
[0346] According to an embodiment of the present disclosure, operation S210 includes the following operations.
[0347] According to the restoration type corresponding to the first image to be restored, an evaluation interface corresponding to the restoration type is called, and an image quality evaluation operation is performed on the first image to be restored using the evaluation interface to obtain image quality evaluation information.
[0348] According to an embodiment of the present disclosure, a correspondence between different restoration types and evaluation interfaces may be preconfigured. The evaluation interface may be used to evaluate image quality corresponding to the restoration type.
[0349] According to an embodiment of the present disclosure, after obtaining a first image to be restored, a restoration type corresponding to the first image to be restored can be determined. The restoration type can include at least one of the following: image denoising, image deblurring, background / text removal, inpainting, rectangle restoration, color restoration, or foreground / object removal.
[0350] According to an embodiment of the present disclosure, after determining the restoration type corresponding to the first image to be restored, an evaluation interface corresponding to the restoration type can be called based on the restoration type of the first image to be restored, so as to perform an image quality evaluation operation on the first image to be restored using the evaluation interface to obtain image quality evaluation information. The evaluation interface can be used to evaluate the image quality corresponding to the restoration type.
[0351] According to an embodiment of the present disclosure, since the evaluation interface is determined based on the restoration type corresponding to the first image to be restored, by utilizing the evaluation interface to perform an image quality evaluation operation on the first image to be restored, an automated evaluation of the first image to be restored can be achieved to obtain image quality evaluation information, thereby improving the efficiency of information processing.
[0352] FIG11 schematically shows an example diagram of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to an embodiment of the present disclosure.
[0353] As shown in FIG11 , in step 1100 , an evaluation interface 1103 corresponding to the restoration type 1102 of the first image to be restored 1101 may be called according to the restoration type 1102 corresponding to the first image to be restored 1101. The evaluation interface 1103 is used to perform an image quality evaluation operation on the first image to be restored 1101 to obtain image quality evaluation information 1104.
[0354] According to an embodiment of the present disclosure, operation S210 may include one of the following.
[0355] A first feature extraction operation is performed on the first image to be repaired to obtain a shared feature map; based on the shared feature map, a fifty-fourth intermediate feature map corresponding to each of the at least one image quality assessment items is obtained; and based on the at least one fifty-fourth intermediate feature map, an image quality assessment value corresponding to each of the at least one image quality assessment items is obtained. A second feature extraction operation is performed on the first image to be repaired to obtain a fifty-fifth intermediate feature map; and based on the fifty-fifth intermediate feature map, an image quality assessment value corresponding to each of the at least one image quality assessment items is obtained. A third feature extraction operation is performed on the first image to be repaired to obtain a fifty-sixth intermediate feature map corresponding to each of the at least one image quality assessment items; and based on the at least one fifty-sixth intermediate feature map, an image quality assessment value corresponding to each of the at least one image quality assessment items is obtained.
[0356] According to an embodiment of the present disclosure, the first feature extraction operation, the second feature extraction operation, and the third feature extraction operation can be used to extract representative and highly discriminative feature representations from an image. The specific feature extraction method can be configured according to actual business needs and is not limited here. For example, the feature extraction method may include at least one of the following: a feature extraction method based on SIFT (Scale-Invariant Feature Transform), a feature extraction method based on SURF (Speeded-Up Robust Features), a feature extraction method based on HOG (Histogram of Oriented Gradients), a feature extraction method based on Haar features, a feature extraction method based on CNN, a feature extraction method based on Gabor filters, a feature extraction method based on LBP (Local Binary Patterns), a feature extraction method based on a color histogram, a feature extraction method based on texture features, and a feature extraction method based on depth features.
[0357] According to an embodiment of the present disclosure, after obtaining a first image to be restored, a first feature extraction operation can be performed on the first image to be restored to obtain a shared feature map. A shared feature map may refer to a feature map shared by multiple locations or regions. Based on the shared feature map and at least one image quality assessment item, a fifty-fourth intermediate feature map corresponding to each of the at least one image quality assessment items is determined. Based on this, an image quality assessment value corresponding to each of the at least one image quality assessment items can be obtained based on the at least one fifty-fourth intermediate feature map.
[0358] According to the embodiments of the present disclosure, by obtaining image quality assessment values corresponding to at least one image quality assessment item in different ways, the efficiency and accuracy of obtaining the image quality assessment values can be improved, which is beneficial to the subsequent image restoration of the first image to be restored based on the image quality assessment values, thereby improving the efficiency and accuracy of image restoration.
[0359] FIG12A schematically shows an example of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to an embodiment of the present disclosure.
[0360] As shown in FIG12A , in step 1200A, a first feature extraction operation may be performed on a first image to be restored 1201 to obtain a shared feature map 1202. Based on shared feature map 1202, a fifty-fourth intermediate feature map 1203 corresponding to at least one image quality assessment item is obtained. Based on at least one fifty-fourth intermediate feature map 1203, an image quality assessment value 1204 corresponding to at least one image quality assessment item is obtained.
[0361] FIG12B schematically shows an example diagram of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to another embodiment of the present disclosure.
[0362] As shown in FIG12B , in 1200B, a convolutional layer 1206 can be used to perform a second feature extraction operation on the first image to be restored 1205 to obtain a fifty-fifth intermediate feature map. The fifty-fifth intermediate feature map is then processed using a residual block 1207 to obtain a first intermediate fifty-fifth intermediate feature map. The first intermediate fifty-fifth intermediate feature map is then processed using an average pooling layer 1208 to obtain a second intermediate fifty-fifth intermediate feature map. The second intermediate fifty-fifth intermediate feature map is then processed using a fully connected layer 1209 to obtain an image quality assessment value 1210 corresponding to each of the at least one image quality assessment items.
[0363] FIG12C schematically shows an example diagram of a process of performing an image quality assessment operation on a first image to be restored to obtain image quality assessment information according to another embodiment of the present disclosure.
[0364] As shown in FIG. 12C , in step 1200C, convolutional layers 1212_1, 1212_2, ..., and 1212_A can be used to perform a third feature extraction operation on the first image to be restored 1211, respectively, to obtain fifty-sixth intermediate feature maps corresponding to at least one image quality assessment item. A is a positive integer.
[0365] For the fifty-sixth intermediate feature map obtained by the convolution layer 1212_1, the fifty-sixth intermediate feature map can be processed in sequence using the residual block 1213_1, the average pooling layer 1214_1, and the fully connected layer 1215_1 to obtain an image quality assessment value 1216_1 corresponding to at least one image quality assessment item.
[0366] For the fifty-sixth intermediate feature map obtained by the convolution layer 1212_2, the fifty-sixth intermediate feature map can be processed in sequence using the residual block 1213_2, the average pooling layer 1214_2, and the fully connected layer 1215_2 to obtain image quality assessment values 1216_2 corresponding to at least one image quality assessment item.
[0367] By analogy, for the fifty-sixth intermediate feature map obtained by the convolution layer 1212_A, the fifty-sixth intermediate feature map can be processed in sequence using the residual block 1213_A, the average pooling layer 1214_A, and the fully connected layer 1215_A to obtain an image quality assessment value 1216_A corresponding to at least one image quality assessment item.
[0368] According to an embodiment of the present disclosure, the first image to be restored includes an image of at least one of the following application scenarios: film and television image restoration, cultural relic image restoration, medical image restoration, traffic image restoration, and astronomical image restoration.
[0369] According to an embodiment of the present disclosure, film and television image restoration may include television image restoration, movie image restoration, and library archive image restoration.
[0370] According to the embodiments of the present disclosure, since the video data of relatively old TV programs are usually stored on magnetic tape media, problems such as horizontal lines, field patterns, jagged edges or large areas of mosaics may appear on the screen when it is played again. However, today's high-definition or ultra-high-definition TV programs require that the clarity of the image material meet the standard. Therefore, when the quality of the materials is uneven, especially when the visual quality of the important portrait materials is not up to standard, in order to improve the overall program visual effect, TV image restoration of the TV image is required. The TV image of the TV image restoration may include at least one of the following: current real-shot materials, high-definition and pseudo-high-definition (for example, stretching), standard definition (for example, news), low definition (for example, tape, telecine, film, network materials), image materials (for example, scanning, special effects production, inventory and network compression).
[0371] According to the embodiments of the present disclosure, in the early days of film, movie video data was typically stored on film media. Due to imperfect preservation technology and unsuitable storage environments, the film was subject to certain damage during storage, such as scratches, stains, and mold spots. This also reduced the quality of the stored images, leading to frequent problems such as poor sound quality and blurred images. Therefore, movie image restoration can be performed to ensure that the clarity of the movie image meets standard standards.
[0372] According to the embodiments of the present disclosure, archives contain a large amount of precious historical materials, mostly stored on tapes, films, and paper. Regular repair consumes a lot of manpower and material resources and is costly. Therefore, archives can be restored and digitally scanned and archived to improve efficiency and maximize the integrity of the content.
[0373] According to an embodiment of the present disclosure, cultural relic image restoration may refer to the process of repairing and reconstructing damaged, faded, or damaged images of ancient cultural relics. The goal of cultural relic image restoration is to restore the integrity, color, and details of the image.
[0374] According to embodiments of the present disclosure, medical image restoration can refer to the process of removing noise, eliminating artifacts, and enhancing the quality of medical images. The goal of medical image restoration is to improve image quality and enable accurate diagnosis and analysis by using restoration techniques to address potential noise, artifacts, or motion blur.
[0375] According to embodiments of the present disclosure, traffic image restoration refers to the process of addressing issues in images captured by traffic surveillance cameras. The goal of traffic image restoration is to improve the quality and visualization of traffic images by using restoration techniques to address issues such as raindrops, motion blur, or lighting variations.
[0376] According to embodiments of the present disclosure, astronomical image restoration refers to the process of addressing issues in astronomical images obtained from telescopes or satellites. The goal of astronomical image restoration is to improve noise, blur, or other contamination in images through restoration techniques, thereby enhancing celestial detail, improving image clarity and contrast, and enabling more accurate celestial observation and analysis.
[0377] FIG13 schematically shows an example diagram of an image restoration process according to an embodiment of the present disclosure.
[0378] As shown in FIG13 , in step 1300 , an image quality assessment operation may be performed on a first image to be restored 1301 to obtain image quality assessment information 1302 . A restoration operation may be performed on the first image to be restored 1301 based on the image quality assessment information 1302 to obtain a first restored image 1303 .
[0379] FIG14 schematically shows a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure.
[0380] As shown in FIG14 , the training 1400 of the deep learning model includes operations S1410 to S1430 .
[0381] In operation S1410, an image quality assessment operation is performed on the first sample image to be repaired to obtain first sample image quality assessment information, wherein the first sample image quality assessment information includes first sample image quality assessment values corresponding to at least one first sample image quality assessment item, and the first sample image quality assessment values represent the degree of interference of the first sample image quality assessment item on the first sample image to be repaired.
[0382] In operation S1420 , a restoration operation is performed on the first sample image to be restored according to the first sample image quality evaluation information to obtain a first restored sample image.
[0383] In operation S1430, a first deep learning model is trained using the first restored sample image and a first predetermined sample image corresponding to the first sample image to be restored to obtain a first image restoration model.
[0384] According to an embodiment of the present disclosure, for descriptions of the first sample image to be repaired, the first sample image quality assessment information, the first sample image quality assessment value, the first sample image quality assessment item, and the first repaired sample image, reference can be made to the above-mentioned relevant contents regarding the first image to be repaired, the image quality assessment information, the image quality assessment value, the image quality assessment item, and the first repaired image, which will not be repeated here.
[0385] According to an embodiment of the present disclosure, the first predetermined sample image may include at least one of the following: a first real sample image and a first reference sample image. Specifically, the first real sample image may be a first sample image to be restored. The first reference sample image may be obtained by processing the first sample image to be restored using a teacher network to obtain an auxiliary sample image, and fusing the first sample image to be restored with the auxiliary sample image.
[0386] According to an embodiment of the present disclosure, since the first sample image quality assessment information is obtained by performing an image quality assessment operation on the first sample image to be repaired, the first sample image quality assessment value included in the first sample image quality assessment information can be used to characterize the degree of interference of the first sample image quality assessment item on the first sample image to be repaired. On this basis, by performing a repair operation on the first sample image to be repaired based on the first sample image quality assessment information to obtain a first repaired sample image, and using the first repaired sample image and a first predetermined sample image corresponding to the first sample image to be repaired to train a first deep learning model, the quality of the first repaired sample image can be improved, thereby improving the repair capability of the first image repair model.
[0387] According to an embodiment of the present disclosure, operation S1430 may include the following operations.
[0388] Determine first quantized sample information corresponding to a first sample image to be repaired, wherein the first sample image to be repaired includes at least one first sample image region to be repaired, and the first quantized sample information includes a first quantized value corresponding to each of the at least one first sample image region to be repaired, wherein the first quantized value corresponding to the first sample image region to be repaired represents the importance of the first sample image region to be repaired. Train a first deep learning model using the first repair sample image, a first predetermined sample image corresponding to the first sample image to be repaired, and the first quantized sample information corresponding to the first sample image to be repaired to obtain a first image repair model.
[0389] According to an embodiment of the present disclosure, after obtaining a first sample image to be restored, the first sample image to be restored can be divided into at least one first sample image region to be restored. Based on this, the importance of each of the at least one first sample image region to be restored can be evaluated separately to obtain a first quantization value corresponding to each of the at least one first sample image region to be restored, and the first quantization value corresponding to each of the at least one first sample image region to be restored can be determined as the first quantized sample information.
[0390] According to the embodiments of the present disclosure, since the first quantized sample information corresponds to the first sample image to be repaired, the first quantized value corresponding to the region of the first sample image to be repaired, included in the first quantized sample information, can be used to characterize the importance of the region of the first sample image to be repaired. On this basis, by training a first deep learning model using the first repair sample image, a first predetermined sample image corresponding to the first sample image to be repaired, and the first quantized sample information corresponding to the first sample image to be repaired, the characteristics of the first sample image to be repaired can be better retained during the training process, thereby improving the repair capability of the first image repair model.
[0391] According to an embodiment of the present disclosure, training a first deep learning model using a first repaired sample image, a first predetermined sample image corresponding to the first sample image to be repaired, and first quantized sample information corresponding to the first sample image to be repaired to obtain a first image repair model may include the following operations.
[0392] Based on the first loss function, a first loss function value is obtained based on the first inpainted sample image, a first predetermined sample image corresponding to the first sample image to be inpainted, and first quantized sample information corresponding to the first sample image to be inpainted. Model parameters of the first deep learning model are adjusted based on the first loss function value until a first predetermined termination condition is satisfied. The first deep learning model obtained when the first predetermined termination condition is satisfied is determined as the first image inpainting model.
[0393] According to an embodiment of the present disclosure, a first loss function can be used to obtain a first loss function value using a first restored sample image, a first predetermined sample image corresponding to the first sample image to be restored, and first quantized sample information corresponding to the first sample image to be restored. Model parameters of the first deep learning model can be adjusted based on the first loss function value until a first predetermined condition is satisfied. For example, the model parameters of the first deep learning model can be adjusted based on a backpropagation algorithm or a stochastic gradient descent algorithm until the first predetermined condition is satisfied. The first deep learning model obtained when the first predetermined condition is satisfied is determined as the trained first image restoration model.
[0394] According to the embodiment of the present disclosure, the specific form of the first loss function can be configured according to actual business needs and is not limited here. For example, the first loss function can be shown as the following formula (2).
[0395] Among them, Loss 1_qua represents the first loss function, C′ represents the number of channels of the first repaired sample image, H′ represents the height of the first repaired sample image, W′ represents the width of the first repaired sample image, img_qua i′,j′ Characterize the first quantized sample information, y′ i′,j′,c′ Characterize the first predetermined sample image, o′ i′,j′,c′ Characterize the first inpainted sample image.
[0396] Figure 15 schematically shows an example diagram of a process of training a first deep learning model using a first restoration sample image and a first predetermined sample image corresponding to a first sample image to be restored to obtain a first image restoration model according to an embodiment of the present disclosure.
[0397] As shown in FIG15 , at step 1500, an image quality assessment operation may be performed on a first sample image to be restored 1501 to obtain first sample image quality assessment information 1502. The first sample image quality assessment information 1502 and the first sample image to be restored 1501 are input into a first deep learning model 1503 to obtain a first restored sample image 1504. First quantized sample information 1505 corresponding to the first sample image to be restored 1501 is determined.
[0398] Based on the first loss function 1507, a first loss function value 1508 is obtained according to the first restored sample image 1501, the first predetermined sample image 1506 corresponding to the first sample image to be restored 1501, and the first quantized sample information 1505 corresponding to the first sample image to be restored 1501. Based on this, the model parameters of the first deep learning model 1503 can be adjusted according to the first loss function value 1508 until a first predetermined termination condition is satisfied. The first deep learning model 1503 obtained when the first predetermined termination condition is satisfied is determined as the first image restoration model.
[0399] According to an embodiment of the present disclosure, determining first quantized sample information corresponding to a first sample image to be restored may include the following operations.
[0400] Performing a first edge extraction operation on the first sample image to be repaired to obtain a first edge sample image. Performing a first quantization operation on the first edge sample image to obtain first quantized sample information corresponding to the first sample image to be repaired.
[0401] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, a first edge extraction operation can be performed on the first sample image to be repaired to obtain a first edge sample image. The first edge extraction operation can be used to capture the edge and structural information of the image. The specific first edge extraction method can be configured according to actual business needs and is not limited here. For example, the first edge extraction method can include at least one of the following: a first edge extraction method based on Canny edge detection, a first edge extraction method based on the Sobel operator, a first edge extraction method based on the Roberts operator, a first edge extraction method based on the Laplacian operator, and a first edge extraction method based on the Prewitt operator.
[0402] According to an embodiment of the present disclosure, after obtaining the first edge sample image, a first quantization operation can be performed on the first edge sample image to obtain first quantized sample information corresponding to the first sample image to be repaired. The first quantization operation can be used to map pixel values in the image to discrete values of lower resolution. The specific first quantization method can be configured according to actual business needs and is not limited here. For example, the first quantization method may include at least one of the following: lossless quantization (i.e., Lossless Quantization), uniform quantization (i.e., Uniform Quantization), vector quantization (i.e., Vector Quantization), color quantization (i.e., Color Quantization), and adaptive quantization (i.e., Adaptive Quantization).
[0403] According to the embodiments of the present disclosure, by performing a first edge extraction operation on a first sample image to be restored, edge information of the image to be restored can be effectively extracted, thereby obtaining a first edge sample image. Furthermore, by performing a first quantization operation on the first edge sample image, first quantized sample information corresponding to the first sample image to be restored is obtained, enabling the quantized first quantized sample information to more accurately reflect the edge features of the first image to be restored.
[0404] According to an embodiment of the present disclosure, the first to-be-repaired sample image includes at least one pixel, and the first edge sample image includes first pixel values corresponding to each of the at least one first pixel.
[0405] According to an embodiment of the present disclosure, performing a first quantization operation on a first edge sample image to obtain first quantized sample information corresponding to a first sample image to be repaired may include the following operations.
[0406] The pixel value interval to which the first pixel value of the pixel belongs is determined, and a quantization value corresponding to the pixel value interval is determined as the first quantization value corresponding to the pixel.
[0407] According to an embodiment of the present disclosure, at least one pixel value interval can be preconfigured. For each pixel value interval in the at least one pixel value interval, a quantization value corresponding to each pixel value interval can be preconfigured. The pixel value interval and the quantization value corresponding to the pixel value interval can be configured according to actual business needs and are not limited here. For example, as shown in Table 1 below, the pixel value interval may include: [0, E min ),[E min , 2*E min ),[2*E min , 3*E min ),[3*E min , 4*E min ) and [4*Emin , E max On this basis, with the pixel value interval [0, E min ) can be 1, which is consistent with the pixel value interval [E min , 2*E min ) The corresponding quantization value can be 2, which is consistent with the pixel value interval [2*E min , 3*E min ) The corresponding quantization value can be 3, which is consistent with the pixel value interval [3*E min , 4*E min ) The corresponding quantization value can be 4, which is consistent with the pixel value interval [4*E min , E max ]The corresponding quantization value can be 5.
[0408] Table 1
[0409] According to an embodiment of the present disclosure, for each of at least one pixel, a first pixel value corresponding to the pixel can be determined. Based on the first pixel value, a pixel value interval to which the first pixel value of the pixel belongs can be determined. Based on this, a quantized value corresponding to the pixel value interval can be determined as the first quantized value corresponding to the pixel.
[0410] According to an embodiment of the present disclosure, since the pixel value interval is determined according to the first pixel value of the pixel, the efficiency of determining the first quantization value can be improved by determining the quantization value corresponding to the pixel value interval as the first quantization value corresponding to the pixel.
[0411] FIG16A schematically shows an example schematic diagram of performing a first quantization operation on a first edge sample image to obtain first quantized sample information corresponding to a first sample image to be repaired according to an embodiment of the present disclosure.
[0412] As shown in FIG16A , in 1600A, an example of linear quantization with 5 quantization levels is schematically shown, wherein the horizontal axis is the pixel value and the vertical axis is the quantization value. min Can represent the minimum level value, E max Can represent the maximum level value, E min =E max / 5.
[0413] The pixel value can be placed in [0, E min ) interval of the pixel corresponding to the quantization value is set to 1, the pixel value is between [E min , 2*E min ) interval of pixels corresponding to the quantization value is set to 2, and the pixel value is between [2*E min , 3*E min ) interval of pixels corresponding to the quantization value is set to 3, and the pixel value is between [3*Emin , 4*E min ) interval of pixels corresponding to the quantization value is set to 4, the pixel value is between [4*E min , E max ] The quantization value corresponding to the pixels in the interval is set to 5,
[0414] FIG16B schematically shows an example schematic diagram of performing a first quantization operation on a first edge sample image to obtain first quantized sample information corresponding to a first sample image to be repaired according to another embodiment of the present disclosure.
[0415] As shown in FIG16B , in 1600B, an example of quantization according to a log curve with 5 quantization levels is schematically shown, wherein the horizontal axis is the pixel value and the vertical axis is the quantization value. max Indicates the maximum level value, the pixel value can be divided into levels according to the power of b, that is, b 5 =E max , then E1=b 1 、E2=b 2 、E3=b 3 、E4=b 4 .
[0416] The quantization value corresponding to the pixel whose pixel value is in the interval [0, E1) can be set to 1, the quantization value corresponding to the pixel whose pixel value is in the interval [E1, E2) can be set to 2, the quantization value corresponding to the pixel whose pixel value is in the interval [E2, E3) can be set to 3, the quantization value corresponding to the pixel whose pixel value is in the interval [E3, E4) can be set to 4, and the quantization value corresponding to the pixel whose pixel value is in the interval [E4, E5) can be set to 5. max ] The quantization value corresponding to the pixels in the interval is set to 5,
[0417] FIG16C schematically shows an example diagram of first quantized sample information according to an embodiment of the present disclosure.
[0418] As shown in FIG16C , first quantized sample information 1601, first quantized sample information 1602, first quantized sample information 1603, and first quantized sample information 1604 schematically illustrate an example of quantization according to four levels, with importance arranged from small to large. The white portions in first quantized sample information 1601, first quantized sample information 1602, first quantized sample information 1603, and first quantized sample information 1604 may be used to represent areas associated with that level, while the black portions may be used to represent other areas.
[0419] The details of the first sample image to be restored corresponding to the first quantized sample information 1601 , the first quantized sample information 1602 , the first quantized sample information 1603 and the first quantized sample information 1604 are concentrated in areas such as windows of the building, building material textures and connections between floors.
[0420] According to an embodiment of the present disclosure, performing a first edge extraction operation on a first sample image to be repaired to obtain a first edge sample image may include the following operations.
[0421] Performing global context processing on the first sample image to be repaired to obtain a global sample feature map, and obtaining a first edge sample image based on the global sample feature map.
[0422] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, global context processing can be performed on the first sample image to be repaired to obtain a global sample feature map. Global context processing can take into account the information and structure of the entire image, and analyze and process the whole. The specific global context processing method can be configured according to actual business needs and is not limited here. For example, the global image features of the first sample image to be repaired can be extracted, and the global image features can include at least one of the following: color histogram, texture features, and shape descriptors. On this basis, the extracted global image features can be used to perform global analysis and modeling, so as to obtain a global sample feature map for characterizing the global description of the first sample image to be repaired.
[0423] According to the embodiments of the present disclosure, since the global sample feature map is obtained by performing global context processing on the first sample image to be inpainted, it can more comprehensively express the information of the first sample image to be inpainted. On this basis, by obtaining the first edge sample image based on the global sample feature map, the accuracy of the first edge sample image can be improved.
[0424] According to an embodiment of the present disclosure, obtaining a first edge sample image according to a global sample feature map may include the following operations.
[0425] Performing local context processing on the first sample image to be repaired to obtain a local sample feature map. Obtaining a first edge sample image based on the global sample feature map and the local sample feature map.
[0426] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, local context processing can be performed on the first sample image to be repaired to obtain a local sample feature map. Local context processing can focus on the neighborhood information around each pixel in the image. The specific local context processing method can be configured according to actual business needs and is not limited here. For example, the image area adjacent to each pixel can be defined for each pixel, and the local image features of the image area can be extracted. The local image features may include at least one of the following: grayscale value, texture feature and gradient information. On this basis, the extracted local image features can be used to perform analysis and processing operations to obtain a local sample feature map for characterizing the local description of the first sample image to be repaired.
[0427] According to the embodiments of the present disclosure, since the local sample feature map is obtained by performing local context processing on the first sample image to be repaired, the local sample feature map can better represent the detailed information of the first sample image to be repaired. On this basis, by combining the global sample feature map and the local sample feature map, a first edge sample image with better expressiveness can be obtained, which is conducive to improving the repair effect of the image repair model and the quality of image repair.
[0428] According to an embodiment of the present disclosure, performing global context processing on the first sample image to be repaired to obtain a global sample feature map may include the following operations.
[0429] The first sample image to be repaired is divided into a plurality of first sample image block groups to be repaired, wherein the first sample image block groups to be repaired include a first sample image block to be repaired. Based on the first intermediate sample feature vectors and first sample position vectors corresponding to each of the plurality of first sample image block groups to be repaired, a second intermediate sample feature vector corresponding to each of the plurality of first sample image block groups to be repaired is obtained. The plurality of second intermediate sample feature vectors are processed based on a sixth self-attention strategy to obtain a third intermediate sample feature vector corresponding to each of at least one seventh cascade level. A global sample feature map is obtained based on the third intermediate sample feature vector corresponding to each of the at least one seventh cascade level.
[0430] According to an embodiment of the present disclosure, after obtaining a first sample image to be restored, an appropriate window size and shape can be selected to divide the first sample image to be restored into a plurality of first sample image blocks to be restored. Each of the plurality of first sample image blocks to be restored is respectively regarded as a first sample image block group to be restored.
[0431] According to an embodiment of the present disclosure, after obtaining a plurality of first sample image block groups to be restored, feature extraction can be performed on the first sample image block to be restored corresponding to each of the plurality of first sample image block groups to be restored, thereby obtaining a second intermediate sample feature vector corresponding to the first sample image block group to be restored. A first sample position vector corresponding to the first sample image block to be restored is determined based on the position of the first sample image block to be restored in the first sample image to be restored.
[0432] According to the embodiment of the present disclosure, the description of the sixth self-attention strategy can refer to the relevant content of the fifth self-attention strategy mentioned above, and will not be repeated here. After obtaining the second intermediate sample feature vectors corresponding to each of the multiple first sample image block groups to be repaired, the multiple second intermediate sample feature vectors can be processed based on the sixth self-attention strategy to obtain the third intermediate sample feature vectors corresponding to each of the at least one seventh cascade level. On this basis, the third intermediate sample feature vectors corresponding to each of the at least one seventh cascade level can be cascaded to obtain a global sample feature map.
[0433] According to the embodiments of the present disclosure, by dividing the first sample image to be repaired into multiple first sample image block groups and processing each first sample image block group separately, it is beneficial to better utilize computing resources and improve the training efficiency of the deep learning model. On this basis, by processing multiple second intermediate sample feature vectors based on the sixth self-attention strategy to obtain a global sample feature map, it is possible to improve the accuracy of the first repair sample image, thereby improving the repair capability of the image repair model.
[0434] According to an embodiment of the present disclosure, the at least one seventh cascade level includes O, where O is an integer greater than 1.
[0435] According to an embodiment of the present disclosure, obtaining a global sample feature map according to the third intermediate sample feature vectors respectively corresponding to at least one seventh cascade level may include the following operations.
[0436] Based on the third intermediate sample feature vectors corresponding to each of the O seventh cascade levels, a first intermediate sample feature map corresponding to each of the O seventh cascade levels is obtained. When 1 < o ≤ O, based on the o-1th second intermediate sample feature map and the oth first intermediate sample feature map, an oth second intermediate sample feature map is obtained, where the first second intermediate sample feature map is the first intermediate sample feature map corresponding to the first seventh cascade level. Based on the second intermediate sample feature map corresponding to each of the O seventh cascade levels, a global sample feature map is obtained. o is an integer greater than or equal to 1 and less than or equal to O.
[0437] According to an embodiment of the present disclosure, the third intermediate sample feature vectors corresponding to each of the O seventh cascade levels can be processed separately to obtain the first intermediate sample feature maps corresponding to each of the O seventh cascade levels. After obtaining the first intermediate sample feature maps corresponding to each of the O seventh cascade levels, for the oth seventh cascade level among the O seventh cascade levels, when o=1, the first intermediate sample feature map corresponding to the 1st seventh cascade level can be determined as the 1st second intermediate sample feature map. When 1<o≤O, the oth second intermediate sample feature map can be obtained based on the o-1th second intermediate sample feature map and the oth first intermediate sample feature map.
[0438] For example, the 1st second intermediate sample feature map and the 2nd first intermediate sample feature map can be decoded to obtain the 2nd second intermediate sample feature map. The 2nd second intermediate sample feature map and the 3rd first intermediate sample feature map can be decoded to obtain the 3rd second intermediate sample feature map. Similarly, the o-1th second intermediate sample feature map and the oth first intermediate sample feature map can be decoded to obtain the oth second intermediate sample feature map.
[0439] According to an embodiment of the present disclosure, obtaining a global sample feature map based on the second intermediate sample feature maps corresponding to each of the O seventh cascade levels may include the following operations.
[0440] In the case of 1≤o<O, the oth third intermediate sample feature map is obtained based on the o+1th third intermediate sample feature map and the oth second intermediate sample feature map, wherein the oth third intermediate sample feature map is the second intermediate sample feature map corresponding to the oth seventh cascade level. The global sample feature map is obtained based on the second intermediate sample feature map and the third intermediate sample feature map corresponding to each of the o seventh cascade levels.
[0441] According to an embodiment of the present disclosure, when o=0, the second intermediate sample feature map corresponding to the oth seventh cascade level can be determined as the oth third intermediate sample feature map. When 1≤o<O, the oth third intermediate sample feature map can be obtained based on the o+1th third intermediate sample feature map and the oth second intermediate sample feature map.
[0442] For example, the 2nd third intermediate sample feature map and the 1st second intermediate sample feature map can be concatenated to obtain the 1st third intermediate sample feature map. The 3rd second intermediate sample feature map and the 2nd first intermediate sample feature map can be concatenated to obtain the 2nd second intermediate sample feature map. Similarly, the (o+1th) second intermediate sample feature map and the (oth) first intermediate sample feature map can be concatenated to obtain the (oth) second intermediate sample feature map.
[0443] According to an embodiment of the present disclosure, after obtaining the second intermediate sample feature map and the third intermediate sample feature map corresponding to each of the O seventh cascade levels, for each of the O seventh cascade levels, the second intermediate sample feature map and the third intermediate sample feature map corresponding to the seventh cascade level can be fused to obtain a fused sample feature map corresponding to the seventh cascade level. After obtaining the fused sample feature map corresponding to each seventh cascade level, the fused sample feature maps corresponding to each seventh cascade level can be cascaded to obtain a global sample feature map.
[0444] According to an embodiment of the present disclosure, a more accurate first intermediate sample feature map corresponding to each of the O seventh cascade levels can be obtained based on the third intermediate sample feature vectors corresponding to each of the O seventh cascade levels. Since the oth second intermediate sample feature map is obtained based on the o-1th second intermediate sample feature map and the oth first intermediate sample feature map, the second intermediate sample feature map corresponding to a specific level can more accurately describe the sample features of that level. On this basis, by integrating and analyzing the second intermediate sample feature maps corresponding to different levels, a more comprehensive and accurate global sample feature map can be obtained.
[0445] According to an embodiment of the present disclosure, performing local context processing on the first sample image to be repaired to obtain a local sample feature map may include the following operations.
[0446] The first sample image to be repaired is divided into a plurality of second sample image block groups to be repaired, wherein the second sample image block groups to be repaired include a plurality of second sample image blocks to be repaired. Based on the fourth intermediate sample feature vectors and the second sample position vector corresponding to each of the plurality of second sample image block groups to be repaired, a fifth intermediate sample feature vector corresponding to each of the plurality of second sample image block groups to be repaired is obtained. Based on the seventh self-attention strategy, the plurality of fifth intermediate sample feature vectors are processed respectively to obtain at least one sixth intermediate sample feature vector corresponding to each of the at least one eighth cascade level. Based on the at least one sixth intermediate sample feature vector corresponding to each of the at least one eighth cascade level, a local sample feature map is obtained.
[0447] According to an embodiment of the present disclosure, after obtaining a first sample image to be restored, an appropriate window size and shape can be selected to divide the first sample image to be restored into a plurality of second sample image blocks to be restored. Multiple second sample image blocks to be restored can be randomly selected from the plurality of second sample image blocks to be restored, so as to determine a second sample image block group based on the randomly selected plurality of second sample image blocks to be restored.
[0448] According to an embodiment of the present disclosure, after obtaining a plurality of second sample image block groups to be restored, feature extraction can be performed on the plurality of second sample image blocks to be restored corresponding to each of the plurality of second sample image block groups to be restored, thereby obtaining a fourth intermediate sample feature vector corresponding to the second sample image block group to be restored. A fifth sample position vector corresponding to the second sample image block group to be restored is determined based on the positions of the plurality of second sample image blocks to be restored corresponding to the second sample image block group in the first sample image to be restored.
[0449] According to the embodiment of the present disclosure, the description of the seventh self-attention strategy can be found in the above-mentioned relevant content about the fifth self-attention strategy, which will not be repeated here. After obtaining the fifth intermediate sample feature vectors corresponding to each of the multiple second sample image block groups to be repaired, the multiple fifth intermediate sample feature vectors can be processed based on the seventh self-attention strategy to obtain the sixth intermediate sample feature vectors corresponding to each of the at least one eighth cascade level. On this basis, the sixth intermediate sample feature vectors corresponding to each of the at least one eighth cascade level can be cascaded to obtain a local sample feature map.
[0450] According to the embodiments of the present disclosure, by dividing the first sample image to be repaired into multiple second sample image block groups to be repaired and processing each second sample image block group separately, it is beneficial to better utilize computing resources and improve the training efficiency of the deep learning model. On this basis, by processing multiple fifth intermediate sample feature vectors based on the seventh self-attention strategy to obtain a global sample feature map, it is possible to improve the accuracy of the first repair sample image, thereby improving the repair capability of the image repair model.
[0451] According to an embodiment of the present disclosure, the at least one eighth cascade level includes P levels, where P is an integer greater than 1.
[0452] According to an embodiment of the present disclosure, obtaining a local sample feature map according to at least one sixth intermediate sample feature vector corresponding to at least one eighth cascade level may include the following operations.
[0453] Determine the target sixth intermediate sample feature vector corresponding to each of the P eighth cascade levels from at least one sixth intermediate sample feature vector corresponding to each of the P eighth cascade levels. Based on the target sixth intermediate sample feature vector corresponding to each of the P eighth cascade levels, obtain the fourth intermediate sample feature map corresponding to each of the P eighth cascade levels. When 1<p≤P, obtain the pth fifth intermediate sample feature map based on the p-1th fifth intermediate sample feature map and the pth fourth intermediate sample feature map, wherein the 1st fifth intermediate sample feature map is the fourth intermediate sample feature map corresponding to the 1st eighth cascade level. Obtain a local sample feature map based on the fifth intermediate sample feature map corresponding to each of the P eighth cascade levels. p is an integer greater than or equal to 1 and less than or equal to P.
[0454] According to an embodiment of the present disclosure, after obtaining at least one sixth intermediate sample feature vector corresponding to each of at least one eighth cascade levels, for each eighth cascade level in the at least one eighth cascade level, a target sixth intermediate sample feature vector corresponding to the eighth cascade level can be determined from the at least one sixth intermediate sample feature vector corresponding to the eighth cascade level.
[0455] According to an embodiment of the present disclosure, the target sixth intermediate sample feature vectors corresponding to each of the P eighth cascade levels can be processed separately to obtain the fourth intermediate sample feature maps corresponding to each of the P eighth cascade levels. After obtaining the fourth intermediate sample feature maps corresponding to each of the P eighth cascade levels, for the pth eighth cascade level among the P eighth cascade levels, when p=1, the fourth intermediate sample feature map corresponding to the 1st eighth cascade level can be determined as the 1st fifth intermediate sample feature map. When 1<p≤P, the pth fifth intermediate sample feature map can be obtained based on the p-1th fifth intermediate sample feature map and the pth fourth intermediate sample feature map.
[0456] For example, the 1st fifth intermediate sample feature map and the 2nd fourth intermediate sample feature map can be decoded to obtain the 2nd fifth intermediate sample feature map. The 2nd fifth intermediate sample feature map and the 3rd fourth intermediate sample feature map can be decoded to obtain the 3rd fifth intermediate sample feature map. Similarly, the p-1th fifth intermediate sample feature map and the pth fourth intermediate sample feature map can be decoded to obtain the pth fifth intermediate sample feature map.
[0457] According to an embodiment of the present disclosure, obtaining a local sample feature map according to the fifth intermediate sample feature maps corresponding to each of the P eighth cascade levels may include the following operations.
[0458] When 1≤p<P, a pth sixth intermediate sample feature map is obtained based on the p+1th sixth intermediate sample feature map and the pth fifth intermediate sample feature map, where the pth sixth intermediate sample feature map is the fifth intermediate sample feature map corresponding to the pth eighth cascade level. A local sample feature map is obtained based on the fifth intermediate sample feature map and the sixth intermediate sample feature map corresponding to each of the p eighth cascade levels.
[0459] According to an embodiment of the present disclosure, when p=P, the fifth intermediate sample feature map corresponding to the Pth eighth cascade level can be determined as the Pth sixth intermediate sample feature map. When 1≤p<P, the pth+1th sixth intermediate sample feature map and the pth fifth intermediate sample feature map can be used to obtain the pth sixth intermediate sample feature map.
[0460] For example, the 2nd sixth intermediate sample feature map and the 1st fifth intermediate sample feature map can be encoded to obtain the 1st sixth intermediate sample feature map. The 3rd sixth intermediate sample feature map and the 2nd fifth intermediate sample feature map can be encoded to obtain the 2nd sixth intermediate sample feature map. Similarly, the p+1th sixth intermediate sample feature map and the pth fifth intermediate sample feature map can be encoded to obtain the pth sixth intermediate sample feature map.
[0461] According to an embodiment of the present disclosure, after obtaining the fifth intermediate sample feature maps corresponding to the P eighth cascade levels, for each of the P eighth cascade levels, the fifth intermediate sample feature map and the sixth intermediate sample feature map corresponding to the eighth cascade level can be fused to obtain an auxiliary sample feature map corresponding to the eighth cascade level. After obtaining the auxiliary sample feature map corresponding to each eighth cascade level, the auxiliary sample feature map corresponding to each eighth cascade level can be decoded to obtain a local sample feature map.
[0462] According to the embodiments of the present disclosure, since the pth sixth intermediate sample feature map is obtained based on the p+1th sixth intermediate sample feature map and the pth fifth intermediate sample feature map, the fifth intermediate sample feature map of the current layer can be generated by utilizing the sixth intermediate sample feature map of the previous layer, thereby better achieving feature extraction and representation. On this basis, since the local sample feature map is obtained based on the fifth intermediate sample feature map and the sixth intermediate sample feature map corresponding to each of the P eighth cascade levels, by utilizing a combination of feature maps at different levels, the ability to express local information can be enhanced, further improving the performance of the image restoration model.
[0463] According to an embodiment of the present disclosure, obtaining a first edge sample image according to a global sample feature map and a local sample feature map may include the following operations.
[0464] Based on the global sample feature map, a seventh intermediate sample feature map and an eighth intermediate sample feature map are obtained. Based on the local sample feature map and the seventh intermediate sample feature map, a ninth intermediate sample feature map is obtained. Based on the eighth intermediate sample feature map and the ninth intermediate sample feature map, a tenth intermediate sample feature map is obtained. Based on the tenth intermediate sample feature map, a first edge sample image is obtained.
[0465] According to an embodiment of the present disclosure, after obtaining a global sample feature map, a fifth convolution operation may be performed on the global sample feature map to obtain a seventh intermediate sample feature map. Furthermore, after obtaining the seventh intermediate sample feature map and the local sample feature map, a fusion operation may be performed on the seventh intermediate sample feature map and the local sample feature map to obtain a ninth intermediate sample feature map.
[0466] According to an embodiment of the present disclosure, a sixth convolution operation may be performed on the global sample feature map to obtain an eighth intermediate sample feature map. On this basis, after obtaining the eighth intermediate sample feature map and the ninth intermediate sample feature map, a fusion operation may be performed on the eighth intermediate sample feature map and the ninth intermediate sample feature map to obtain a tenth intermediate sample feature map. After obtaining the tenth intermediate sample feature map, a seventh convolution operation may be performed on the tenth intermediate sample feature map to obtain a first edge sample image.
[0467] According to the embodiments of the present disclosure, by fusing the local sample feature map with the seventh intermediate sample feature map, a more representative ninth intermediate sample feature map can be obtained. By fusing the eighth intermediate sample feature map with the ninth intermediate sample feature map, the expressive power of the sample feature map can be further improved. Through multi-level feature extraction, the input global sample and local sample can be converted into a more representative and distinguishing tenth intermediate sample feature map, thereby facilitating improvement in the accuracy of the first edge sample image.
[0468] According to an embodiment of the present disclosure, performing a first edge extraction operation on a first sample image to be repaired to obtain a first edge sample image may include the following operations.
[0469] A downsampling operation of Q ninth cascade levels is performed on the first sample image to be repaired to obtain eleventh intermediate sample feature maps corresponding to each of the Q ninth cascade levels, where Q is an integer greater than 1. An upsampling operation is performed on the Qth eleventh intermediate sample feature map to obtain a Qth twelfth intermediate sample feature map. When 1<q<Q, a qth twelfth intermediate sample feature map is obtained based on the q+1th twelfth intermediate sample feature map and the qth eleventh intermediate sample feature map. When q=1, a first edge sample image is obtained based on the 2nd twelfth intermediate sample feature map and the 1st eleventh intermediate sample feature map. q is an integer greater than or equal to 1 and less than or equal to Q.
[0470] According to an embodiment of the present disclosure, the Q ninth cascade levels include Q first ninth cascade levels and Q second ninth cascade levels. After obtaining the first sample image to be repaired, the first sample image to be repaired can be downsampled using the qth first ninth cascade level among the Q first ninth cascade levels to obtain the qth eleventh intermediate sample feature map corresponding to the qth first ninth cascade level. On this basis, the Qth second ninth cascade levels can be sequentially upsampled on the Qth eleventh intermediate sample feature map to obtain the Qth twelfth intermediate sample feature map.
[0471] According to an embodiment of the present disclosure, after obtaining the 1st to the Qth intermediate sample feature maps, when 1<q<Q, the qth 11th intermediate sample feature map among the Q 11th intermediate sample feature maps may be fused with the q+1th 12th intermediate sample feature map and the qth 11th intermediate sample feature map to obtain the qth 12th intermediate sample feature map. When q=1, the 2nd 12th intermediate sample feature map and the 1st 11th intermediate sample feature map may be fused to obtain a first edge sample image.
[0472] According to an embodiment of the present disclosure, by performing Q ninth-cascade-level downsampling operations on the first sample image to be repaired, a plurality of eleventh intermediate sample feature maps of different resolutions can be extracted from the first sample image to be repaired. By performing an upsampling operation on the Qth eleventh intermediate sample feature map, the eleventh intermediate sample feature map with a lower resolution can be restored to the Qth twelfth intermediate sample feature map with a higher resolution. On this basis, by performing multi-level processing on the q+1th twelfth intermediate sample feature map and the qth eleventh intermediate sample feature map, a first edge sample image with a better edge feature expression capability can be obtained, which is beneficial to improving the repair effect of the subsequent image repair model.
[0473] According to an embodiment of the present disclosure, obtaining the qth twelfth intermediate sample feature map according to the q+1th twelfth intermediate sample feature map and the qth eleventh intermediate sample feature map may include the following operations.
[0474] Based on the q+1th intermediate sample feature map, the qth thirteenth intermediate sample feature map is obtained. Based on the qth eleventh intermediate sample feature map, the qth fourteenth intermediate sample feature map is obtained. Based on the qth thirteenth intermediate sample feature map and the qth fourteenth intermediate sample feature map, the qth fifteenth intermediate sample feature map is obtained. Based on the qth fifteenth intermediate sample feature map, the qth twelfth intermediate sample feature map is obtained.
[0475] According to an embodiment of the present disclosure, after obtaining the q+1th intermediate sample feature map, an eighth convolution operation may be performed on the q+1th intermediate sample feature map to obtain the qth thirteenth intermediate sample feature map. After obtaining the qth eleventh intermediate sample feature map, a ninth convolution operation may be performed on the qth eleventh intermediate sample feature map to obtain the qth fourteenth intermediate sample feature map.
[0476] According to an embodiment of the present disclosure, after obtaining the qth thirteenth intermediate sample feature map and the qth fourteenth intermediate sample feature map, a fusion operation can be performed on the qth thirteenth intermediate sample feature map and the qth fourteenth intermediate sample feature map to obtain the qth fifteenth intermediate sample feature map. Based on this, a tenth convolution operation can be performed on the qth fifteenth intermediate sample feature map to obtain the qth twelfth intermediate sample feature map.
[0477] According to an embodiment of the present disclosure, obtaining a first edge sample image according to the second twelfth intermediate sample feature map and the first eleventh intermediate sample feature map may include the following operations.
[0478] Based on the second twelfth intermediate sample feature map, a first sixteenth intermediate sample feature map is obtained. Based on the first eleventh intermediate sample feature map, a first seventeenth intermediate sample feature map is obtained. Based on the first sixteenth intermediate sample feature map and the first seventeenth intermediate sample feature map, a first eighteenth intermediate sample feature map is obtained. Based on the first eighteenth intermediate sample feature map, a first edge sample image is obtained.
[0479] According to an embodiment of the present disclosure, after obtaining the second twelfth intermediate sample feature map, the eleventh convolution operation can be performed on the second twelfth intermediate sample feature map to obtain the first sixteenth intermediate sample feature map. After obtaining the first eleventh intermediate sample feature map, the twelfth convolution operation can be performed on the first eleventh intermediate sample feature map to obtain the first seventeenth intermediate sample feature map.
[0480] According to an embodiment of the present disclosure, after obtaining the first sixteenth intermediate sample feature map and the first seventeenth intermediate sample feature map, a fusion operation can be performed on the first sixteenth intermediate sample feature map and the first seventeenth intermediate sample feature map to obtain the first eighteenth intermediate sample feature map. Based on this, a thirteenth convolution operation can be performed on the first eighteenth intermediate sample feature map to obtain a first edge sample image.
[0481] According to an embodiment of the present disclosure, performing a first edge extraction operation on a first sample image to be repaired to obtain a first edge sample image may include the following operations.
[0482] The first sample image to be repaired is processed by R tenth cascade levels to obtain nineteenth intermediate sample feature maps corresponding to each of the R tenth cascade levels, where R is an integer greater than 1. When r = R, the R twentieth intermediate sample feature map is obtained based on the R-th nineteenth intermediate sample feature map. When 1 ≤ r < R, the r-th twentieth intermediate sample feature map is obtained based on the r+1th twentieth intermediate sample feature map and the r-th nineteenth intermediate sample feature map. The twenty-first intermediate sample feature map is obtained based on the first sample image to be repaired. The first edge sample image is obtained based on the 1st twentieth intermediate sample feature map and the twenty-first intermediate sample feature map. r is an integer greater than or equal to 1 and less than or equal to R.
[0483] According to an embodiment of the present disclosure, when r=R, after obtaining the Rth nineteenth intermediate sample feature map, a fourteenth convolution operation can be performed on the Rth nineteenth intermediate sample feature map to obtain the Rth twentieth intermediate sample feature map.
[0484] According to an embodiment of the present disclosure, when 1≤r<R, after obtaining the (r+1)th twentieth intermediate sample feature map and the (r)th nineteenth intermediate sample feature map, a fusion operation can be performed on the (r+1)th twentieth intermediate sample feature map and the (r)th nineteenth intermediate sample feature map to obtain the (r)th twentieth intermediate sample feature map.
[0485] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, a fifteenth convolution operation can be performed on the first sample image to be repaired to obtain a twenty-first intermediate sample feature map. Based on this, a fusion operation can be performed on the first twentieth intermediate sample feature map and the twenty-first intermediate sample feature map to obtain a first edge sample image.
[0486] According to an embodiment of the present disclosure, obtaining the rth twentieth intermediate sample feature map according to the (r+1)th twentieth intermediate sample feature map and the rth nineteenth intermediate sample feature map may include the following operations.
[0487] The rth 22nd intermediate sample feature map is obtained based on the r+1th 20th intermediate sample feature map and the rth 19th intermediate sample feature map. The rth 20th intermediate sample feature map is obtained based on the rth 22nd intermediate sample feature map.
[0488] According to an embodiment of the present disclosure, after obtaining the (r+1)th twentieth intermediate sample feature map and the (r)th nineteenth intermediate sample feature map, a fusion operation can be performed on the (r+1)th twentieth intermediate sample feature map and the (r)th nineteenth intermediate sample feature map to obtain the (r)th twenty-second intermediate sample feature map. Based on this, a sixteenth convolution operation can be performed on the (r)th twenty-second intermediate sample feature map to obtain the (r)th twenty-second intermediate sample feature map.
[0489] According to the embodiments of the present disclosure, since the rth twenty-second intermediate sample feature map is obtained based on the r+1th twentieth intermediate sample feature map and the rth nineteenth intermediate sample feature map, the rth twentieth intermediate sample feature map with better expression ability can be obtained based on the rth twenty-second intermediate sample feature map, which is beneficial to improving the restoration effect of the subsequent image restoration model.
[0490] According to an embodiment of the present disclosure, performing a first edge extraction operation on a first sample image to be repaired to obtain a first edge sample image may include the following operations.
[0491] The first sample image to be repaired is processed using an edge detection operator to obtain a first edge sample image.
[0492] According to an embodiment of the present disclosure, the edge detection operator may include at least one of the following: a Sobel operator, a Roberts operator, a Laplacian operator, and a Prewitt operator. After obtaining the first sample image to be repaired, the first sample image to be repaired may be processed using the edge detection operator to obtain a first edge sample image.
[0493] According to an embodiment of the present disclosure, by processing the first sample image to be restored using an edge detection operator, a first edge sample image with more prominent edge information can be obtained.
[0494] According to an embodiment of the present disclosure, determining first quantized sample information corresponding to a first sample image to be restored may include the following operations.
[0495] A first depth estimation operation is performed on the first sample image to be restored to obtain a first depth sample image. A second quantization operation is performed on the first depth sample image to obtain first quantized sample information corresponding to the first sample image to be restored.
[0496] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, a first depth estimation operation can be performed on the first sample image to be repaired to obtain a first depth sample image. The first depth estimation operation can infer the depth or distance information of each pixel in the image by analyzing the content of the image. The specific first depth estimation method can be configured according to actual business needs and is not limited here. For example, the first depth estimation method may include at least one of the following: a method based on stereo matching, a method based on structured light, a time-based method, a neural network-based method, and a method based on optical flow.
[0497] According to an embodiment of the present disclosure, after obtaining the first depth sample image, a second quantization operation can be performed on the first depth sample image to obtain first quantized sample information corresponding to the first sample image to be repaired. For an explanation of the second quantization operation, refer to the above-mentioned description of the first quantization operation and will not be repeated here.
[0498] FIG17A schematically shows an example of a process of determining first quantized sample information corresponding to a first sample image to be restored according to another embodiment of the present disclosure.
[0499] As shown in FIG17A , in step 1700A, a first depth estimation operation may be performed on a first sample image 1701 to be restored, obtaining a first depth sample image 1702. A second quantization operation may be performed on the first depth sample image 1702 to obtain first quantized sample information 1703 corresponding to the first sample image 1701 to be restored.
[0500] According to an embodiment of the present disclosure, the first to-be-repaired sample image includes at least one pixel, and the first depth sample image includes second pixel values corresponding to each of the at least one pixel.
[0501] According to an embodiment of the present disclosure, performing a second quantization operation on the first depth sample image to obtain first quantized sample information corresponding to the first sample image to be restored may include the following operations.
[0502] Determining, based on the second pixel value of the pixel and a predetermined pixel threshold, a depth sample image region to which the pixel belongs in the second depth sample image, wherein the depth sample image region includes a background sample image region and a foreground sample image region. Determining a second quantization value corresponding to the pixel based on the depth sample image region to which the pixel belongs.
[0503] According to an embodiment of the present disclosure, a predetermined pixel threshold may be pre-configured based on the background sample image region and the foreground sample image region. The predetermined pixel threshold may be used to determine whether a pixel in the first sample image to be restored belongs to the background sample image region or the foreground sample image region.
[0504] According to an embodiment of the present disclosure, after obtaining a first sample image to be repaired, a second pixel value corresponding to each of at least one pixel can be determined. After obtaining the second pixel value corresponding to the pixel, the second pixel value can be compared with a predetermined pixel threshold to obtain a pixel value comparison result. Based on the pixel value comparison result, it is determined that the pixel belongs to a depth sample image region of the second depth sample image. Based on this, the second quantization value corresponding to the pixel can be determined based on pre-configured second quantization values corresponding to different depth sample image regions.
[0505] According to an embodiment of the present disclosure, since the depth sample image area is determined based on the second pixel value of the pixel and a predetermined pixel threshold, the second quantization value corresponding to the pixel is determined based on the depth sample image area to which the pixel belongs. The second quantization value can be automatically determined, which is beneficial to improving the repair effect of the subsequent image repair model.
[0506] According to an embodiment of the present disclosure, determining a second quantization value corresponding to a pixel according to the depth sample image region to which the pixel belongs may include the following operations.
[0507] If it is determined that the depth sample image region to which the pixel belongs is a background sample image region, the second quantized value corresponding to the pixel is determined to be a first predetermined value. If it is determined that the depth sample image region to which the pixel belongs is a foreground sample image region, the second quantized value corresponding to the pixel is determined to be a second predetermined value.
[0508] According to an embodiment of the present disclosure, the second quantization value corresponding to the background sample image region of the second depth sample image can be preconfigured as a first predetermined value, and the second quantization value corresponding to the foreground sample image region of the second depth sample image can be preconfigured as a second predetermined value. The first predetermined value and the second predetermined value can be configured according to actual business needs and are not limited here. For example, the first predetermined value can be set to 1, and the second predetermined value can be set to 2, 3, or 4.
[0509] According to an embodiment of the present disclosure, when it is determined that the depth sample image region to which a pixel belongs is a background sample image region, the second quantized value corresponding to the pixel can be determined as the first predetermined value corresponding to the background sample image region. When it is determined that the depth sample image region to which the pixel belongs is a foreground sample image region, the second quantized value corresponding to the pixel can be determined as the second predetermined value corresponding to the foreground sample image region.
[0510] According to an embodiment of the present disclosure, determining whether a pixel belongs to a depth sample image region of the second depth sample image according to the second pixel value of the pixel and a predetermined pixel threshold may include the following operations.
[0511] If the second pixel value of the pixel is greater than or equal to the predetermined pixel threshold, the pixel is determined to belong to the background sample image area of the first depth sample image. If the second pixel value of the pixel is less than the predetermined pixel threshold, the pixel is determined to belong to the foreground sample image area of the second depth sample image.
[0512] According to an embodiment of the present disclosure, the predetermined pixel threshold can be configured according to actual business needs and is not limited here. For example, the predetermined pixel threshold can be set to 200. After obtaining the second pixel value corresponding to the pixel, the second pixel value can be compared with the predetermined pixel threshold to obtain a pixel value comparison result. For example, when the pixel value comparison result represents that the second pixel value of the pixel is greater than or equal to the predetermined pixel threshold, it can be determined that the pixel belongs to the background sample image area of the first depth sample image. Alternatively, when the pixel value comparison result represents that the second pixel value of the pixel is less than the predetermined pixel threshold, it can be determined that the pixel belongs to the foreground sample image area of the second depth sample image.
[0513] According to an embodiment of the present disclosure, performing a first depth estimation operation on a first sample image to be restored to obtain a first depth sample image may include the following operations.
[0514] A coarse-scale depth estimation operation is performed on the first image to be restored to obtain a coarse-scale depth sample image. A fine-scale depth estimation operation is performed on the first sample image to be restored and the coarse-scale depth sample image to obtain a first depth sample image.
[0515] According to embodiments of the present disclosure, a coarse-scale depth estimation operation may refer to an operation that uses a global or semi-global approach to perform depth estimation on the entire image or a specific region of the image to obtain relatively coarse depth information. After obtaining a first image to be inpainted, a coarse-scale depth estimation operation may be performed on the first image to be inpainted to obtain a coarse-scale depth sample image with a lower resolution.
[0516] According to embodiments of the present disclosure, a fine-scale depth estimation operation may refer to an operation that uses a more refined feature representation method and a more complex model to extract depth information from an image, thereby obtaining relatively fine depth information. Based on this, a fine-scale depth estimation operation may be performed on the first sample image to be inpainted and the coarse-scale depth sample image to improve the lower-resolution coarse-scale depth sample image and obtain a more refined first depth sample image.
[0517] According to an embodiment of the present disclosure, by combining the coarse-scale depth estimation operation and the fine-scale depth estimation operation to obtain the first depth sample image, more accurate depth information can be provided, which is conducive to fully reflecting the depth changes of different areas in the first image to be repaired.
[0518] According to an embodiment of the present disclosure, performing a fine-scale depth estimation operation on a first sample image to be restored and a coarse-scale depth sample image to obtain a first depth sample image may include the following operations.
[0519] A twenty-third intermediate sample feature map is obtained based on the first sample image to be repaired. A twenty-fourth intermediate sample feature map is obtained based on the twenty-third intermediate sample feature map and the coarse-scale depth sample image. A first depth sample image is obtained based on the twenty-fourth intermediate sample feature map.
[0520] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, a seventeenth convolution operation may be performed on the first sample image to be repaired to obtain a twenty-third intermediate sample feature map. After obtaining the twenty-third intermediate sample feature map, a fusion operation may be performed on the twenty-third intermediate sample feature map and the coarse-scale depth sample image to obtain a twenty-fourth intermediate sample feature map. After obtaining the twenty-fourth intermediate sample feature map, an eighteenth convolution operation may be performed on the twenty-fourth intermediate sample feature map to obtain the first depth sample image.
[0521] According to embodiments of the present disclosure, a twenty-third intermediate sample feature map related to depth information can be obtained based on the first sample image to be restored. Based on the twenty-third intermediate sample feature map and the coarse-scale depth sample image, the coarse-scale depth sample image and the twenty-third intermediate sample feature map can be jointly processed to obtain a twenty-fourth intermediate sample feature map with richer feature representation. Furthermore, based on the twenty-fourth intermediate sample feature map, a first depth sample image reflecting the depth information of the first sample image to be restored can be obtained.
[0522] According to an embodiment of the present disclosure, performing a first depth estimation operation on a first sample image to be restored to obtain a first depth sample image may include the following operations.
[0523] A twenty-fifth intermediate sample feature map is obtained based on the first sample image to be repaired, and the twenty-fifth intermediate sample feature map is processed based on the fifth attention strategy to obtain a first depth sample image.
[0524] According to an embodiment of the present disclosure, after obtaining the first sample image to be repaired, the first sample image to be repaired can be subjected to a nineteenth convolution operation to obtain a twenty-fifth intermediate sample feature map. After obtaining the twenty-fifth intermediate sample feature map, the twenty-fifth intermediate sample feature map can be processed based on the fifth attention strategy to obtain a first depth sample image. The fifth attention strategy includes one of the following: a fifth channel attention strategy, a fifth spatial attention strategy, a fifth mixed attention strategy, and a fifth self-attention strategy. For the description of the fifth channel attention strategy, the fifth spatial attention strategy, the fifth mixed attention strategy, and the fifth self-attention strategy, please refer to the above-mentioned relevant content on the first channel attention strategy, the first spatial attention strategy, the first mixed attention strategy, and the first self-attention strategy, which will not be repeated here.
[0525] According to an embodiment of the present disclosure, since the first depth sample image is obtained by processing the twenty-fifth intermediate sample feature map based on the fifth attention strategy, by introducing the attention mechanism when processing the feature map, it is possible to better focus on important features.
[0526] According to an embodiment of the present disclosure, obtaining a twenty-fifth intermediate sample feature map based on the first sample image to be repaired may include the following operations.
[0527] When 1<s≤S, the s-th 26th intermediate sample feature map is obtained based on the s-1th 26th intermediate sample feature map, where the 1st 26th intermediate sample feature map is obtained based on the first sample image to be repaired. When 1≤s<S, the s-th 27th intermediate sample feature map is obtained based on the s+1th 27th intermediate sample feature map and the s-th 26th intermediate sample feature map, where the s-th 27th intermediate sample feature map is the s-th 26th intermediate sample feature map. The 25th intermediate sample feature map is obtained based on the 1st 27th intermediate sample feature map. S is an integer greater than 1, and s is an integer greater than or equal to 1 and less than or equal to S.
[0528] According to an embodiment of the present disclosure, when s=1, the first sample image to be repaired can be coded to obtain the 1st 26th intermediate sample feature map. When 1<s≤S, the s-1th 26th intermediate sample feature map can be coded to obtain the sth 26th intermediate sample feature map.
[0529] According to an embodiment of the present disclosure, when 1≤s<S, a decoding operation can be performed on the s+1th 27th intermediate sample feature map and the sth 26th intermediate sample feature map to obtain the sth 27th intermediate sample feature map. When s=S, the sth 26th intermediate sample feature map can be determined as the sth 27th intermediate sample feature map. Based on this, a convolution operation can be performed on the 1st 27th intermediate sample feature map to obtain the 25th intermediate sample feature map.
[0530] According to an embodiment of the present disclosure, performing a first depth estimation operation on a first sample image to be restored to obtain a first depth sample image may include the following operations.
[0531] Perform T stages of the fourth feature extraction operation on the first image to be repaired to obtain at least one twenty-eighth intermediate sample feature map corresponding to the Tth stage. Obtain a first depth sample image based on the at least one twenty-eighth intermediate sample feature map corresponding to the Tth stage. tThe image resolution of the twenty-eighth intermediate sample feature maps of the same fourth parallel level is the same, and the image resolution of the twenty-eighth intermediate sample feature maps of different fourth parallel levels is different. T is an integer greater than or equal to 1, t is an integer greater than or equal to 1 and less than or equal to T, U t is an integer greater than or equal to 1.
[0532] According to an embodiment of the present disclosure, the T stages may include an input stage, an intermediate stage, and an output stage. The input stage may refer to stage 1. The output stage may refer to stage T. The intermediate stage may refer to stages 2 to T-1. The number of parallel levels in each stage may be the same or different. In stages 1 to T-1, the current stage may have at least one more parallel level than the previous stage. The T stage may have the same number of parallel levels as the T-1 stage. T may be configured according to actual business needs and is not limited here. For example, T=4. In stages 1 to 3, the current stage may have at least one more parallel level than the previous stage. Stage 1 has U1=2 parallel levels. Stage 2 has U2=3 parallel levels. Stage 3 has U3=4 parallel levels. Stage 4 has U4=4 parallel levels.
[0533] According to an embodiment of the present disclosure, the image resolutions of the twenty-eighth intermediate sample feature maps of the same fourth parallel level are the same. The image resolutions of the twenty-eighth intermediate sample feature maps of different fourth parallel levels are different. For example, the image resolution of the twenty-eighth intermediate sample feature map of the current fourth parallel level is smaller than the image resolution of the twenty-eighth intermediate sample feature map of the previous fourth parallel level. The image resolution of the twenty-eighth intermediate sample feature map of the current fourth parallel level at the current stage may be determined based on the image resolution of the twenty-eighth intermediate sample feature map of the previous fourth parallel level at the previous stage. For example, the image resolution of the twenty-eighth intermediate sample feature map of the current fourth parallel level at the current stage may be obtained by downsampling the image resolution of the twenty-eighth intermediate sample feature map of the previous fourth parallel level at the previous stage.
[0534] According to an embodiment of the present disclosure, when T>1, performing T stages of fourth feature extraction operations on the first image to be repaired to obtain at least one twenty-eighth intermediate sample feature map corresponding to the Tth stage may include: in response to t=1, performing the fourth feature extraction operation on the first image to be repaired to obtain the first twenty-eighth intermediate sample feature map of at least one scale corresponding to the first stage. Obtaining the twenty-eighth intermediate sample feature map of at least one scale corresponding to the first stage based on the first twenty-eighth intermediate sample feature map of at least one scale corresponding to the first stage. In response to 1<t≤T, performing the fourth feature extraction operation on the key point feature map of at least one scale corresponding to the t-1th stage to obtain the first twenty-eighth intermediate sample feature map of at least one scale corresponding to the tth stage. Obtaining the twenty-eighth intermediate sample feature map of at least one scale corresponding to the tth stage based on the first twenty-eighth intermediate sample feature map of at least one scale corresponding to the tth stage.
[0535] According to an embodiment of the present disclosure, when T=1, performing feature extraction for T stages on the first image to be restored to obtain at least one twenty-eighth intermediate sample feature map corresponding to the T stages may include: performing feature extraction on the first image to be restored to obtain a first twenty-eighth intermediate sample feature map of at least one scale corresponding to stage 1. Obtaining a twenty-eighth intermediate sample feature map of at least one scale corresponding to stage 1 based on the first twenty-eighth intermediate sample feature map of at least one scale corresponding to stage 1.
[0536] According to an embodiment of the present disclosure, when T>1, performing T stages of fourth feature extraction operations on the first image to be repaired to obtain at least one twenty-eighth intermediate sample feature map corresponding to the Tth stage may also include: in response to t=1, performing a feature extraction operation based on a dense connection mechanism on the first image to be repaired to obtain a second twenty-eighth intermediate sample feature map of at least one scale corresponding to the first stage. Performing a convolutional attention operation on the second twenty-eighth intermediate sample feature map of at least one scale corresponding to the first stage to obtain a twenty-eighth intermediate sample feature map of at least one scale corresponding to the first stage. In response to 1<t≤T, performing a feature extraction operation based on a dense connection mechanism on the key point feature map of at least one scale corresponding to the t-1th stage to obtain a second twenty-eighth intermediate sample feature map of at least one scale corresponding to the tth stage. Performing a convolutional attention operation on the second twenty-eighth intermediate sample feature map of at least one scale corresponding to the tth stage to obtain a twenty-eighth intermediate sample feature map of at least one scale corresponding to the tth stage.
[0537] According to an embodiment of the present disclosure, since the image resolutions of the twenty-eighth intermediate sample feature maps of the same fourth parallel level are the same, and the image resolutions of the twenty-eighth intermediate sample feature maps of different fourth parallel levels are different, it is possible to maintain high-resolution feature representation throughout the entire feature extraction process, and gradually increase the high resolution to the low-resolution fourth parallel level. Directly extracting deep semantic information from the high-resolution feature representation instead of supplementing the low-level feature information of the image enables it to have sufficient classification capabilities and avoids the loss of effective spatial resolution. At least one fourth parallel level can take into account the capture of contextual information and obtain rich global and local information. In addition, repeatedly exchanging information at the fourth parallel level to achieve multi-scale fusion of features can obtain more accurate depth information, thereby improving the accuracy of the first depth sample image.
[0538] According to an embodiment of the present disclosure, when T is an integer greater than 1, performing T stages of fourth feature extraction operations on the first image to be repaired to obtain at least one twenty-eighth intermediate sample feature map corresponding to the Tth stage may include the following operations.
[0539] When 1<t≤T, a fourth convolution operation is performed on the at least one twenty-eighth intermediate sample feature map corresponding to the t-1th stage to obtain at least one twenty-ninth intermediate sample feature map corresponding to the tth stage. A nineteenth fusion operation is performed on the at least one twenty-ninth intermediate sample feature map corresponding to the tth stage to obtain at least one twenty-eighth intermediate sample feature map corresponding to the tth stage.
[0540] According to an embodiment of the present disclosure, for the t-1 stage, for the 28th intermediate sample feature map in at least one 28th intermediate sample feature map, a fourth convolution operation can be performed on the 28th intermediate sample feature map to obtain the 29th intermediate sample feature map of the t stage, thereby obtaining at least one 29th intermediate sample feature map corresponding to the t stage.
[0541] According to an embodiment of the present disclosure, performing a nineteenth fusion operation on at least one twenty-ninth intermediate sample feature map corresponding to the t-th stage to obtain at least one twenty-eighth intermediate sample feature map corresponding to the t-th stage may include: for the twenty-ninth intermediate sample feature map among the at least one twenty-ninth intermediate sample feature map corresponding to the t-th stage, performing a nineteenth fusion operation on the twenty-ninth intermediate sample feature map of the t-th stage and the twenty-ninth intermediate sample feature maps of other parallel levels other than the parallel level where the twenty-ninth intermediate sample feature map is located in the t-th stage to obtain the twenty-eighth intermediate sample feature map of the t-th stage corresponding to the twenty-ninth intermediate sample feature map. The other parallel levels may refer to at least some of the parallel levels of the t-th stage other than the parallel level where the twenty-ninth intermediate sample feature map is located.
[0542] According to an embodiment of the present disclosure, performing a nineteenth fusion operation on at least one twenty-ninth intermediate sample feature map corresponding to the t-th stage to obtain at least one twenty-eighth intermediate sample feature map corresponding to the t-th stage may include the following operations.
[0543] Targeting U t The vth fourth parallel level in the fourth parallel levels, according to the other 29th intermediate sample feature graph corresponding to the vth fourth parallel level and the 29th intermediate sample feature graph corresponding to the vth fourth parallel level, obtains the 28th intermediate sample feature graph corresponding to the vth fourth parallel level. The other 29th intermediate sample feature graph corresponding to the vth fourth parallel level is the same as U t The twenty-ninth intermediate sample feature graph corresponding to at least some of the fourth parallel levels except the vth fourth parallel level in the fourth parallel levels, v is greater than or equal to 1 and less than or equal to U t An integer.
[0544] According to an embodiment of the present disclosure, the other twenty-ninth intermediate sample feature map corresponding to the vth fourth parallel level is t The twenty-ninth intermediate sample feature map corresponding to at least part of the fourth parallel levels except the vth fourth parallel level in the fourth parallel levels. v can be greater than or equal to 1 and less than or equal to U t An integer.
[0545] According to an embodiment of the present disclosure, in the case of 1<v<V, at least one first other intermediate sample feature map may be upsampled to obtain an upsampled sample feature map corresponding to at least one first other intermediate sample feature map. At least one second other intermediate sample feature map may be downsampled to obtain a downsampled sample feature map corresponding to at least one second other intermediate sample feature map. The first other intermediate sample feature map may refer to U t The second other intermediate sample feature map may refer to U t The image resolution of the upsampled sample feature map is the same as the resolution of the intermediate sample feature map of the vth fourth parallel level. The resolution of the downsampled sample feature map is the same as the resolution of the intermediate sample feature map of the vth fourth parallel level.
[0546] According to an embodiment of the present disclosure, when v=1, upsampling is performed on at least one first other intermediate sample feature map to obtain an upsampled sample feature map corresponding to the at least one first other intermediate sample feature map. The first other intermediate sample feature map may refer to U tThe upsampled sample feature map has an image resolution that is the same as the resolution of the intermediate sample feature map of the first fourth parallel level.
[0547] According to an embodiment of the present disclosure, when v=V, at least one second other intermediate sample feature map is downsampled to obtain a downsampled sample feature map corresponding to at least one second other intermediate sample feature map. The second other intermediate sample feature map may refer to U t The resolution of the downsampled sample feature map is the same as the resolution of the intermediate sample feature map of the v fourth parallel levels.
[0548] According to an embodiment of the present disclosure, based on the up-sampled sample feature map corresponding to at least one first other intermediate sample feature map, the down-sampled sample feature map corresponding to at least one second other intermediate sample feature map, and the intermediate sample feature map of the vth fourth parallel level, a twenty-eighth intermediate sample feature map corresponding to the vth fourth parallel level is obtained. For example, the up-sampled sample feature map corresponding to at least one first other intermediate sample feature map, the down-sampled sample feature map corresponding to at least one second other intermediate sample feature map, and the intermediate sample feature map of the vth fourth parallel level can be fused to obtain the twenty-eighth intermediate sample feature map corresponding to the vth fourth parallel level. The fusion may include at least one of the following: splicing and addition.
[0549] According to an embodiment of the present disclosure, determining first quantized sample information corresponding to a first sample image to be restored may include the following operations.
[0550] A second edge extraction operation is performed on the first sample image to be repaired to obtain a second edge sample image. A second depth estimation operation is performed on the first sample image to be repaired to obtain a second depth sample image. First quantized sample information corresponding to the first sample image to be repaired is obtained based on the second edge sample image and the second depth sample image.
[0551] According to an embodiment of the present disclosure, for the description of the second edge extraction operation and the second depth estimation operation, please refer to the above-mentioned relevant content about the first edge extraction operation and the first depth estimation operation, which will not be repeated here. After obtaining the first sample image to be repaired, a second edge extraction operation can be performed on the first sample image to be repaired to obtain a second edge sample image. A second depth estimation operation is performed on the first sample image to be repaired to obtain a second depth sample image. On this basis, the second edge sample image and the second depth sample image can be combined to determine the first quantized sample information corresponding to the first sample image to be repaired.
[0552] According to embodiments of the present disclosure, a second edge sample image containing edge information can be obtained by performing a second edge extraction operation on the first sample image to be inpainted. A second depth sample image containing depth information can be obtained by performing a second depth estimation operation on the first sample image to be inpainted. Furthermore, by combining the second edge sample image and the second depth sample image, first quantized sample information corresponding to the first sample image to be inpainted can be obtained, thereby improving the efficiency of image inpainting.
[0553] FIG17B schematically shows an example diagram of a process of determining first quantized sample information corresponding to a first sample image to be restored according to another embodiment of the present disclosure.
[0554] As shown in FIG17B , in step 1700B, a second depth estimation operation can be performed on the first sample image to be restored 1704 to obtain a second depth sample image 1705. A second edge extraction operation can be performed on the first sample image to be restored 1704 to obtain a second edge sample image 1706. Based on this, first quantized sample information 1707 corresponding to the first sample image to be restored 1704 can be obtained based on the second edge sample image 1706 and the second depth sample image 1705.
[0555] According to an embodiment of the present disclosure, obtaining first quantized sample information corresponding to the first sample image to be restored according to the second edge sample image and the second depth sample image may include the following operations.
[0556] A third quantization operation is performed on the second edge sample image to obtain first quantized sample sub-information. A fourth quantization operation is performed on the second depth sample image to obtain second quantized sample sub-information. First quantized sample information corresponding to the first sample image to be repaired is obtained based on the first quantized sample sub-information and the second quantized sample sub-information.
[0557] According to the embodiments of the present disclosure, the description of the third and fourth quantization operations can be found in the above-mentioned description of the first quantization operation, and will not be repeated here. After obtaining the second edge sample image, the third quantization operation can be performed on the second edge sample image to obtain the first quantized sample sub-information. For example, the third quantization operation can be a 4-equal quantization operation, based on which the second edge sample image can be quantized into the first quantized sample sub-information, and the first quantized sample sub-information can be 1, 2, 3, or 4.
[0558] According to an embodiment of the present disclosure, after obtaining the second depth sample image, a fourth quantization operation can be performed on the second depth sample image to obtain second quantized sample sub-information. For example, the fourth quantization operation can be a two-equal quantization operation. Based on this, the second quantized sample sub-information corresponding to the background sample image region can be determined to be 1, and the second quantized sample sub-information corresponding to the foreground sample image region can be determined to be 2, 3, or 4. Based on this, the first quantized sample sub-information and the second quantized sample sub-information can be combined to obtain the first quantized sample information corresponding to the first sample image to be restored.
[0559] According to an embodiment of the present disclosure, since the first quantized sample sub-information is obtained by performing a third quantization operation on the second edge sample image, and the second quantized sample sub-information is obtained by performing a fourth quantization operation on the second depth sample image, on this basis, by combining the first quantized sample sub-information and the second quantized sample sub-information, the first quantized sample information that integrates the edge information and the depth information can be obtained while retaining the main information of the first sample image to be repaired.
[0560] According to an embodiment of the present disclosure, the first sample image to be repaired includes at least one pixel, the first quantized sample sub-information includes a first quantized sub-value corresponding to the at least one pixel, and the second quantized sample sub-information includes a second quantized sub-value corresponding to the at least one pixel.
[0561] According to an embodiment of the present disclosure, obtaining first quantized sample information corresponding to a first sample image to be restored based on first quantized sub-information and second quantized sub-information may include the following operations: determining a statistical value based on the first quantized sub-value and the second quantized sub-value corresponding to a pixel; and determining the statistical value as the first quantized value corresponding to the pixel.
[0562] According to an embodiment of the present disclosure, a second edge sample image including edge information obtained by the second edge extraction operation and a second depth sample image including depth of field information obtained by the second depth estimation operation can be used in combination, that is, a statistical value is comprehensively evaluated based on the first quantization sub-value and the second quantization sub-value corresponding to the pixel.
[0563] According to the embodiment of the present disclosure, the method for determining the statistical value can be configured according to actual business needs and is not limited here. For example, the method for determining the statistical value can be as shown in the following formula (3). qua =max(img qua1 ,img qua2 ) (3)
[0564] Among them, img qua Characterization statistics, img qua1 Represents the first quantized sub-value, img qua2Characterizes the second quantization sub-value, max(img qua1 ,img qua2 ) represents taking the maximum value of the first quantizer value and the second quantizer value.
[0565] According to an embodiment of the present disclosure, performing a restoration operation on a first sample image to be restored according to first sample image quality evaluation information to obtain a first restored sample image may include the following operations.
[0566] A sample image quality evaluation vector is obtained according to the first sample image quality evaluation information, and a restoration operation is performed on the first sample image to be restored according to the sample image quality evaluation vector to obtain a first restored sample image.
[0567] According to an embodiment of the present disclosure, after obtaining the first sample image quality assessment information, a third number of sample image quality assessment values can be mapped to a fourth number of first sample image quality assessment values for each sample image quality assessment value corresponding to at least one sample image quality assessment item in the first sample image quality assessment information. Based on this, a sample image quality assessment vector can be determined based on the fourth number of first sample image quality assessment values in the handover document. After obtaining the sample image quality assessment vector, a restoration operation can be performed on the first sample image to be restored based on the sample image quality assessment vector to obtain a first restored sample image.
[0568] According to the embodiment of the present disclosure, the fourth number can be set according to actual business needs and is not limited here. For example, the fourth number can be set to 64. On this basis, the third number of sample image quality evaluation values is w 21 、w 22 、…、w 2n For example, if the third number is n2, then n2 sample image quality assessment values can be mapped to 64 first sample image quality assessment values s1', s2', ..., s 64 '.
[0569] According to an embodiment of the present disclosure, since the sample image quality assessment vector is obtained based on the first sample image quality assessment information, by repairing the first sample image to be repaired based on the sample image quality assessment vector, the quality of the first repaired sample image can be improved and the repair effect can be enhanced.
[0570] FIG18 schematically shows an example schematic diagram of a process of performing a restoration operation on a first sample image to be restored according to first sample image quality assessment information to obtain a first restored sample image according to an embodiment of the present disclosure.
[0571] As shown in Figure 18, in step 1800, an image quality assessment operation may be performed on a first sample image to be restored 1801 to obtain first sample image quality assessment information 1802. Alternatively, the first sample image quality assessment information 1802 of the first sample image to be restored 1801 may also be directly obtained.
[0572] The first deep learning model may include a first backbone network 1803 and a first mapping network 1804. A convolutional layer 1803_1 may be used to process the first sample image to be repaired 1801 to obtain first intermediate sample image features to be repaired. A residual block 1803_21 may be used to process the first intermediate sample image features to obtain second intermediate sample image features to be repaired. A fully connected layer 1804_2 may be used to process the first sample image quality assessment information 1802 to obtain first intermediate first sample image quality assessment information. A split module 1804_1 may be used to process the first intermediate first sample image quality assessment information to obtain second intermediate first sample image quality assessment information.
[0573] On this basis, the channel fusion module 1803_22 can be used to perform channel fusion on the second intermediate sample image feature to be restored and the second intermediate first sample image quality assessment information to obtain the third intermediate sample image feature to be restored.
[0574] The third intermediate sample image to be restored is processed using residual block 1803_31 to obtain the fourth intermediate sample image to be restored. The first sample image quality assessment information 1802 is processed using fully connected layer 1804_4 to obtain third intermediate first sample image quality assessment information. The third intermediate first sample image quality assessment information is processed using split module 1804_3 to obtain fourth intermediate first sample image quality assessment information.
[0575] On this basis, the channel fusion module 1803_32 can be used to perform channel fusion on the fourth intermediate sample image feature to be restored and the fourth intermediate first sample image quality assessment information to obtain the fourth intermediate sample image feature to be restored.
[0576] The convolution layer 1803_4 can be used to process the features of the fourth intermediate sample image to be repaired to obtain a sample image quality assessment vector. The first sample image to be repaired 1801 is repaired according to the sample image quality assessment vector to obtain a first repaired sample image 1805.
[0577] According to an embodiment of the present disclosure, the first deep learning model includes a first generator and a first discriminator.
[0578] According to an embodiment of the present disclosure, performing a restoration operation on a first sample image to be restored according to first sample image quality evaluation information to obtain a first restored sample image may include the following operations.
[0579] The first generator is used to process the first sample image quality assessment information and the sample image to be repaired to obtain a first repaired sample image.
[0580] According to an embodiment of the present disclosure, training a first deep learning model using a first restored sample image and a first predetermined sample image corresponding to a first sample image to be restored to obtain a first image restoration model may include the following operations.
[0581] The first generator and the first discriminator are alternately trained using the first restored sample image and a first predetermined sample image corresponding to the first sample image to be restored until a first predetermined termination condition is satisfied. The first deep learning model obtained when the first predetermined termination condition is satisfied is determined as the first image restoration mode...
Claims
1. An image restoration method, comprising: Performing an image quality assessment operation on the first image to be restored to obtain image quality assessment information, wherein the image quality assessment information includes image quality assessment values corresponding to at least one image quality assessment item, and the image quality assessment value represents a degree of interference of the image quality assessment item on the first image to be restored; and A restoration operation is performed on the first image to be restored according to the image quality assessment information to obtain a first restored image.
2. The method according to claim 1, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment information to obtain a first restored image includes: Obtaining an image quality assessment vector according to the image quality assessment information; and A restoration operation is performed on the first image to be restored according to the image quality assessment vector to obtain the first restored image.
3. The method according to claim 2, wherein: The step of obtaining an image quality assessment vector according to the image quality assessment information comprises: Performing a dimension transformation operation on the image quality assessment information to obtain the image quality assessment vector, wherein the number of dimensions of the image quality assessment vector is determined according to the first image to be restored.
4. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: In the case of 1<e≤E, According to the e-1th first intermediate feature map, an eth second intermediate feature map is obtained, wherein the first fusion feature map is obtained according to the first first intermediate feature map and the image quality assessment vector, and the first first intermediate feature map is obtained according to the first image to be repaired; Performing a first fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain an e-th first intermediate feature map; and Obtaining the first repaired image according to the Eth first intermediate feature map; Here, E is an integer greater than 1, and e is an integer greater than or equal to 1 and less than or equal to E.
5. The method according to claim 4, wherein: The step of performing a first fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map comprises: A first channel fusion operation is performed on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map.
6. The method according to claim 5, wherein: The image quality assessment vector includes image quality assessment dimension information corresponding to each of F dimensions, the number of channels of the e-th second intermediate feature map includes F, the number of channels of the first intermediate feature map includes F, and F is an integer greater than or equal to 1; The step of performing a first channel fusion operation on the e-th second intermediate feature map and the image quality assessment vector to obtain the e-th first intermediate feature map includes: The f-th channel of the e-th second intermediate feature map is multiplied by the f-th image quality assessment dimension information to obtain the f-th channel of the e-th first intermediate feature map, wherein the f-th image quality assessment dimension information represents the image quality assessment dimension information corresponding to the f-th dimension.
7. The method according to any one of claims 4 to 6, wherein: The step of obtaining the e-th second intermediate feature map according to the e-1th first intermediate feature map comprises: Obtaining an e-th third intermediate feature map according to the e-1th first intermediate feature map; and The e-th second intermediate feature map is obtained according to the e-1th first intermediate feature map and the e-th third intermediate feature map.
8. The method according to claim 7, wherein: The step of obtaining an e-th third intermediate feature map according to the e-1th first intermediate feature map comprises: Performing a first conversion operation on the e-1th first intermediate feature map to obtain an eth fourth intermediate feature map; performing a first channel grouping operation on the e-th fourth intermediate feature map to obtain a plurality of e-th fifth intermediate feature maps; Performing a second conversion operation on the plurality of e-th fifth intermediate feature maps to obtain a plurality of e-th sixth intermediate feature maps; and A second fusion operation is performed on the multiple e-th sixth intermediate feature maps to obtain the e-th third intermediate feature map.
9. The method according to claim 7, wherein: The step of obtaining an e-th third intermediate feature map according to the e-1th first intermediate feature map comprises: Performing a second channel grouping operation on the e-1th first intermediate feature map to obtain a plurality of eth seventh intermediate feature maps; Performing a third conversion operation on the plurality of e-th seventh intermediate feature maps to obtain a plurality of e-th eighth intermediate feature maps; and A third fusion operation is performed on the multiple e-th eighth intermediate feature maps to obtain the e-th third intermediate feature map.
10. The method according to any one of claims 7 to 9, wherein: The step of obtaining the e-th second intermediate feature map according to the e-1th first intermediate feature map and the e-th third intermediate feature map comprises: Processing the e-th third intermediate feature map based on a first attention strategy to obtain an e-th ninth intermediate feature map, wherein the first attention strategy includes one of the following: a first channel attention strategy, a first spatial attention strategy, a first mixed attention strategy, and a first self-attention strategy; and The e-th second intermediate feature map is obtained according to the e-1th first intermediate feature map and the e-th ninth intermediate feature map.
11. The method according to any one of claims 4 to 10, wherein: The step of obtaining the first repaired image according to the Eth first intermediate feature map comprises: Processing the Eth first intermediate feature map based on a second attention strategy to obtain an eth tenth intermediate feature map, wherein the second attention strategy includes one of the following: a second channel attention strategy, a second spatial attention strategy, a second mixed attention strategy, and a second self-attention strategy; and The first repaired image is obtained according to the Eth first intermediate feature map and the eth tenth intermediate feature map.
12. The method according to any one of claims 4 to 6, wherein: The step of obtaining the e-th second intermediate feature map according to the e-1th first intermediate feature map comprises: In the case of 2<g≤G, According to the e1-th twenty-ninth intermediate feature map to the eg-1-th eleventh intermediate feature map, an eg-th eleventh intermediate feature map is obtained, wherein the e2-th eleventh intermediate feature map is obtained according to the e1-th eleventh intermediate feature map, and the e1-th eleventh intermediate feature map is the e-1-th first intermediate feature map; and According to the eGth eleventh intermediate feature map, obtaining the eth second intermediate feature map; Here, G is an integer greater than 2, and g is an integer greater than or equal to 1 and less than or equal to G.
13. The method according to any one of claims 4 to 6, wherein: The step of obtaining the e-th second intermediate feature map according to the e-1th first intermediate feature map comprises: Performing a third convolution operation of multiple first parallel levels on the e-1th first intermediate feature map to obtain an e-th twelfth intermediate feature map corresponding to each of the multiple first parallel levels; and The e-th second intermediate feature map is obtained according to multiple e-th twelfth intermediate feature maps.
14. The method according to claim 13, wherein: The step of obtaining the e-th second intermediate feature map according to the plurality of the e-th twelfth intermediate feature maps comprises: Performing a fourth fusion operation on the plurality of the e-th twelfth intermediate feature maps to obtain the e-th thirteenth intermediate feature map; Processing the e-th thirteenth intermediate feature map based on a fourth attention strategy to obtain an e-th fourteenth intermediate feature map, wherein the fourth attention strategy includes one of the following: a fourth channel attention strategy, a fourth spatial attention strategy, a fourth mixed attention strategy, and a fourth self-attention strategy; and The e-th second intermediate feature map is obtained according to the e-th thirteenth intermediate feature map and the e-th fourteenth intermediate feature map.
15. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: Performing a plurality of first cascade level processing on the target first image to be restored and the image quality assessment vector to obtain a fifteenth intermediate feature map corresponding to each of the plurality of first cascade levels, wherein the target first image to be restored is obtained based on the first image to be restored; and Perform multiple second cascade level processing on the fifteenth intermediate feature maps corresponding to each of the multiple first cascade levels to obtain the first repaired image.
16. The method according to claim 15, wherein: The target first image to be restored is obtained according to the first image to be restored, including one of the following: The target first image to be restored is the first image to be restored; The target first image to be restored is obtained according to sixteenth intermediate feature maps corresponding to each of the plurality of second parallel levels, and the sixteenth intermediate feature maps corresponding to each of the plurality of second parallel levels are obtained by performing the plurality of second parallel level processing on the first image to be restored; The target first image to be restored is obtained based on a first intermediate restoration image and a first structural image, the first intermediate restoration image is obtained based on the first image to be restored, and the first structural image is obtained based on the first intermediate restoration image; as well as The target first image to be repaired is obtained based on the first image to be repaired and a second intermediate repair image, and the second intermediate repair image is obtained by performing a structural repair operation on the first image to be repaired.
17. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: Performing J third cascade level processing on the first image to be repaired to obtain seventeenth intermediate feature maps corresponding to the J third cascade levels, respectively, where J is an integer greater than 1; Obtaining an eighteenth intermediate feature map according to the J-th seventeenth intermediate feature map and the image quality assessment vector; and The first repaired image is obtained according to the eighteenth intermediate feature map and the seventeenth intermediate feature maps corresponding to each of the J third cascade levels.
18. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: According to the first image to be restored and the image quality assessment vector, obtain nineteenth intermediate feature maps corresponding to I fourth cascade levels respectively, where I is an integer greater than 1; According to the Lth nineteenth intermediate feature map, a twentieth intermediate feature map is obtained; and The first repaired image is obtained according to the twentieth intermediate feature map and the nineteenth intermediate feature maps corresponding to the I fourth cascade levels respectively.
19. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: In the case of 1<m≤M, Obtaining an mth 22nd intermediate feature map according to the first 21st intermediate feature map to the m-1th 21st intermediate feature map, wherein the first 21st intermediate feature map is obtained according to the first 22nd intermediate feature map and the image quality assessment vector, and the first 22nd intermediate feature map is obtained according to the first image to be restored; Performing a fifth fusion operation on the mth twenty-second intermediate feature map and the image quality assessment vector to obtain an mth twenty-first intermediate feature map; and Obtaining the first repaired image according to the Mth twenty-first intermediate feature map; Here, M is an integer greater than 1, and m is an integer greater than or equal to 1 and less than or equal to M.
20. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: Obtaining at least one sub-image to be restored at a certain scale according to the first image to be restored; and A step-by-step restoration operation is performed on the sub-image to be restored of the at least one scale according to the image quality evaluation vector to obtain the first restored image.
21. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: Perform a first restoration sub-operation on the first image to be restored according to the image quality evaluation vector to obtain fourth intermediate repaired image; Obtaining a second fused image according to the first image to be restored and the fourth intermediate restored image; Performing a second restoration sub-operation on the second fused image according to the image quality assessment vector to obtain a fifth intermediate restoration image; Obtaining a third fused image according to the first image to be restored and the fifth intermediate restored image; and A third restoration sub-operation is performed on the third fused image according to the image quality assessment vector to obtain the first restored image.
22. The method according to claim 2 or 3, wherein: The performing a restoration operation on the first image to be restored according to the image quality assessment vector to obtain the first restored image includes: Obtaining a sixth intermediate restoration image according to the first image to be restored; Obtaining a second structural image according to the sixth intermediate restoration image; Obtaining a seventh fused image according to the sixth intermediate restoration image and the second structural image; and A fourth restoration sub-operation is performed on the seventh fused image according to the image quality assessment vector to obtain the first restored image.
23. The method according to any one of claims 1 to 22, wherein: The performing of an image quality assessment operation on the first image to be repaired to obtain image quality assessment information includes: According to the restoration type corresponding to the first image to be restored, calling an evaluation interface corresponding to the restoration type; and An image quality assessment operation is performed on the first image to be restored using the assessment interface to obtain the image quality assessment information.
24. The method according to any one of claims 1 to 22, wherein: The performing of an image quality assessment operation on the first image to be restored to obtain image quality assessment information includes one of the following: Performing a first feature extraction operation on the first image to be repaired to obtain a shared feature map; obtaining a fifty-fourth intermediate feature map corresponding to each of the at least one image quality assessment items based on the shared feature map; and obtaining an image quality assessment value corresponding to each of the at least one image quality assessment items based on at least one of the fifty-fourth intermediate feature maps. performing a second feature extraction operation on the first image to be repaired to obtain a fifty-fifth intermediate feature map; and obtaining, according to the fifty-fifth intermediate feature map, an image quality assessment value corresponding to each of the at least one image quality assessment items; as well as A third feature extraction operation is performed on the first image to be repaired to obtain a feature that is consistent with the at least one image quality. the fifty-sixth intermediate feature map corresponding to each evaluation item; And according to at least one of the fifty-sixth intermediate feature maps, an image quality assessment value corresponding to each of the at least one image quality assessment items is obtained.
25. The method according to any one of claims 1 to 24, wherein: The first image to be restored includes an image of at least one of the following application scenarios: film and television image restoration, cultural relic image restoration, medical image restoration, traffic image restoration, and astronomical image restoration.
26. A method for training a deep learning model, comprising: Obtaining a first restored sample image according to the sample image to be restored; Determine quantized sample information corresponding to the sample image to be repaired, wherein the sample image to be repaired includes at least one sample image region to be repaired, and the quantized sample information includes quantized values corresponding to each of the at least one sample image region to be repaired, and the quantized values corresponding to the sample image region to be repaired represent the importance of the sample image region to be repaired; as well as The deep learning model is trained using the first repair sample image, a predetermined sample image corresponding to the sample image to be repaired, and quantized sample information corresponding to the sample image to be repaired to obtain a first image repair model.
27. The method according to claim 26, wherein: The determining of the quantized sample information corresponding to the sample image to be repaired includes one of the following: Performing a first edge extraction operation on the sample image to be repaired to obtain a first edge sample image; and performing a first quantization operation on the first edge sample image to obtain quantized sample information corresponding to the sample image to be repaired; Performing a first depth estimation operation on the sample image to be repaired to obtain a first depth sample image; and performing a second quantization operation on the first depth sample image to obtain quantized sample information corresponding to the sample image to be repaired; as well as Performing a second edge extraction operation on the image to be repaired to obtain a second edge sample image; Performing a second depth estimation operation on the image to be repaired to obtain a second depth sample image; And according to the second edge sample image and the second depth sample image, quantized sample information corresponding to the sample image to be repaired is obtained.
28. The method according to claim 26 or 27, wherein: The method of training a deep learning model using the first restoration sample image, a predetermined sample image corresponding to the sample image to be restored, and quantized sample information corresponding to the sample image to be restored to obtain a first image restoration model includes: Based on the loss function, according to the first repair sample image, the predicted Determine the sample image and the quantized sample information corresponding to the sample image to be repaired, and obtain the loss function value; and Adjusting the model parameters of the deep learning model according to the loss function value until a predetermined end condition is met; and The deep learning model obtained when the predetermined end condition is met is determined as the first image restoration model.
29. The method according to claim 26 or 27, wherein: The second deep learning model includes a generator and a discriminator; The step of obtaining a first restored sample image according to the sample image to be restored includes: Processing the sample image to be repaired by using the generator to obtain the first repaired sample image; The method of training a deep learning model using the first restoration sample image, a predetermined sample image corresponding to the sample image to be restored, and quantized sample information corresponding to the sample image to be restored to obtain a first image restoration model includes: Using the first repair sample image, a predetermined sample image corresponding to the sample image to be repaired, and quantized sample information corresponding to the sample image to be repaired, the generator and the discriminator are alternately trained until a predetermined end condition is met; and The deep learning model obtained when the predetermined end condition is met is determined as the first image restoration model.
30. The method according to any one of claims 26 to 29, wherein: The step of obtaining a first restored sample image according to the sample image to be restored comprises: performing an image quality assessment operation on the sample image to be repaired to obtain sample image quality assessment information, wherein the sample image quality assessment information includes a sample image quality assessment value corresponding to each of at least one sample image quality assessment items, and the sample image quality assessment value represents a degree of interference of the sample image quality assessment item on the sample image to be repaired; and A restoration operation is performed on the sample image to be restored according to the sample image quality assessment information to obtain the first restored sample image.
31. The method according to claim 30, wherein: The predetermined sample image includes one of the following: a reference sample image and a real sample image; Wherein, in the case where the predetermined sample image includes the reference sample image, the method further includes: According to the second restored sample image corresponding to the sample image to be restored, the real sample image corresponding to the sample image to be restored and the sample image quality evaluation information, a The corresponding reference sample image, wherein the second repaired sample image is obtained by processing the sample image to be repaired using a second image repair model.
32. An image restoration method, comprising: Acquire a second image to be restored; as well as Inputting the second image to be repaired into the first image repair model to obtain a second repaired image; Wherein, the first image restoration model is trained using the method according to any one of claims 26 to 31.
33. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 32.
34. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 32.
35. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 32.