Training method of image reconstruction model, image reconstruction method and related equipment
By using a fully convolutional image reconstruction model, the problems of high computational resource consumption and real-time requirements in real-time video transmission are solved, achieving efficient and lightweight image reconstruction that is suitable for servers and mobile devices.
Patent Information
- Application Number
- CN202410956635.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2026-01-20
AI Technical Summary
In weak network or mobile communication scenarios, the quality of real-time video transmission is easily affected, resulting in blurred images and increased latency. Existing super-resolution technologies consume a lot of computing resources, are difficult to meet real-time requirements, and lack optimization for transmission scenarios.
An image reconstruction model employing a fully convolutional structure maps low-resolution images to high-resolution images through convolutional processing and pixel rearrangement. The model is trained to learn the image reconstruction process and is suitable for servers and mobile devices.
It achieves efficient and lightweight image reconstruction, meets real-time requirements, is suitable for various scenarios, especially mobile devices, and improves image quality and resolution.
Smart Images

Figure CN121366085A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a training method of an image reconstruction model, an image reconstruction method and related equipment. BACKGROUND
[0002] With the rapid development of digital technology, video calls have become an important demand for people's daily communication, work and entertainment. However, in a weak network or mobile communication scenario, due to factors such as bandwidth limitation and network fluctuation, the transmission of real-time video faces significant challenges. In particular, in an environment with poor network conditions, the quality of real-time video transmission is easily affected, resulting in blurred pictures and increased delay, thereby seriously affecting the user experience.
[0003] Therefore, it is necessary to perform image reconstruction processing on the received video picture to improve the video picture quality. How to efficiently and accurately perform image reconstruction to improve the picture resolution and reduce the delay has become a hot research direction. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a training method of an image reconstruction model, an image reconstruction method and related equipment, which are used to realize super-resolution reconstruction of an image in an efficient, high-quality and lightweight manner, obtain a high-resolution video picture and details, and consume few resources in the image reconstruction process and are not limited by application scenarios.
[0005] To achieve the above purpose, the embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a training method of an image reconstruction model, comprising:
[0007] obtaining a first image and a second image, the first image and the second image being obtained based on the same image, the resolution of the second image being higher than that of the first image;
[0008] performing convolution processing on the first image by using an image reconstruction model to obtain a first intermediate image;
[0009] performing pixel recombination on the first intermediate image to obtain a reconstructed image of the first image;
[0010] training the image reconstruction model based on the reconstructed image and the second image to obtain a target image reconstruction model.
[0011] In a second aspect, the embodiments of the present application provide an image reconstruction method, comprising:
[0012] convolve the to-be-processed image by using the target image reconstruction model to obtain a third intermediate image, wherein the target image reconstruction model is obtained by training the image reconstruction model based on the training method of the image reconstruction model provided in the first aspect;
[0013] perform pixel recombination on the third intermediate image to obtain a reconstructed image of the to-be-processed image.
[0014] In a third aspect, an embodiment of the present application provides a training device of an image reconstruction model, comprising:
[0015] an acquisition module configured to acquire a first image and a second image, wherein the first image and the second image are obtained based on a same image, and the resolution of the second image is higher than that of the first image;
[0016] a convolution processing module configured to convolve the first image by using an image reconstruction model to obtain a first intermediate image;
[0017] a recombination module configured to perform pixel recombination on the first intermediate image to obtain a reconstructed image of the first image;
[0018] a training module configured to train the image reconstruction model based on the reconstructed image and the second image to obtain a target image reconstruction model.
[0019] In a fourth aspect, an embodiment of the present application provides an image reconstruction device, comprising:
[0020] a convolution processing module configured to convolve a to-be-processed image by using a target image reconstruction model to obtain a third intermediate image;
[0021] a recombination module configured to perform pixel recombination on the third intermediate image to obtain a reconstructed image of the to-be-processed image, wherein the target image reconstruction model is obtained by training the image reconstruction model based on the training method of the image reconstruction model provided in the first aspect.
[0022] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising:
[0023] a processor;
[0024] a memory configured to store instructions executable by the processor;
[0025] The processor is configured to execute the instructions to implement the training method of the image reconstruction model provided in the first aspect, or the processor is configured to execute the instructions to implement the image reconstruction method provided in the second aspect.
[0026] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the training method of the image reconstruction model according to the first aspect, or when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the image reconstruction method according to the second aspect.
[0027] In a seventh aspect, an embodiment of the present application provides a computer program product, the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform some or all of the steps of the training method of the image reconstruction model according to the first aspect or the image reconstruction method according to the second aspect.
[0028] The above at least one technical scheme adopted by the embodiments of the present application can achieve the following beneficial effects:
[0029] The low-resolution first image and the high-resolution second image are taken as training data, the first image is processed by the image reconstruction model to obtain a reconstructed image, and the image reconstruction model is trained based on the reconstructed image and the second image, so that the image reconstruction model can sufficiently learn the mapping relationship from the low-resolution image to the high-resolution image, can realize image reconstruction of an image of any resolution, and can output a high-quality image. On this basis, the image reconstruction model adopts a full convolution structure, and realizes the mapping from the low-resolution image to the high-resolution image by performing convolution processing and pixel recombination on the input image. Since the convolution structure is simple and involves fewer parameters, the image reconstruction process based on the convolution structure has the advantages of high efficiency, low computational complexity, low resource consumption, and light weight. Therefore, the target image reconstruction model trained in this way can not only quickly reconstruct a high-resolution image to meet the real-time requirement, but also consumes less resources in the image reconstruction process, is not limited by the application scenario, can run not only on a server but also on a mobile device with limited computing resources, and can be applied to special operating environments with limited performance such as mobile terminals. BRIEF DESCRIPTION OF DRAWINGS
[0030] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, and the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:
[0031] Figure 1 A flowchart of a training method of an image reconstruction model provided by an embodiment of the present application is shown in the figure;
[0032] Figure 2 A schematic diagram of a first image acquisition method provided by an embodiment of the present application is shown in the figure;
[0033] Figure 3 A schematic diagram of a method for obtaining a first feature image according to an embodiment of the present application;
[0034] Figure 4 A schematic diagram of a method for determining a second sub-convolution kernel according to an embodiment of the present application;
[0035] Figure 5 A structural schematic diagram of an image reconstruction model according to an embodiment of the present application;
[0036] Figure 6 A structural schematic diagram of a super-resolution network according to an embodiment of the present application;
[0037] Figure 7 A structural schematic diagram of a convolution layer according to an embodiment of the present application;
[0038] Figure 8 A structural schematic diagram of an image reconstruction device according to an embodiment of the present application;
[0039] Figure 9 A structural schematic diagram of a target image reconstruction model according to an embodiment of the present application;
[0040] Figure 10 A schematic diagram of a method for obtaining a third feature image according to an embodiment of the present application;
[0041] Figure 11 A structural schematic diagram of a training device of an image reconstruction model according to an embodiment of the present application;
[0042] Figure 12 A structural schematic diagram of an image reconstruction device according to an embodiment of the present application;
[0043] Figure 13 A structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in conjunction with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0045] The terms "first", "second", etc. in the specification and claims are used to distinguish similar objects, and are not used to describe a particular sequential or chronological order. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the front and rear associated objects are "or" relationship.
[0046] As described in the background, in a weak network or mobile communication scenario, the video transmission quality is easily affected by the network conditions, resulting in blurred pictures and increased delay, thereby seriously affecting the user experience.
[0047] Super resolution technology (SR) provides an innovative solution to this problem, which can reconstruct high-resolution images from low-resolution images, significantly improving image quality and details. This technology can break the strong dependence between traditional video transmission quality and network conditions, ensuring the quality of video calls and content transmission even in bandwidth-limited or unstable network conditions.
[0048] The applicant found that although the super resolution technology has obvious advantages, there are still many problems and challenges in realizing real-time super resolution:
[0049] (1) The computational resource consumption is too large, which limits the application scenarios of the above super resolution technology. For example, the computing power, memory and battery life of mobile devices are far inferior to servers and personal computers, which limits the possibility of running the above super resolution technology on mobile devices. At the same time, high computing requirements will significantly increase the power consumption of the running device, affecting the endurance time of the running device.
[0050] (2) It is difficult to meet the real-time requirement. Real-time video call scenarios have low tolerance for latency, which means that image reconstruction processing based on super resolution technology needs to be completed in a very short time, which puts higher requirements on image reconstruction efficiency. The algorithm complexity of the super resolution technology in the related art is high, and the processing time is long, which makes it difficult to meet the real-time requirement.
[0051] (3) Operation operation restriction. Most devices, especially mobile devices, do not have good support for the operation operations involved in super resolution technology.
[0052] (4) Lack of transmission scenario optimization. In weak network transmission conditions, video streams are often processed as low bit rate for transmission, and additional image quality loss will occur during the encoding compression process. However, the super resolution technology in the related art lacks compensation for the encoding loss.
[0053] In view of this, the embodiments of the present application comprehensively consider algorithm efficiency, calculation resource optimization, power consumption management and the like, and propose a training method of an image reconstruction model. The training method takes a first image with low resolution and a second image with high resolution as training data, performs image reconstruction processing on the first image through an image reconstruction model to obtain a reconstructed image, and trains the image reconstruction model based on the reconstructed image and the second image, so that the image reconstruction model can sufficiently learn the mapping relationship from the low-resolution image to the high-resolution image, can realize image reconstruction of an image with any resolution, and can output a high-quality image. On this basis, the image reconstruction model adopts a fully convolutional structure, and realizes the mapping from the low-resolution image to the high-resolution image through convolution processing and pixel recombination on the input image. Since the convolution structure is simple and involves fewer parameters, the image reconstruction process based on the convolution structure has the advantages of high efficiency, low calculation complexity, low resource consumption and light weight. Therefore, the target image reconstruction model trained in this way can not only quickly reconstruct a high-resolution image to meet real-time requirements, but also consumes less resources in the image reconstruction process and is not limited by application scenarios. The target image reconstruction model can not only run on a server, but also run on a mobile device with limited calculation resources, and can be applied to special operating environments with limited performance such as mobile terminals.
[0054] The embodiments of the present application also propose an image reconstruction method, which uses the target image reconstruction model trained by the above training method to reconstruct any image to obtain a high-resolution image. Since the target image reconstruction model has a fully convolutional structure, it realizes image reconstruction from a low-resolution image to a high-resolution image through convolution processing and pixel recombination on the input image, and thus has the advantages of high efficiency, low calculation complexity, low resource consumption and light weight. Therefore, the target image reconstruction model can not only meet real-time requirements, but also be applicable to special operating environments with limited performance such as mobile terminals.
[0055] It should be noted that the training method of the image reconstruction model and the image reconstruction method proposed by the embodiments of the present application can be applied to various scenarios with image reconstruction requirements, such as but not limited to satellite and aerial image processing, security monitoring, medical image enhancement, digital film restoration, and specific applications of consumer electronics.
[0056] It should be understood that the training method of the image reconstruction model and the image reconstruction method provided by the embodiments of the present application can be executed by an electronic device. The electronic device referred to herein can include a terminal device, such as a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a smart home appliance, a smart watch, a vehicle-mounted terminal, an aircraft, and the like. Alternatively, the electronic device can also include a server, such as a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0057] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the drawings.
[0058] Referring to Figure 1 A flowchart of a training method of an image reconstruction model is provided for an embodiment of the present application. The method comprises the following steps:
[0059] S102, a first image and a second image are obtained.
[0060] The first image and the second image are obtained based on the same image. The resolution of the second image is higher than that of the first image. Exemplarily, the first image can be a low-resolution image obtained by performing degradation processing on a sample image, and the second image can be a high-resolution image obtained by performing enhancement processing on the sample image. The sample image can be a video frame extracted from a high-definition video. In actual applications, the sample image is an RGB image, and the first image and the second image generated therefrom are also RGB images.
[0061] The degradation processing on the sample image can be implemented by various degradation techniques, which are not limited in the embodiments of the present application. In an embodiment, as shown in Figure 2 The first image is obtained by processing the sample image in the following manner: adding noise data to the sample image to obtain a noise image; performing a replication operation on the noise image to obtain a first noise video; performing compression encoding on the first noise video to obtain a second noise video; and generating the first image based on at least one video frame in the second noise video.
[0062] Exemplarily, adding noise data to the sample image can be implemented in the following manner: performing blur processing on the sample image by using a Gaussian kernel function with random parameters, then performing noise processing on the sample image subjected to the blur processing, the noise intensity being random each time, and then generating an image with a specified size being a multiple (such as 1 / 2, 1 / 3, 1 / 4, etc.) by using bilinear interpolation, as the noise image.
[0063] Further, the frame repetition technique is used to convert the noise image into the first noise video, and then the H264 encoding is used to encode and compress the first noise video to obtain a second noise video with low resolution and low code rate. Finally, the middle frame of the second noise video is extracted as the first image.
[0064] The enhancement processing on the sample image can be implemented by various data enhancement techniques, which are not limited in the embodiments of the present application. In an embodiment, the sample image is sharpened and enhanced by using a non-sharpening masking technique to obtain the second image.
[0065] S104, the first image is convoluted by using the image reconstruction model to obtain a first intermediate image.
[0066] The image reconstruction model has super-resolution image reconstruction capability and can reconstruct a low-resolution image into a high-resolution image. The image reconstruction model can be used to perform convolution processing on the first image, so as to extract effective features from the first image and obtain a first intermediate image, thereby facilitating subsequent high-quality image reconstruction.
[0067] In an embodiment, in order to ensure that the obtained first intermediate image contains effective image features and provides data support for subsequent reconstruction of a high-resolution image, color space conversion can be performed on the first image, and then convolution processing can be performed on different image channels. Specifically, the first intermediate image includes a first channel image and a first feature image, and S104 can include the following steps.
[0068] S141, performing color space conversion on the first image to obtain a converted image in a target color space.
[0069] The target color space can be set according to actual needs, and embodiments of the present application do not limit this. For example, the target color space can be a YCbCr color space. Since the Y channel of the YCbCr color space represents the brightness component of the color, the human eye is more sensitive to the Y channel of the image. Therefore, by performing convolution processing on the first image converted to the target color space, the sensitivity of the human eye to the Y channel can be reduced, and thus the change in image quality will not be perceived by the human eye.
[0070] Color space conversion of the first image can be achieved by various conversion techniques, and embodiments of the present application do not limit this. As an example, the color space conversion of the first image can be performed by using the channel mapping relationship between the color space of the first image and the target color space.
[0071] For example, the color space of the first image is an RGB color space, and the target color space is a YCbCr color space. The first image can be converted from the RGB color space to the YCbCr color space by using the channel mapping relationship shown in the following formulas (1)-(3).
[0072] I Y = 0.299·R + 0.587·G + 0.114·B (1)
[0073] I Cb = -0.1687·R - 0.3313·G + 0.15·B + 128 (2)
[0074] I Cr = 0.5·R - 0.4187·G - 0.0813·B + 128 (3)
[0075] wherein the first image is represented as I RGB= [R, G, B], R, G, B represent R channel image, G channel image and B channel image respectively; the conversion image is represented as I YCbCr = [I Y , I Cb , I Cr ], I Y , I Cb and I Cr represent Y channel image, Cb channel image and Cr channel image respectively.
[0076] As another example, the conversion between different color spaces can be regarded as a special convolution process along the input channel. For this purpose, the corresponding convolution kernel can be configured for the image reconstruction model in advance based on the channel mapping relationship between different color spaces, and then the convolution operation is performed on the input image by using the convolution kernel, so that the color space conversion of the input image can be realized. Specifically, the first image can be convolved by the second convolution layer of the image reconstruction model to obtain the conversion image of the target color space. The second convolution layer is determined based on the channel mapping relationship between the color space of the first image and the target color space.
[0077] For example, taking the color space of the first image as the RGB color space and the target color space as the YCbCr color space, the channel mapping relationship shown in the above formulas (1)-(3) can be described as the convolution process shown in the following formula (4):
[0078] I YCbCr = I RGB ·C1+B1 (4)
[0079] Wherein, C1 represents the parameter of the second convolution layer, B1 represents the bias vector, B1 = [0 12 12 28] T .
[0080] Since the convolution has better support than the matrix operation, the color space conversion of the first image is performed by the second convolution layer configured for the image reconstruction model, which occupies less memory than other methods based on matrix operation color space conversion, and is more efficient.
[0081] S142, based on the channel parameters of the target color space, the conversion image is separated into a first channel image and a second channel image.
[0082] The channel parameters represent the channels constituting the target color space and the value range of each channel.
[0083] S143, based on the first convolution layer of the image reconstruction model, the second channel image is convolved to obtain a first feature image.
[0084] For example, if the target color space is YCbCr, then the converted image is separated into Y-channel and CbCr-channel images. Since the human eye is highly sensitive to the Y-channel, the Y-channel image has better depth and detail. Therefore, the CbCr-channel image can be used as the first channel image, and the Y-channel image as the second channel image. Based on this, convolution processing is performed on the second channel image using the first convolutional layer, which can extract more effective features and help to reconstruct a higher-resolution image subsequently.
[0085] To extract effective features from the first image more efficiently and fully, two parallel branches can be used to perform convolution processing on the second channel image. Specifically, the first convolutional layer includes a first sub-convolutional kernel and a second sub-convolutional kernel, with the size of the second sub-convolutional kernel being larger than the size of the first sub-convolutional kernel. Accordingly, S143 includes: performing convolution processing on the second channel image based on the first sub-convolutional kernel to obtain a first feature sub-image; performing convolution processing on the second channel image based on the second sub-convolutional kernel to obtain a second feature sub-image; and merging the first feature sub-image and the second feature sub-image to obtain a first feature image.
[0086] For example, such as Figure 3 As shown, the first sub-convolutional kernel is an M×N×1×1 convolutional kernel, indicating that the size of the convolutional kernel is 1×1, the number of input channels is M, and the number of output channels is N. The first sub-convolutional kernel is used to map M input channels to N output channels. The second sub-convolutional kernel is an M×N×K×K convolutional kernel, indicating that the size of the convolutional kernel is K×K (K is an odd number greater than 1), the number of input channels is M, and the number of output channels is N. The second sub-convolutional kernel is used to map M input channels to N output channels. The second channel image with M channels is processed by the first sub-convolutional kernel, as shown in the following formula (5), to obtain the first feature sub-image with N channels; at the same time, the second channel image is processed by the second sub-convolutional kernel, as shown in the following formula (6), to obtain the second feature sub-image with N channels; finally, the first feature sub-image and the second feature sub-image are summed, as shown in the following formula (7), to achieve the merging of the two and obtain the first feature image with N channels.
[0087]
[0088] Y = Y 1×1 +Y K×K (7)
[0089] Among them, Y 1×1 Y represents the first feature sub-image. K×K Y represents the second feature sub-image, Y represents the first feature image, and X represents the second channel image. Indicates the first sub-convolution kernel. This indicates the second sub-convolution kernel, * indicates convolution processing, and + indicates summation processing.
[0090] In this embodiment, the second sub-convolutional kernel can be constructed using weight collapse technology, enabling it to have a complex update formula during gradient backpropagation in the model training process, thereby improving the training effect of the image reconstruction model. Specifically, the first convolutional layer also includes a third and a fourth sub-convolutional kernel. The second sub-convolutional kernel is obtained by convolving the third and fourth sub-convolutional kernels. The third and second sub-convolutional kernels have the same size, and the fourth sub-convolutional kernel has the same size as the first sub-convolutional kernel. The number of input channels of the third sub-convolutional kernel is the same as that of the second sub-convolutional kernel, the number of output channels of the third sub-convolutional kernel is the same as that of the fourth sub-convolutional kernel, and the number of output channels of the fourth sub-convolutional kernel is the same as that of the second sub-convolutional kernel.
[0091] For example, such as Figure 4 As shown, if the second sub-convolutional kernel has M input channels, N output channels, and a size of K×K, then the third sub-convolutional kernel has M input channels, F output channels (F>>N), and a size of K×K. The fourth sub-convolutional kernel has F input channels, N output channels, and a size of 1×1. The second sub-convolutional kernel is obtained by convolving the third and fourth sub-convolutional kernels based on the following formula (8).
[0092]
[0093] in, This represents the second sub-convolution kernel. Indicates the third sub-convolution kernel. This indicates the fourth sub-convolution kernel, and * indicates convolution processing.
[0094] This application embodiment illustrates a partial implementation of S104 described above. Of course, it should be understood that S104 can also be implemented in other ways, and this application embodiment does not limit this. For example, instead of performing channel-by-channel processing on the first image, the first image can be converted to a different color space, and then the entire converted image can be convolved based on the first convolutional layer of the image reconstruction model to obtain a first intermediate image, etc.
[0095] In practical applications, S141 to S143 described above can be implemented using different networks in the image reconstruction model. For example, such as... Figure 5 As shown, the image reconstruction model includes a first conversion network and a super-resolution network. S141 to S142 are implemented through the first conversion network, and S143 is implemented through the super-resolution network.
[0096] Specifically, such as Figure 5As shown, the first image is input to the first conversion network, the first image is color space converted by a second convolutional layer in the first conversion network to obtain a converted image in a target color space, and the converted image is channel separated to obtain a first channel image and a second channel image. Further, the second channel image is input to the super-resolution network, and the second channel image is convoluted by a pre-configured first convolutional layer in the super-resolution network to obtain a first feature image.
[0097] More specifically, as shown in Figure 6 The super-resolution network includes a plurality of first convolutional layers and a pixel recombination layer. The second channel image is convoluted by the plurality of first convolutional layers to obtain a first feature image, and the first feature image is processed by the pixel recombination layer to obtain a reconstructed image of the second channel (hereinafter referred to as a second channel reconstructed image). Each first convolutional layer has a corresponding size of first convolutional layer built-in, the size of the first first convolutional layer is equal to the size of the last first convolutional layer (i.e., both are K1xK1), the sizes of the intermediate first convolutional layers are equal (i.e., both are K2xK2, K2 Figure 7 As shown, the first convolutional layer in the KxK convolutional layer can be stacked by a first sub-convolutional kernel 1x1 and a second sub-convolutional kernel KxK, and the second sub-convolutional kernel KxK is realized by the weight collapse technique.
[0098] It is worth noting that, Figure 5 to Figure 7 Only one structure of the image reconstruction model is shown. It should be understood that the image reconstruction model can also be other appropriate structures, which can be specifically set according to actual needs, and the embodiments of the present application do not limit this.
[0099] S106, pixel recombination is performed on the first intermediate image to obtain a reconstructed image of the first image.
[0100] The resolution of the reconstructed image of the first image is higher than that of the first image.
[0101] In the case where the first image is not channel separated, the above S106 includes the following steps: pixel recombination is performed on the first intermediate image to obtain a candidate reconstructed image, and then color space conversion is performed on the candidate reconstructed image to obtain the reconstructed image of the first image.
[0102] In an embodiment, in order to improve the resolution of the reconstructed image, in the case that the first intermediate image comprises a first channel image and a first feature image, the S106 comprises the following steps: performing pixel reorganization on the first channel image to obtain a first channel reconstructed image; performing pixel reorganization on the first feature image to obtain a second channel reconstructed image; performing channel merging on the first channel reconstructed image and the second channel reconstructed image to obtain a candidate reconstructed image of the target color space; and performing color space conversion on the candidate reconstructed image to obtain the reconstructed image of the first image.
[0103] The pixel reorganization on the first intermediate image can be implemented by various pixel reorganization algorithms in the art, and the embodiments of the present application do not limit this. By performing pixel reorganization on the first intermediate image, the specified multiple (such as 2 times, 3 times, 4 times) up-sampling of the first intermediate image can be realized, the image reconstruction of the first image can be realized, and a high-resolution reconstructed image can be obtained. In this way, any low-resolution image input can be mapped to a high-resolution image of a specified multiple.
[0104] The color space conversion on the candidate reconstructed image can be implemented by various conversion techniques, and the embodiments of the present application do not limit this. As an example, the color space conversion on the candidate reconstructed image can be performed by using the channel mapping relationship between the color space of the first image and the target image color.
[0105] For example, the color space of the first image is the RGB color space, and the target color space is the YCbCr color space. The candidate reconstructed image can be converted from the YCbCr color space to the RGB color space by the channel mapping relationship shown in the following formulas (9)-(11) to obtain the reconstructed image of the first image.
[0106] I R = Y + 1.402·(Cr-128) (9)
[0107] I G = Y - 0.34414·(Cb-128) - 0.71414·(Cr-128) (10)
[0108] I b = Y + 1.772·(Cb-128) (11)
[0109] Wherein, the candidate reconstructed image is represented as I YCbCr = [Y, Cb, Cr], Y, Cb, Cr represent the Y channel image, the Cb channel image, and the Cr channel image respectively; the reconstructed image is represented as I RGB = [I R , I G , I B ], I R , IG , I B respectively represent R channel image, G channel image, B channel image.
[0110] As another example, the conversion between different color spaces can be regarded as a special convolution process along the input channel. To this end, a corresponding convolution kernel can be configured for the image reconstruction model in advance based on the channel mapping relationship between different color spaces, and then the input image is convolved using the convolution kernel, so that the color space conversion of the input image can be realized. Specifically, the third convolution layer of the image reconstruction model can be used to convolve the candidate reconstructed image to obtain the reconstructed image of the first image. The third convolution layer is determined based on the channel mapping relationship between the color space of the first image and the target color space.
[0111] For example, taking the color space of the first image as the RGB color space and the target color space as the YCbCr color space, the channel mapping relationship shown in the above formulas (9)-(11) can be described as the convolution process shown in the following formula (12):
[0112] I' RGB =(I' YCbCr -B2)·C2 (12)
[0113] wherein C2 represents the parameters of the third convolution layer, B2 represents the bias vector, B2 = [0-128-128] T .
[0114] Since convolution has better support than matrix operation, the third convolution layer configured for the image reconstruction model is used to convert the color space of the candidate reconstructed image, which occupies less memory than other methods based on matrix operation color space conversion, and is more efficient.
[0115] In practical applications, as shown in Figure 5 , the image reconstruction model can further include a second conversion network. The color space conversion of the candidate reconstructed image can be realized through the second conversion network. Specifically, as shown in Figure 5 , the first channel reconstructed image and the second channel reconstructed image are merged to obtain the candidate reconstructed image, and the candidate reconstructed image is input into the second conversion network. The third convolution layer in the second conversion network is used to convert the color space of the candidate reconstructed image to obtain the reconstructed image of the first image.
[0116] S108, based on the reconstructed image of the first image and the second image, the image reconstruction model is trained to obtain the target image reconstruction model.
[0117] The difference between the reconstructed image of the first image and the second image reflects the effect of the image reconstruction model on the image reconstruction of the first image. Therefore, the image reconstruction model is trained based on the difference, so that the image reconstruction model can continuously reduce the difference in the reconstruction process and fully learn the accurate mapping relationship from the low-resolution image to the high-resolution image, thereby improving the reconstruction effect of the image reconstruction model.
[0118] In an embodiment, the above 108 can include the following steps:
[0119] S181, based on the reconstructed image and the second image, updating the model parameters of the image reconstruction model to obtain a pre-trained image reconstruction model.
[0120] Illustratively, the absolute square error (Mean Absolute Error, MAE) between the reconstructed image and the second image is calculated to obtain the reconstruction loss of the image reconstruction model; then, the back propagation algorithm is used to update the model parameters of the image reconstruction model with the goal of reducing the reconstruction loss; the above S102-S181 is repeated multiple times until the preset training stop condition is met, thereby obtaining the pre-trained image reconstruction model. The training stop condition can be set according to actual needs, for example, the iteration number reaches a preset number threshold, the reconstruction loss of the image reconstruction model converges, etc., which is not limited by the embodiments of the present application.
[0121] For another example, the reconstructed image and the second image are converted to a target color space, and a Y channel image is separated from the converted reconstructed image and a Y channel image is separated from the converted second image, and the model parameters of the image reconstruction model are updated based on the absolute square error between the two Y channel images; the above S102-S181 is repeated multiple times until the preset training stop condition is met, thereby obtaining the pre-trained image reconstruction model.
[0122] S182, determining a target image reconstruction model based on the pre-trained image reconstruction model.
[0123] As an example, in the above S182, the pre-trained image reconstruction model can be used as the target image reconstruction model.
[0124] As another example, in order to adapt the target image reconstruction model to business needs, in the above S182, the first object image and the second object image are obtained; then, the first interest region of the first object image is convoluted by the pre-trained image reconstruction model to obtain a second intermediate image, and the second intermediate image is pixel-recombined to obtain a reconstructed image of the first interest region; the pre-trained image reconstruction model is trained based on the reconstructed image of the first interest region and the second interest region of the second object image to obtain the target image reconstruction model.
[0125] The first object image and the second object image are obtained based on the same object image, and the resolution of the second object image is higher than that of the first object image. For example, the first object image can be a low-resolution image obtained by performing degradation processing on a sample object image, and the second object image can be a high-resolution image obtained by performing enhancement processing on the sample object image. The sample object image can be a video frame containing a sample object extracted from a high-definition video.
[0126] The first region of interest refers to an image region containing the sample object in the first object image, and the second region of interest refers to an image region containing the sample object in the second object image.
[0127] For example, if a target image reconstruction model suitable for a face reconstruction scenario is to be obtained, the sample object is a face, the first object image is a low-resolution face image, and the second object image is a high-resolution face image. Key point detection is performed on the first object image to obtain key point data of the eyes, mouth, nose, and eyebrows, etc. Then, based on the key point data, the first region of interest (ROI) of the first object image is marked, and the coordinate information of the first region of interest is stored. Similarly, by performing key point detection on the second object image, the second region of interest of the second object image is marked, and the coordinate information of the second region of interest is stored.
[0128] Then, based on the coordinate information of the first region of interest, the first region of interest is convoluted using the pre-trained image reconstruction model to obtain a second intermediate image, and the pixels of the second intermediate image are reorganized to obtain a reconstructed image of the first region of interest. In this way, image reconstruction of the first region of interest is realized.
[0129] Further, the absolute square error between the reconstructed image of the first region of interest and the second region of interest of the second object image is calculated to calculate the reconstruction loss of the pre-trained image reconstruction model. Then, the back propagation algorithm is used to update the model parameters of the pre-trained image reconstruction model with the goal of reducing the reconstruction loss. The above steps are repeated multiple times until a preset training stop condition is met, and the target image reconstruction model is obtained.
[0130] In the above manner, the image reconstruction model is first pre-trained using the general first image and the second image to obtain a pre-trained image reconstruction model, and then the pre-trained image reconstruction model is trained using the first object image and the second object image in the business scenario to be applied. The target image reconstruction model obtained can better adapt to the business scenario, and improve the image reconstruction effect of the target image reconstruction model in the business scenario.
[0131] The training method of the image reconstruction model provided by one or more embodiments of the present application takes a low-resolution first image and a high-resolution second image as training data, performs image reconstruction processing on the first image through the image reconstruction model to obtain a reconstructed image, and trains the image reconstruction model based on the reconstructed image and the second image, so that the image reconstruction model can sufficiently learn the mapping relationship from a low-resolution image to a high-resolution image, can realize image reconstruction of an image of any resolution, and can output a high-quality image. On this basis, the image reconstruction model adopts a fully convolutional structure, and realizes the mapping from a low-resolution image to a high-resolution image through convolution processing and pixel recombination on the input image. Since the convolution structure is simple and involves fewer parameters, the image reconstruction process based on the convolution structure has the advantages of high efficiency, low computational complexity, low resource consumption, and lightweight. Therefore, the target image reconstruction model trained in this way can not only quickly reconstruct a high-resolution image to meet the real-time requirement, but also consumes less resources in the image reconstruction process, is not limited by the application scenario, can run not only on a server but also on a mobile device with limited computing resources, and can be applied to special operating environments with limited performance such as mobile terminals.
[0132] The target image reconstruction model trained based on the above training method can be deployed on various devices such as mobile devices and servers. If it is to be deployed on a mobile device, the target image reconstruction model can be converted into a corresponding format, such as onnx, tflite, nunn, etc., to facilitate the deployment of the mobile device.
[0133] Based on the above training method, the target image reconstruction model trained by the embodiments of the present application also proposes an image reconstruction method. Please refer to Figure 8 The flowchart of the image reconstruction method provided by one embodiment of the present application comprises the following steps:
[0134] S802, performing convolution processing on the to-be-processed image through the target image reconstruction model to obtain a third intermediate image.
[0135] The to-be-processed image refers to a low-resolution image to be processed for image reconstruction. Exemplarily, before or during a video call of a mobile device, a target image reconstruction model is loaded and instantiated for instant super-resolution image reconstruction. During the video call of the mobile device, when a low-resolution low-bitrate video stream is received due to network reasons, a VideoFrame object in the video stream is replaced by a Bitmap object as an input of the target image reconstruction model, and after processing by the target image reconstruction model, a high-resolution Bitmap object is obtained. Thus, a processor of the mobile device, such as a graphics processing unit (GPU), can be used to realize accelerated calculation to ensure real-time processing. Finally, the high-resolution Bitmap object is converted back to a VideoFrame object, the original VideoFrame object is replaced, and the high-resolution image is displayed on the screen of the mobile device, thereby realizing real-time super-resolution enhancement of a low-resolution low-bitrate video stream in a mobile terminal call scenario.
[0136] The specific implementation of S802 is similar to that of S104 in the embodiment shown in Figure 1 The specific implementation of S104 in the embodiment shown in
[0137] S804, pixel recombination is performed on the third intermediate image to obtain a reconstructed image of the to-be-processed image.
[0138] The specific implementation of S802 is similar to that of S106 in the embodiment shown in Figure 1 The specific implementation of S106 in the embodiment shown in
[0139] In an embodiment, S804 includes the following steps: S841, color space conversion is performed on the to-be-processed image to obtain a converted image in a target color space; S842, channel separation is performed on the converted image based on channel parameters of the target color space to obtain a third channel image and a fourth channel image; and S843, convolution processing is performed on the third channel image and a first convolution layer of the target image reconstruction model to obtain a third feature image.
[0140] As shown in Figure 9As shown, the target image reconstruction model is input with the to-be-processed image, the to-be-processed image is subjected to convolution processing by a second convolution layer in the first conversion network, color space conversion of the to-be-processed image is realized, a converted image is obtained, and the converted image is subjected to channel separation to obtain a third channel image and a fourth channel image. Further, the third channel image is subjected to pixel recombination to obtain a third channel reconstruction image; the fourth channel image is input into the super-resolution network, and the fourth channel image is subjected to convolution processing by a first convolution layer pre-configured in the super-resolution network to obtain a third feature image, and the third feature image is subjected to pixel recombination to obtain a fourth channel reconstruction image. Further, the third channel reconstruction image and the fourth channel image are subjected to channel merging to obtain a candidate reconstruction image, and the candidate reconstruction image is input into the second conversion network, the candidate reconstruction image is subjected to color space conversion by a third convolution layer in the second conversion network to obtain a reconstruction image of the to-be-processed image.
[0141] More specifically, the structure of the super-resolution network is as shown in Figure 6 The structure of the super-resolution network is as shown in Figure 7 The second sub-convolution kernel adopts a weight collapse technology as shown in
[0142] Optionally, in order to simplify the structure of each convolution layer, the second sub-convolution kernel can be expanded and merged with the first sub-convolution kernel to form an equivalent single-branch convolution structure, which has both the high-quality processing effect of the parallel branch structure and the high-efficiency processing speed of the single-branch structure.
[0143] Specifically, the first convolution layer includes a first sub-convolution kernel and a second sub-convolution kernel, and the size of the second sub-convolution kernel is larger than the size of the first sub-convolution kernel. Correspondingly, the convolution processing of the first convolution layer of the target image reconstruction model on the third channel image to obtain the third feature image includes: performing size expansion processing on the first sub-convolution kernel to obtain an expanded sub-convolution kernel, the size of the expanded sub-convolution kernel being the same as the size of the second sub-convolution kernel; performing merging processing on the second sub-convolution kernel and the expanded sub-convolution kernel to obtain a merged sub-convolution kernel; and performing convolution processing on the third channel image based on the merged sub-convolution kernel to obtain the third feature image.
[0144] For example, as shown in Figure 10As shown, the first sub-convolution kernel is an MxNxlx1 convolution kernel, and the second sub-convolution kernel is an MxNxKxK convolution kernel. First, the first sub-convolution kernel is expanded to the same size as the second sub-convolution kernel, as shown in equation (13) below, to obtain an expanded sub-convolution kernel MxNxKxK. Then, the expanded sub-convolution kernel and the second sub-convolution kernel are added, as shown in equation (14) below, to obtain a merged sub-convolution kernel MxNxKxK. Finally, the third channel image is convolved based on the expanded sub-convolution kernel to obtain a third feature image, which has an input channel number of M, an output channel number of N, and a size of KxK.
[0145]
[0146] wherein, represents the expanded sub-convolution kernel, represents the first sub-convolution kernel, and expand represents an expansion process that keeps the weight of the center position of the 1x1 convolution kernel unchanged and adds 0 value padding around it until the convolution kernel size reaches KxK. This means that, in addition to the center position, the weights of other positions are all 0. represents the merged sub-convolution kernel, represents the second sub-convolution kernel.
[0147] The image reconstruction method provided by one or more embodiments of the present application uses the target image reconstruction model trained by the training method described above to reconstruct any image to obtain a high-resolution image. Since the target image reconstruction model has a full convolution structure, image reconstruction from a low-resolution image to a high-resolution image is achieved by convolving the input image and recombining the pixels, so it has the advantages of high efficiency, low computational complexity, low resource consumption, and lightweight. Not only can it meet the real-time requirements, but it can also be applied to special operating environments with limited performance such as mobile terminals.
[0148] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve desirable results. In addition, the processes depicted in the drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be utilized or can be advantageous.
[0149] Based on the same inventive concept, the present application also provides a training device for an image reconstruction model. Figure 11 is a structural schematic diagram of a training device 1100 for an image reconstruction model according to an embodiment of the present application. Please refer to Figure 11In a software implementation, the training apparatus of the image reconstruction model can include an acquisition module 1110, a convolution processing module 1120, a recombination module 1130, and a training module 1140.
[0150] The acquisition module 1110 is configured to acquire a first image and a second image, the first image and the second image being obtained based on the same image, and the resolution of the second image being higher than that of the first image.
[0151] The convolution processing module 1120 is configured to perform convolution processing on the first image by using an image reconstruction model to obtain a first intermediate image.
[0152] The recombination module 1130 is configured to perform pixel recombination on the first intermediate image to obtain a reconstructed image of the first image.
[0153] The training module 1140 is configured to train the image reconstruction model based on the reconstructed image and the second image to obtain a target image reconstruction model.
[0154] In another embodiment, the first intermediate image includes a first channel image and a first feature image.
[0155] The convolution processing module is configured to:
[0156] perform color space conversion on the first image to obtain a converted image in a target color space;
[0157] perform channel separation on the converted image based on channel parameters of the target color space to obtain a first channel image and a second channel image;
[0158] perform convolution processing on the second channel image based on a first convolution layer of the image reconstruction model to obtain the first feature image.
[0159] In another embodiment, the first convolution layer includes a first sub-convolution kernel and a second sub-convolution kernel, and the size of the second sub-convolution kernel is greater than that of the first sub-convolution kernel.
[0160] When the convolution processing module performs convolution processing on the second channel image based on the first convolution layer of the image reconstruction model to obtain the first feature image, the following steps are performed:
[0161] perform convolution processing on the second channel image based on the first sub-convolution kernel to obtain a first feature sub-image;
[0162] perform convolution processing on the second channel image based on the second sub-convolution kernel to obtain a second feature sub-image;
[0163] merge the first feature sub-image and the second feature sub-image to obtain the first feature image.
[0164] In another embodiment, the first convolutional layer further comprises a third sub-convolutional kernel and a fourth sub-convolutional kernel, the second sub-convolutional kernel is obtained by performing convolutional processing on the third sub-convolutional kernel and the fourth sub-convolutional kernel, the third sub-convolutional kernel has the same size as the second sub-convolutional kernel, and the fourth sub-convolutional kernel has the same size as the first sub-convolutional kernel.
[0165] The third sub-convolutional kernel has the same number of input channels as the second sub-convolutional kernel, the third sub-convolutional kernel has the same number of output channels as the input channels of the fourth sub-convolutional kernel, and the fourth sub-convolutional kernel has the same number of output channels as the output channels of the second sub-convolutional kernel.
[0166] In another embodiment, the convolutional processing module performs the following steps when performing color space conversion on the first image to obtain a converted image in a target color space:
[0167] performing convolutional processing on the first image by a second convolutional layer of the image reconstruction model to obtain the converted image in the target color space;
[0168] The second convolutional layer is determined based on a channel mapping relationship between the color space of the first image and the target color space.
[0169] In another embodiment, the reorganization module is configured to:
[0170] performing pixel reorganization on the first channel image to obtain a first channel reconstructed image;
[0171] performing pixel reorganization on the first feature image to obtain a second channel reconstructed image;
[0172] performing channel merging on the first channel reconstructed image and the second channel reconstructed image to obtain a candidate reconstructed image in the target color space;
[0173] performing color space conversion on the candidate reconstructed image to obtain a reconstructed image of the first image.
[0174] In another embodiment, the training module is configured to:
[0175] updating model parameters of the image reconstruction model based on the reconstructed image and the second image to obtain a pre-trained image reconstruction model;
[0176] determining a target image reconstruction model based on the pre-trained image reconstruction model.
[0177] In another embodiment, the training module, when determining the target image reconstruction model based on the pre-trained image reconstruction model, performs the following steps:
[0178] obtaining a first object image and a second object image, the first object image and the second object image being obtained based on a same object image, the second object image having a higher resolution than the first object image;
[0179] performing convolution processing on a first region of interest of the first object image by the pre-trained image reconstruction model to obtain a second intermediate image, and performing pixel recombination on the second intermediate image to obtain a reconstructed image of the first region of interest;
[0180] training the pre-trained image reconstruction model based on the reconstructed image of the first region of interest and a second region of interest of the second object image to obtain the target image reconstruction model.
[0181] In another embodiment, the first image is obtained by processing a sample image in the following manner:
[0182] adding noise data to the sample image to obtain a noise image;
[0183] performing a copy operation on the noise image to obtain a first noise video;
[0184] performing compression encoding on the first noise video to obtain a second noise video;
[0185] generating the first image based on at least one video frame in the second noise video.
[0186] Obviously, the training device of the image reconstruction model provided by the embodiments of the present application can be used as an execution subject of the training method of the image reconstruction model as shown in Figure 1 for example, in the training method of the image reconstruction model as shown in Figure 1 Step S102 can be performed by the acquisition module 1110 in the training device of the image reconstruction model as shown in Figure 11 Step S104 can be performed by the convolution processing module 1120 in the training device of the image reconstruction model as shown in Figure 11 Step S106 can be performed by the recombination module 1130 in the training device of the image reconstruction model as shown in Figure 11 Step S108 can be performed by the recombination module 1130 in the training device of the image reconstruction model as shown in Figure 11 Step S108 can be performed by the recombination module 1130 in the training device of the image reconstruction model as shown in
[0187] According to another embodiment of the present application, Figure 11The various modules in the training apparatus of the image reconstruction model shown can be combined into one or several other modules respectively or entirely, or some of the modules can be further split into a plurality of modules with smaller functions to constitute, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual application, the function of one module can also be implemented by a plurality of modules, or the functions of a plurality of modules can be implemented by one module. The training apparatus of the image reconstruction model in the embodiments of the present application can also include other modules. In actual application, these modules can also be implemented by other modules, and can be implemented by a plurality of modules in cooperation.
[0188] According to another embodiment of the present application, the training apparatus of the image reconstruction model shown can be constructed by running a computer program (including program codes) capable of performing the steps involved in the corresponding method on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), a Read-Only Memory (ROM), etc. Figure 1 The training apparatus of the image reconstruction model shown can be constructed by running a computer program (including program codes) capable of performing the steps involved in the corresponding method on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), a Read-Only Memory (ROM), etc. Figure 11 The training apparatus of the image reconstruction model shown can be constructed by running a computer program (including program codes) capable of performing the steps involved in the corresponding method on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), a Read-Only Memory (ROM), etc.
[0189] Based on the same inventive concept, the embodiments of the present application also provide an image reconstruction apparatus. Figure 12 FIG. 1 is a structural schematic diagram of an image reconstruction apparatus 1200 according to an embodiment of the present application. As shown in FIG. 1, the image reconstruction apparatus 1200 can include a convolution processing module 1210 and a recombination module 1220. Figure 12 In a software implementation, the image reconstruction apparatus can include a convolution processing module 1210 and a recombination module 1220.
[0190] The convolution processing module 1210 is configured to perform convolution processing on the to-be-processed image by using a target image reconstruction model to obtain a third intermediate image.
[0191] The recombination module 1220 is configured to perform pixel recombination on the third intermediate image to obtain a reconstructed image of the to-be-processed image, wherein the target image reconstruction model is trained by using the image reconstruction model training method provided in the embodiments of the present application.
[0192] In another embodiment, the third intermediate image includes a third channel image and a third feature image.
[0193] The convolution processing module is configured to:
[0194] perform color space conversion on the image to be processed to obtain a converted image in a target color space;
[0195] perform channel separation on the converted image based on channel parameters of the target color space to obtain a third channel image and a fourth channel image;
[0196] perform convolution processing on the third channel image and a first convolution layer of the target image reconstruction model to obtain the third feature image.
[0197] In another embodiment, the first convolution layer includes a first sub-convolution kernel and a second sub-convolution kernel, and the size of the second sub-convolution kernel is greater than the size of the first sub-convolution kernel;
[0198] When the convolution processing module performs convolution processing on the third channel image and the first convolution layer of the target image reconstruction model to obtain the third feature image, the following steps are performed:
[0199] perform size expansion processing on the first sub-convolution kernel to obtain an expanded sub-convolution kernel, and the size of the expanded sub-convolution kernel is the same as the size of the second sub-convolution kernel;
[0200] perform merging processing on the second sub-convolution kernel and the expanded sub-convolution kernel to obtain a merged sub-convolution kernel;
[0201] perform convolution processing on the third channel image based on the merged sub-convolution kernel to obtain the third feature image.
[0202] Obviously, the image reconstruction device provided by the embodiments of the present application can be used as an execution subject of the training method of the image reconstruction model as shown in Figure 8 for example, the image reconstruction method as shown in Figure 8 In the image reconstruction method as shown in Figure 11 the convolution processing module 1210 in the image reconstruction device as shown in Figure 12 the recombination module 1220 in the image reconstruction device as shown in
[0203] According to another embodiment of the present application, Figure 12The modules in the image reconstruction apparatus shown can be combined into one or several other modules respectively or all, or some of the modules can be further split into a plurality of modules with smaller functions to constitute, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual application, the function of one module can also be implemented by a plurality of modules, or the functions of a plurality of modules are implemented by one module. The image reconstruction apparatus in the embodiments of the present application can also include other modules. In actual application, these modules can also be implemented by other modules, and can be implemented by a plurality of modules.
[0204] According to another embodiment of the present application, the image reconstruction apparatus as shown can be constructed by running a computer program (including program codes) capable of performing the steps of the corresponding method as shown on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), a Read-Only Memory (ROM), etc. Figure 1 The computer program (including program codes) capable of performing the steps of the corresponding method as shown can be used to construct the image reconstruction apparatus as shown, and to implement the image reconstruction method of the embodiments of the present application. The computer program can be recorded on a computer readable storage medium such as a computer readable storage medium, and can be copied to an electronic device through the computer readable storage medium, and run in the electronic device. Figure 12 The computer program (including program codes) capable of performing the steps of the corresponding method as shown can be used to construct the image reconstruction apparatus as shown, and to implement the image reconstruction method of the embodiments of the present application. The computer program can be recorded on a computer readable storage medium such as a computer readable storage medium, and can be copied to an electronic device through the computer readable storage medium, and run in the electronic device.
[0205] Figure 13 is a structural schematic diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 13 At the hardware level, the electronic device includes a processor, and optionally further includes an internal bus, a network interface, a memory. The memory can include a memory such as a high-speed Random-Access Memory (RAM), and can also include a non-volatile memory such as at least one disk memory. Of course, the electronic device can also include other hardware required by the business.
[0206] The processor, the network interface and the memory can be connected with each other through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 13 Only one bidirectional arrow is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.
[0207] The memory is used to store a program. Specifically, the program can include program code including computer operation instructions. The memory can include an internal memory and a non-volatile memory, and provide instructions and data for the processor.
[0208] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs, and forms a training device of an image reconstruction model at a logical level. The processor executes the program stored in the memory, and specifically is configured to perform the following operations:
[0209] Obtain a first image and a second image, the first image and the second image are obtained based on the same image, and the resolution of the second image is higher than that of the first image;
[0210] Convolving the first image through an image reconstruction model to obtain a first intermediate image;
[0211] Pixel recombining the first intermediate image to obtain a reconstructed image of the first image;
[0212] Training the image reconstruction model based on the reconstructed image and the second image to obtain a target image reconstruction model.
[0213] Alternatively, the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs, and forms an image reconstruction device at a logical level. The processor executes the program stored in the memory, and specifically is configured to perform the following operations:
[0214] Convolving a to-be-processed image through a target image reconstruction model to obtain a third intermediate image, wherein the target image reconstruction model is obtained by training the image reconstruction model based on the training method provided in the embodiments of the present application;
[0215] Pixel recombining the third intermediate image to obtain a reconstructed image of the to-be-processed image.
[0216] The method performed by the training device of the image reconstruction model disclosed in the embodiments of the present application Figure 1 The method performed by the image reconstruction device disclosed in the embodiments of the present application Figure 8 The method can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip having a processing capability. In the implementation, each step of the method can be completed by integrated logic circuits or instructions in the form of software in the processor. The processor can be a general processor, including a central processing unit (CPU), a network processor (NP), etc. It can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the method.
[0217] The electronic device can also execute the method Figure 1 and realize the functions of the training device of the image reconstruction model in the embodiments of the present application Figure 1 to Figure 7 , or the electronic device can also execute the method Figure 8 and realize the functions of the image reconstruction device in the embodiments of the present application Figure 8 to Figure 10 , and the embodiments of the present application will not be repeated here.
[0218] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic device.
[0219] The embodiment of the present application further provides a computer readable storage medium storing one or more programs, the one or more programs including instructions which, when executed by a portable electronic device including a plurality of application programs, enable the portable electronic device to perform the method of the embodiment of the present application shown in the above, and specifically to perform the following operations: Figure 1 The method of the embodiment of the present application shown in the above, and specifically to perform the following operations:
[0220] obtaining a first image and a second image, the first image and the second image being obtained based on the same image, the resolution of the second image being higher than that of the first image;
[0221] performing convolution processing on the first image by using an image reconstruction model to obtain a first intermediate image;
[0222] performing pixel recombination on the first intermediate image to obtain a reconstructed image of the first image;
[0223] training the image reconstruction model based on the reconstructed image and the second image to obtain a target image reconstruction model.
[0224] The one or more programs include instructions which, when executed by a portable electronic device including a plurality of application programs, enable the portable electronic device to perform the method of the embodiment of the present application shown in the above, and specifically to perform the following operations: Figure 8 The method of the embodiment of the present application shown in the above, and specifically to perform the following operations:
[0225] performing convolution processing on a to-be-processed image by using a target image reconstruction model to obtain a third intermediate image, wherein the target image reconstruction model is obtained by training the image reconstruction model based on the training method of the image reconstruction model provided by the embodiment of the present application;
[0226] performing pixel recombination on the third intermediate image to obtain a reconstructed image of the to-be-processed image.
[0227] The embodiment of the present application further provides a computer program product, which includes a non-transitory computer readable storage medium storing a computer program, the computer program being operable to cause a computer to execute part or all of the steps of the training method of the image reconstruction model or the image reconstruction method provided by the embodiment of the present application.
[0228] In summary, the above only describes the preferred embodiments of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0229] The systems, apparatuses, modules, or units in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0230] Computer readable media includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer readable media does not include transitory media such as modulated data signals and carriers.
[0231] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0232] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant part can be referred to the part of the method embodiment.
Claims
1. A training method for an image reconstruction model, characterized in that, include: A first image and a second image are acquired, wherein the first image and the second image are obtained based on the same image, and the resolution of the second image is higher than that of the first image; The first intermediate image is obtained by convolutional processing of the first image using an image reconstruction model. The first intermediate image is reconstructed by pixel rearrangement to obtain the reconstructed image of the first image; Based on the reconstructed image and the second image, the image reconstruction model is trained to obtain the target image reconstruction model.
2. The method according to claim 1, characterized in that, The first intermediate image includes a first channel image and a first feature image; The step of performing convolution processing on the first image using an image reconstruction model to obtain a first intermediate image includes: The first image is converted to a different color space to obtain a converted image in the target color space. Based on the channel parameters of the target color space, the converted image is separated into channels to obtain a first channel image and a second channel image; Based on the first convolutional layer of the image reconstruction model, the second channel image is convolved to obtain the first feature image.
3. The method according to claim 2, characterized in that, The step of converting the first image to a color space to obtain a converted image in the target color space includes: The first image is convolved by the second convolutional layer of the image reconstruction model to obtain the converted image of the target color space; The second convolutional layer is determined based on the channel mapping relationship between the color space of the first image and the target color space.
4. The method according to claim 2, characterized in that, The step of reconstructing pixels from the first intermediate image to obtain a reconstructed image of the first image includes: The first channel image is reconstructed by pixel recombination to obtain the first channel reconstructed image; The first feature image is reconstructed pixel by pixel to obtain the second channel reconstructed image; The first channel reconstructed image and the second channel reconstructed image are merged to obtain a candidate reconstructed image of the target color space; The candidate reconstructed image is converted to a different color space to obtain the reconstructed image of the first image.
5. The method according to claim 1, characterized in that, The step of training the image reconstruction model based on the reconstructed image and the second image to obtain the target image reconstruction model includes: Based on the reconstructed image and the second image, the model parameters of the image reconstruction model are updated to obtain a pre-trained image reconstruction model; Based on the pre-trained image reconstruction model, the target image reconstruction model is determined.
6. The method according to claim 5, characterized in that, The step of determining the target image reconstruction model based on the pre-trained image reconstruction model includes: A first object image and a second object image are obtained, wherein the first object image and the second object image are obtained based on the same object image, and the resolution of the second object image is higher than that of the first object image; The first region of interest in the first object image is convolved by the pre-trained image reconstruction model to obtain a second intermediate image, and the second intermediate image is pixel recombined to obtain a reconstructed image of the first region of interest. The pre-trained image reconstruction model is trained based on the reconstructed image of the first region of interest and the second region of interest of the second object image to obtain the target image reconstruction model.
7. The method according to any one of claims 1 to 6, characterized in that, The first image is obtained by processing the sample image in the following way: Add noise data to the sample image to obtain a noisy image; The noisy image is copied to obtain a first noisy video; The first noisy video is compressed and encoded to obtain the second noisy video; The first image is generated based on at least one video frame from the second noisy video.
8. An image reconstruction method, characterized in that, include: The target image reconstruction model is used to perform convolution processing on the image to be processed to obtain a third intermediate image, wherein the target image reconstruction model is trained based on the training method of the image reconstruction model according to any one of claims 1 to 7; The third intermediate image is reconstructed pixel by pixel to obtain the reconstructed image of the image to be processed.
9. The method according to claim 8, characterized in that, The third intermediate image includes a third channel image and a third feature image; The step of performing convolution processing on the image to be processed using the target image reconstruction model to obtain a third intermediate image includes: The image to be processed is converted to a different color space to obtain a converted image in the target color space. Based on the channel parameters of the target color space, the converted image is separated into third and fourth channel images. The third feature image is obtained by performing convolution processing on the third channel image and the first convolutional layer of the target image reconstruction model.
10. The method according to claim 9, characterized in that, The first convolutional layer includes a first sub-convolutional kernel and a second sub-convolutional kernel, wherein the size of the second sub-convolutional kernel is larger than the size of the first sub-convolutional kernel; The process of performing convolution processing on the first convolutional layer of the reconstruction model of the third channel image and the target image to obtain the third feature image includes: The first sub-convolutional kernel is enlarged to obtain an enlarged sub-convolutional kernel, the size of which is the same as the size of the second sub-convolutional kernel; The second sub-convolutional kernel and the expanded sub-convolutional kernel are merged to obtain a merged sub-convolutional kernel; The third channel image is convolved based on the merged sub-convolution kernel to obtain the third feature image.
11. A training device for an image reconstruction model, characterized in that, include: An acquisition module is used to acquire a first image and a second image, wherein the first image and the second image are obtained based on the same image, and the resolution of the second image is higher than that of the first image; The convolution processing module is used to perform convolution processing on the first image through an image reconstruction model to obtain a first intermediate image; The reconstruction module is used to reconstruct pixels of the first intermediate image to obtain a reconstructed image of the first image; The training module is used to train the image reconstruction model based on the reconstructed image and the second image to obtain the target image reconstruction model.
12. An image reconstruction apparatus, characterized in that, include: The convolution processing module is used to perform convolution processing on the image to be processed through the target image reconstruction model to obtain a third intermediate image. The reconstruction module is used to perform pixel reconstruction on the third intermediate image to obtain a reconstructed image of the image to be processed, wherein the target image reconstruction model is trained based on the image reconstruction model training method according to any one of claims 1 to 7.
13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the training method of the image reconstruction model as described in any one of claims 1 to 7; or, the processor is configured to execute the instructions to implement the image reconstruction method as described in any one of claims 8 to 10.
14. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the training method of the image reconstruction model as described in any one of claims 1 to 7; or, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the image reconstruction method as described in any one of claims 8 to 10.
15. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of the training method of the image reconstruction model as claimed in any one of claims 1 to 7 or the image reconstruction method as claimed in any one of claims 8 to 10.