Image Processing Method, Apparatus, Device, and Readable Storage Medium

The panoramic camera image is processed through the conditional neural network model, and super-resolution images are generated using multiple regional images and zoom magnification conditions, solving the problem of limited quality improvement of panoramic camera image super-score technology in the prior art, and achieving higher quality super-resolution images.

CN114092324BActive Publication Date: 2025-06-24RICOH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010859085.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-24
Publication Date
2025-06-24
Estimated Expiration
2040-08-24

AI Technical Summary

Technical Problem

In the prior art, the quality improvement of image super-scoring technology for panoramic cameras is limited, especially in terms of quality improvement of super-resolution images.

Method used

By using the conditional neural network model, at least two area images of the first sampled image are input with the corresponding condition, and a super-resolution image is generated. The conditional neural network model is obtained by training using conditions corresponding to the second sampled image at least two area images at different zoom ratios, and the quality of the second sampled image is higher than that of the first sampled image.

Benefits of technology

When the lens focal length range is exceeded, high-quality super-resolution images can be obtained, which improves the quality of super-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092324B_ABST
    Figure CN114092324B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method, apparatus, device, and readable storage medium, which relate to the field of communication technologies and improve the quality of super-resolution images. The method includes: obtaining an image to be processed; using the image to be processed and the conditions corresponding to the image to be processed as inputs to a conditional neural network model to obtain a super-resolution image of the image to be processed; wherein, the conditional neural network model is trained using at least two regional images of a first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, and the quality of the second sampled images is higher than that of the first sampled image. Embodiments of the present invention can improve the quality of super-resolution images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement, and particularly to an image processing method, apparatus, device and readable storage medium. Background Art

[0002] In the prior art, enhancement and super-resolution (hereinafter referred to as super-resolution) technologies for 360° images are proposed. Among them, enhancement means improving the image quality without changing the image size, and super-resolution means improving the image quality when both the length and width of the original image are enlarged to N times the original, for example, super-resolution by 4 times means improving the image quality when both the length and width of the original image are enlarged to 4 times the original.

[0003] The current image super-resolution technology for panoramic cameras is based on cameras with fixed focal lengths (fixed-focus lenses), and this method has limited improvement in the quality of super-resolution images. Summary of the Invention

[0004] Embodiments of the present invention provide an image processing method, apparatus, device and readable storage medium to improve the quality of super-resolution images.

[0005] In a first aspect, an embodiment of the present invention provides an image processing method, including:

[0006] Obtaining an image to be processed;

[0007] Taking the image to be processed and the conditions corresponding to the image to be processed as inputs of a conditional neural network model to obtain a super-resolution image of the image to be processed;

[0008] Wherein, the conditional neural network model is trained using at least two regional images of a first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, wherein the quality of the second sampled image is higher than the quality of the first sampled image.

[0009] Wherein, the method further includes:

[0010] Training the conditional neural network model.

[0011] Wherein, the training of the conditional neural network model includes:

[0012] Extracting at least two regional images from the first sampled image;

[0013] Taking the at least two regional images and the conditions corresponding to the at least two regional images as inputs into a generator of a conditional generative adversarial network to obtain super-resolution images corresponding to the at least two regional images and the conditional neural network model.

[0014] After extracting at least two regional images from the first sampled image, the method further includes:

[0015] Matching the at least two regional images with second sampled images at at least two zoom ratios respectively;

[0016] Inputting the at least two regional images and the conditions corresponding to the at least two regional images into a generator of a conditional generative adversarial network includes:

[0017] Inputting the at least two matched regional images and the conditions corresponding to the at least two regional images into a generator of a conditional generative adversarial network.

[0018] Wherein, training the conditional neural network model further includes:

[0019] Inputting the second sampled images at the at least two zoom ratios, the super-resolution images corresponding to the at least two regional images, and the conditions corresponding to the at least two regional images into a discriminator of the conditional generative adversarial network to obtain a discrimination result.

[0020] Wherein, the condition includes a one-hot vector.

[0021] Wherein, the one-hot vector is determined by any one of the following information:

[0022] The zoom ratio of a zoom lens;

[0023] The depth of field of a zoom lens;

[0024] The focal length of a zoom lens.

[0025] Wherein, the condition includes: VGG features.

[0026] In a second aspect, an embodiment of the present invention provides an image processing apparatus, including:

[0027] A first acquisition module, configured to acquire an image to be processed;

[0028] A second acquisition module, configured to use the image to be processed and the conditions corresponding to the image to be processed as inputs of a conditional neural network model to obtain a super-resolution image of the image to be processed;

[0029] Wherein, the conditional neural network model is trained by using at least two regional images of a first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, wherein the quality of the second sampled images is higher than the quality of the first sampled image.

[0030] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the image processing method as described in the first aspect.

[0031] In a fourth aspect, an embodiment of the present invention further provides a readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in the image processing method as described in the first aspect above are implemented.

[0032] In an embodiment of the present invention, a conditional neural network model can be used to process an image to be processed, so as to obtain a super-resolution image of the image to be processed. Since in the solution of the embodiment of the present invention, a super-resolution image can be obtained in a case where the lens focal length range is exceeded, therefore, the quality of the obtained super-resolution image can be improved by using the solution of the embodiment of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is one of the flowcharts of the image processing method provided by an embodiment of the present invention;

[0034] Figure 2 is another flowchart of the image processing method provided by an embodiment of the present invention;

[0035] Figure 3 is a schematic diagram of an image to be processed provided by an embodiment of the present invention;

[0036] Figure 4 is one of the schematic diagrams of the processing process of the generator of the conditional generative adversarial network provided by an embodiment of the present invention;

[0037] Figure 5 is one of the schematic diagrams of the processing process of the discriminator of the conditional generative adversarial network provided by an embodiment of the present invention;

[0038] Figure 6 is one of the schematic diagrams of the processing process of the inference (testing) stage of the conditional generative adversarial network provided by an embodiment of the present invention;

[0039] Figure 7 is another schematic diagram of the processing process of the generator of the conditional generative adversarial network provided by an embodiment of the present invention;

[0040] Figure 8 is another schematic diagram of the processing process of the discriminator of the conditional generative adversarial network provided by an embodiment of the present invention;

[0041] Figure 9 is another schematic diagram of the processing process of the inference (testing) stage of the conditional generative adversarial network provided by an embodiment of the present invention;

[0042] Figure 10 is one of the schematic diagrams of the image processing apparatus provided by an embodiment of the present invention;

[0043] Figure 11 is the second schematic diagram of the image processing apparatus provided by an embodiment of the present invention;

[0044] Figure 12 is one of the schematic diagrams of the training module provided by an embodiment of the present invention;

[0045] Figure 13 is the second schematic diagram of the training module provided by an embodiment of the present invention;

[0046] Figure 14 is the third schematic diagram of the training module provided by an embodiment of the present invention;

[0047] Figure 15 is the structural diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0048] In the embodiments of the present invention, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0049] In the embodiments of the present application, the term "a plurality of" refers to two or more, and other quantifiers are similar thereto.

[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0051] See Figure 1 , Figure 1 is the flowchart of the image processing method provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:

[0052] Step 101, obtain the image to be processed.

[0053] Among them, the image to be processed can be a whole image or a regional image in the whole image. In practical applications, the image to be processed can be a low-quality image, such as an image taken by a mobile phone, a panoramic image, etc. Among them, the quality of the image can be measured by the clarity of texture details, the sharpness of edges, whether there is color cast, etc. For example, if the texture details of the image are clear or the clarity meets a preset condition (the preset condition can be set according to needs), then the image can be considered a high-quality image; otherwise, it can be considered a low-quality image.

[0054] An image taken by a high-quality shooting device can be called a high-quality image. High-quality shooting devices usually use perspective lenses and have the following characteristics: 1) high resolution; 2) low chromatic aberration; 3) low noise. A typical high-quality shooting device is a single-lens reflex camera.

[0055] Step 102: Use the image to be processed and the conditions corresponding to the image to be processed as the input of the conditional neural network model to obtain the super-resolution image of the image to be processed.

[0056] In the embodiment of the present invention, the super-resolution image refers to an image with improved resolution compared to the image to be processed. That is to say, the resolution of the super-resolution image is higher than that of the image to be processed.

[0057] Among them, the conditional neural network model is trained using at least two regional images of the first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to the second sampled images at at least two zoom ratios, where the quality of the second sampled images is higher than that of the first sampled image.

[0058] For example, two regional images A and B of the first sampled image can respectively correspond to the second sampled image at zoom ratio C and the second sampled image at zoom ratio D, where zoom ratio C and zoom ratio D are different.

[0059] It should be noted that the conditional neural network model in the embodiment of the present invention can be pre-trained or trained during the execution of the embodiment of the present invention. Moreover, during the training process, it is not necessary to train different neural network models for different zoom ratios, but only to train a neural network model conditional on different zoom ratios, which can save more costs.

[0060] In the embodiment of the present invention, the conditional neural network model is trained using a conditional generative adversarial network. Among them, the input of the conditional generative adversarial network includes at least two regional images in the first sampled image and the conditions corresponding to the at least two regional images. Among them, the conditions can also be understood as the labels of the at least two regional images.

[0061] In the embodiment of the present invention, during the process of training the first sampled image, the condition can be represented by a one-hot vector, or can be represented by the VGG (Visual Geometry Group) features of the first sampled image. Similarly, during the process of generating the super-resolution image, the condition can be represented by a one-hot vector, or can be represented by the VGG features of the image to be processed.

[0062] If the condition includes a one-hot vector, then the one-hot vector is determined by any one of the following information:

[0063] The zoom ratio of the zoom lens;

[0064] The depth of field of the zoom lens;

[0065] The focal length of the zoom lens.

[0066] For example, if the one-hot vector is represented by the zoom ratio of the zoom lens, then (1, 0, 0, 0, 0), (0, 1, 0, 0, 0), (0, 0, 1, 0, 0), (0, 0, 0, 1, 0), (0, 0, 0, 0, 1) can represent zoom ratios of 1x, 2x, 3x, 4x, and 5x respectively.

[0067] Since in the solution of the embodiment of the present invention, a super-resolution image can be obtained in a case where the lens focal length range is exceeded, therefore, the quality of the obtained super-resolution image can be improved by using the solution of the embodiment of the present invention.

[0068] See Figure 2 , Figure 2 is a flowchart of the image processing method provided by the embodiment of the present invention. As Figure 2 shown, it includes the following steps:

[0069] Step 201, train a conditional neural network model.

[0070] The conditional neural network model is trained using a conditional generative adversarial network. The conditional generative adversarial network can include a generator and a discriminator. Among them, the generator is used to generate the conditional neural network model, and the discriminator is used to discriminate the conditional neural network model, so as to further improve the accuracy of the conditional neural network model.

[0071] Specifically, in this step, at least two regional images are extracted from the first sampled image. Then, the at least two regional images and the conditions corresponding to the at least two regional images are input into the generator of the conditional generative adversarial network to obtain the super-resolution images corresponding to the at least two regional images and the conditional neural network model. Specifically, for each regional image, after connecting the regional image and the condition corresponding to the regional image, it is input into the generator of the conditional generative adversarial network.

[0072] To make the obtained neural network model more accurate, after extracting at least two regional images from the first sampled image, the method further includes: respectively matching the at least two regional images with the second sampled images at at least two zoom ratios. For example, two regional images A and B of the first sampled image respectively correspond to the second sampled image at zoom ratio C and the second sampled image at zoom ratio D. Then, match regional image A with the second sampled image at zoom ratio C, and match regional image B with the second sampled image at zoom ratio D. Here, the meaning of matching is to make the contents included in the two images the same or to make the degree of sameness of the contents included in the two images meet a certain preset condition. For example, the degree of sameness is 90% or the like. Then, after the matching, the at least two regional images after matching and the conditions corresponding to the at least two regional images are input into the generator of the conditional generative adversarial network to obtain the super-resolution images corresponding to the at least two regional images and the conditional neural network model.

[0073] In addition, on the above basis, the second sampled images at the at least two zoom ratios, the super-resolution images corresponding to the at least two regional images, and the conditions corresponding to the at least two regional images can also be input into the discriminator of the conditional generative adversarial network to obtain a discrimination result. The discrimination result is a scalar between 0 and 1, which is used to represent the authenticity of the super-resolution image.

[0074] Step 202: Obtain the image to be processed.

[0075] Step 203: Take the image to be processed and the conditions corresponding to the image to be processed as the input of the conditional neural network model to obtain the super-resolution image of the image to be processed.

[0076] Among them, the descriptions of step 202 and step 203 can refer to the descriptions of the foregoing steps 101 and 102.

[0077] Hereinafter, the processes of generating the conditional neural network model, discriminating using the discriminator, and reasoning will be described in combination with different conditions.

[0078] Figure 3 Figure 1 shows an original low-quality panoramic image 1 taken by a fisheye camera, a partially cropped area 2, and a real image 3(x) taken by a single-lens reflex camera. The dashed area 11 in 1 is the area that needs to be enhanced and super-resolved. Figure 3In it, three regional images are extracted. That is, there are 3 different regions in the dotted area 11, which respectively correspond to the 3 images with different focal lengths taken by a single-lens reflex camera in [3], namely 55 mm, 135 mm, and 250 mm focal lengths. These focal lengths respectively correspond to zoom ratios of 1x, 2x, and 5x (rounded to integers). Figure 2 is an enlarged image of the dotted area in Figure 1. The resolution of the single-lens reflex camera image remains unchanged during the process of zooming in and out, but the resolution of the panoramic camera image changes. In practical applications, the three regional images can also be respectively matched with the images with different focal lengths in [3], so that the content included in the regional images is as similar as possible to the content included in the corresponding images in [3]. For example, match the image under 55 mm in Figure 2 with the image under 55 mm focal length in [3], match the image under 135 mm in Figure 2 with the image under 135 mm focal length in [3], and match the image under 250 mm in Figure 2 with the image under 250 mm focal length in [3].

[0079] Figure 4 Figure [4] shows the generator of the conditional generative adversarial network, where the condition is defined by one-hot vectors representing different zoom ratios. This conditional neural network is a generative adversarial network (GAN), which consists of two parts: a generator and a discriminator. In the training stage [4], the regional images and the labels representing the zoom ratios (where the images are taken at these zoom ratios) are concatenated and sent to the generator for training. The definition of the one-hot vectors is as follows: (1, 0, 0, 0, 0), (0, 1, 0, 0, 0), (0, 0, 1, 0, 0), (0, 0, 0, 1, 0), (0, 0, 0, 0, 1) represent 1x, 2x, 3x, 4x, and 5x zoom ratios respectively. Figure 4 Figure [6] shows two different zoom ratios: 1x (Figure 5) and 5x (Figure 6) and the corresponding input images. Network [7] is a standard convolutional neural network, including convolutional layers, activation layers, and pooling layers, etc. In each iteration during training, the parameters of the concatenated input tensors are updated. The output of the network is the super-resolution images [8] and [9] at different zoom ratios and the trained model. The mathematical expressions of the input image and the label can be I i and y, and the output image can be G(I i |y).

[0080] It should be noted that although the zoom ratio is defined as a one-hot vector in Figure 4 [4], in practical applications, other variables such as focal length and depth of field can also be defined as one-hot vectors.

[0081] Figure 5 Figure

[19] shows the discriminator of the conditional generative adversarial network. The discriminator is part of the entire training stage. It takes the generated image G(I i|y), the real image x, and the label y as the concatenated input 10, and then extract features through the convolutional neural network 11. The result 12 is a scalar between 0 and 1. This scalar represents the probability that the generated image is real (or fake). It should be noted that in Figure 5 only 1 zoom ratio is shown. In practical applications, the concatenated ones can be multiple inputs at different zoom ratios, such as 2 times, 3 times, etc.

[0082] Figure 6 Shows the inference (testing) phase 13 of the conditional generative adversarial network. The input is the panoramic images at different zoom ratios concatenated together. Since only one model is trained for different zoom ratios during the training process, the input of the conditional neural network model can be images at multiple zoom ratios. Images and corresponding labels in the cases of two different zoom ratios: 1 time and 5 times are shown in 14 and 15. The conditional neural network model 16 is the same as the model generated by the Figure 4 generator described in. The output of the conditional neural network model is the super-resolution images 17 and 18 corresponding to the 1-time and 5-time zoom ratios. It should be noted that in the inference (testing) phase, the input of the conditional neural network model can be either the regional image of the low-quality panoramic image or the entire low-quality panoramic image. If the input is the entire low-quality panoramic image, the inputs at each zoom ratio can be understood as one image. At this time, if different conditions (such as one-hot vectors) are input, the network will generate the super-resolution image corresponding to this condition.

[0083] Figure 7 Shows the generator of the conditional generative adversarial network, where the condition is defined by the VGG features of the input images representing different zoom ratios. In the training phase 19, the captured low-quality images and the labels representing the zoom ratios (where the images were captured at these zoom ratios) are concatenated and sent to the generator for training. Figure 7 The VGG features (labels) are defined in. Figure 7 Shows two different zoom ratios: 1 time (20) and 5 times (21) and the corresponding input images. The convolutional neural network 22 is a standard convolutional neural network, including convolutional layers, activation layers, pooling layers, etc. In each iteration of training, the parameters of the tensor of the concatenated input are updated. The output of the convolutional neural network is the super-resolution images 23 and 24 at different zoom ratios and the trained model. The mathematical expressions of the input images and labels can be I i and φ i . The output image can be G(I i |φ i)。VGG19-22 represents the second convolutional layer before the second pooling layer in the VGG network. This value can be obtained empirically. In specific applications, VGG19-22, VGG19-54, etc. can be used as VGG features.

[0084] Figure 8 Shows the discriminator of the conditional generative adversarial network. The discriminator is part of the entire training phase. It takes the generated image G(I i |φ i ), the real image x, and the label φ i as concatenated inputs 25, and then extracts features through a convolutional neural network 26. The discriminant result 27 is a scalar between 0 and 1, which represents the probability that the generated image is real (or fake). It should be noted that Figure 8 only shows 1 zoom ratio. In practical applications, the concatenated ones can be multiple inputs at different zoom ratios, such as 2 times, 3 times, etc.

[0085] Figure 9 Shows the inference (testing) phase 28 of the conditional generative adversarial network. The input is the images at different zoom ratios concatenated together. Since only one model is trained for different zoom ratios, the input in the embodiments of the present invention can be images at multiple zoom ratios. Figure 9 Shows the images and corresponding labels in the cases of two different zoom ratios: 1 time and 5 times respectively in 29 and 30. The conditional neural network model 31 is the same as the model generated by the Figure 7 generator described in. The output of the conditional neural network model is the super-resolution images 32 and 33 corresponding to the 1-time and 5-time zoom ratios.

[0086] As can be seen from the above description, in the embodiments of the present invention, by using different zoom ratios as the conditions of the generative adversarial network, only one network model can be trained to achieve the purpose of super-resolution, without training multiple network models. Since a network corresponding to different zoom ratios is trained, the quality of the obtained super-resolution images is better. The solution of the embodiments of the present invention can be applied to fields such as panoramic image and video enhancement, panoramic image and video super-resolution, virtual reality / augmented reality, and three-dimensional reconstruction. Its core technology of training with low-quality and high-quality images can be applied to ordinary image enhancement tasks, such as improving the image quality of smartphone images to that of high-quality cameras.

[0087] The embodiments of the present invention also provide an image processing device. Refer to Figure 10 , Figure 10 is the structural diagram of the image processing device provided by the embodiments of the present invention. As shown in Figure 10As shown in the figure, the image processing device 1000 includes:

[0088] A first acquisition module 1001, configured to acquire an image to be processed; a second acquisition module 1002, configured to use the image to be processed and conditions corresponding to the image to be processed as inputs to a conditional neural network model, and obtain a super-resolution image of the image to be processed. Wherein, the conditional neural network model is trained by using at least two regional images of a first sampled image and conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, wherein the quality of the second sampled image is higher than that of the first sampled image.

[0089] Optionally, as Figure 11 shown, the device may further include: a training module 1003, configured to train the conditional neural network model.

[0090] Optionally, as Figure 12 shown, the training module 1003 includes: an extraction sub-module 10031, configured to extract at least two regional images from the first sampled image; a training sub-module 10032, configured to input the at least two regional images and conditions corresponding to the at least two regional images into a generator of a conditional generative adversarial network, and obtain a super-resolution image corresponding to the at least two regional images and the conditional neural network model.

[0091] Optionally, as Figure 13 shown, the training module 1003 further includes: a matching sub-module 10033, configured to match the at least two regional images with second sampled images at at least two zoom ratios respectively; the training sub-module 10032, configured to input the at least two regional images after matching and conditions corresponding to the at least two regional images into a generator of a conditional generative adversarial network, and obtain a super-resolution image corresponding to the at least two regional images and the conditional neural network model.

[0092] Optionally, as Figure 14 shown, the training module 1003 further includes: a discrimination sub-module 10034, configured to input the second sampled images at the at least two zoom ratios, the super-resolution images corresponding to the at least two regional images, and conditions corresponding to the at least two regional images into a discriminator of the conditional generative adversarial network, and obtain a discrimination result.

[0093] Optionally, the condition includes a one-hot vector. The one-hot vector is determined by any one of the following information:

[0094] The zoom ratio of the zoom lens;

[0095] Depth of field of a zoom lens;

[0096] Focal length of a zoom lens.

[0097] Optionally, the conditions include: VGG features.

[0098] The device provided by the embodiments of the present invention can execute the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated here in this embodiment.

[0099] See Figure 15 , the embodiments of the present invention also provide a hardware structure of an electronic device. As Figure 15 shown, the electronic device 1500 includes:

[0100] A processor 1502; and

[0101] A memory 1504, in which program instructions are stored. When the program instructions are run by the processor, the processor 1502 is caused to execute the following steps:

[0102] Obtain an image to be processed;

[0103] Use the image to be processed and the conditions corresponding to the image to be processed as inputs to a conditional neural network model to obtain a super-resolution image of the image to be processed;

[0104] Wherein, the conditional neural network model is trained by using at least two regional images of a first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, wherein the quality of the second sampled images is higher than that of the first sampled image.

[0105] Further, as Figure 15 shown, the electronic device 1500 may further include a network interface 1501, an input device 1503, a hard disk 1505, and a display device 1506.

[0106] The above-mentioned various interfaces and devices can be interconnected through a bus architecture. The bus architecture can include any number of interconnected buses and bridges. Specifically, various circuits represented by one or more central processing units (CPUs) represented by the processor 1502 and one or more memories represented by the memory 1504 are connected together. The bus architecture can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. It can be understood that the bus architecture is used to realize the connection and communication between these components. In addition to the data bus, the bus architecture also includes a power bus, a control bus, and a status signal bus, which are well known in the art and will not be described in detail herein.

[0107] The network interface 1501 can be connected to a network (such as the Internet, a local area network, etc.), receive data from the network, and can save the received data in the hard disk 1505.

[0108] The input device 1503 can receive various instructions input by an operator and send them to the processor 1502 for execution. The input device 1503 can include a keyboard or a pointing device (for example, a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0109] The display device 1506 can display the result obtained by the processor 1502 executing instructions.

[0110] The memory 1504 is used to store programs and data necessary for the operation of the operating system, as well as data such as intermediate results in the calculation process of the processor 1502.

[0111] It can be understood that the memory 1504 in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. The memory 1504 of the devices and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0112] In some embodiments, the memory 1504 stores the following elements, executable modules or data structures, or subsets thereof, or extended sets thereof: an operating system 15041 and an application program 15042.

[0113] Among them, the operating system 15041 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., and is used to implement various basic services and process hardware-based tasks. The application program 15042 includes various application programs, such as a browser, etc., and is used to implement various application services. The program for implementing the method of the embodiments of the present invention can be included in the application program 15042.

[0114] The image processing method disclosed in the above embodiments of the present invention can be applied to or implemented by the processor 1502. The processor 1502 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above image processing method can be completed by the integrated logic circuit in the hardware of the processor 1502 or instructions in the form of software. The above-mentioned processor 1502 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1504, and the processor 1502 reads the information in the memory 1504 and combines its hardware to complete the steps of the above method.

[0115] It can be understood that these embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For a hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.

[0116] For a software implementation, the technologies described herein can be implemented by modules (such as procedures, functions, etc.) that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.

[0117] Specifically, when the program is executed by the processor 1502, the following steps can also be implemented:

[0118] Train the conditional neural network model.

[0119] Specifically, when the program is executed by the processor 1502, the following steps can also be implemented:

[0120] Extract at least two region images from the first sampled image;

[0121] Input the at least two regional images and the conditions corresponding to the at least two regional images into the generator of the conditional generative adversarial network to obtain the super-resolution images corresponding to the at least two regional images and the conditional neural network model.

[0122] Specifically, when the program is executed by the processor 1502, the following steps may also be implemented:

[0123] Match the at least two regional images with the second sampled images at at least two zoom ratios respectively; input the matched at least two regional images and the conditions corresponding to the at least two regional images into the generator of the conditional generative adversarial network.

[0124] Specifically, when the program is executed by the processor 1502, the following steps may also be implemented:

[0125] Input the second sampled images at the at least two zoom ratios, the super-resolution images corresponding to the at least two regional images, and the conditions corresponding to the at least two regional images into the discriminator of the conditional generative adversarial network to obtain a discrimination result.

[0126] Among them, the condition includes a one-hot vector.

[0127] Among them, the one-hot vector is determined by any one of the following information:

[0128] The zoom ratio of the zoom lens;

[0129] The depth of field of the zoom lens;

[0130] The focal length of the zoom lens.

[0131] Among them, the condition includes: VGG features.

[0132] The electronic device provided in the embodiment of the present invention can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0133] The embodiment of the present invention also provides a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements each process of the above image processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid state drives (SSD)).

[0134] It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0136] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.

Claims

1. An image processing method, characterized in that, Including: Obtain an image to be processed; Use the image to be processed and the conditions corresponding to the image to be processed as the input of a conditional neural network model to obtain a super-resolution image of the image to be processed; Wherein, the conditional neural network model is trained using at least two regional images of a first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, wherein the quality of the second sampled image is higher than that of the first sampled image; Wherein, training the conditional neural network model includes: Extract at least two regional images from the first sampled image; Match the at least two regional images with second sampled images at at least two zoom ratios respectively, such that the content included in the at least two regional images after matching is respectively the same as that of the second sampled images at at least two zoom ratios, or such that the degree of sameness of the content included in the at least two regional images after matching with the content included in the second sampled images at at least two zoom ratios meets a preset condition; Input the at least two regional images after matching and the conditions corresponding to the at least two regional images into the generator of a conditional generative adversarial network to obtain the super-resolution images corresponding to the at least two regional images and the conditional neural network model.

2. The method according to claim 1, characterized in that, Training the conditional neural network model further includes: Input the second sampled images at the at least two zoom ratios, the super-resolution images corresponding to the at least two regional images, and the conditions corresponding to the at least two regional images into the discriminator of the conditional generative adversarial network to obtain a discrimination result.

3. The method according to claim 1, characterized in that, The conditions include one-hot vectors.

4. The method according to claim 3, characterized in that, The one-hot vectors are determined by any one of the following information: The zoom ratio of a zoom lens; The depth of field of a zoom lens; The focal length of a zoom lens.

5. The method according to claim 1, characterized in that, The conditions include: Visual Geometry Group (VGG) features.

6. An image processing apparatus, characterized in that, Including: A first obtaining module, configured to obtain an image to be processed; A second obtaining module, configured to use the image to be processed and the conditions corresponding to the image to be processed as the input of a conditional neural network model to obtain a super-resolution image of the image to be processed; Wherein, the conditional neural network model is trained using at least two regional images of a first sampled image and the conditions corresponding to the at least two regional images; the at least two regional images respectively correspond to second sampled images at at least two zoom ratios, wherein the quality of the second sampled image is higher than that of the first sampled image; Wherein, the device further includes: a training module, including: An extraction sub-module, configured to extract at least two regional images from the first sampled image; a matching sub-module, configured to respectively match the at least two regional images with second sampled images at at least two zoom ratios, such that the content included in the at least two regional images after matching is respectively the same as the second sampled images at at least two zoom ratios, or such that the degree of sameness of the content included in the at least two regional images after matching with the content included in the second sampled images at at least two zoom ratios meets a preset condition; a training sub-module, configured to input the at least two regional images after matching and the conditions corresponding to the at least two regional images into a generator of a conditional generative adversarial network to obtain super-resolution images corresponding to the at least two regional images and the conditional neural network model.

7. An electronic device, comprising: A memory, a processor, and a program stored on the memory and executable on the processor; wherein the processor is configured to read the program in the memory to implement the steps in the image processing method according to any one of claims 1 to 5.

8. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps in the image processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Super-resolution reconstruction method based on conditional generative adversarial network

    CN109978762A

  • Multi-supervision image super-resolution reconstruction method based on generative adversarial network

    CN110322403A