Image processing method and device, electronic equipment and storage medium
By fusing camera intrinsic features into image super-resolution technology, the inter-domain drift problem is solved, the super-resolution effect of images is improved, and the ability to enhance image resolution is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XI AN SENSETIME INTELLIGENT TECH CO LTD
- Filing Date
- 2022-06-01
- Publication Date
- 2026-04-17
AI Technical Summary
Existing image super-resolution techniques suffer from inter-domain drift, which affects the super-resolution effect.
By introducing camera intrinsic features to assist image super-resolution, and using camera intrinsic representation networks and perception networks to fuse image features, the spatial domain difference between the super-resolution image and the original image is reduced, thus solving the inter-domain drift problem.
It improves the super-resolution effect of images, enhances the ability to improve image resolution, and optimizes the training effect of network models.
Smart Images

Figure CN114862682B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device and storage medium. Background Technology
[0002] With the continuous development of imaging technology, it has important applications in industry, agriculture, medicine, and other fields. The quality of images acquired by imaging devices is affected by the resolution of those devices. However, the materials used to manufacture imaging devices are generally expensive, and the manufacturing process is complex. Therefore, improving the resolution of imaging devices through hardware upgrades is costly.
[0003] Image super-resolution technology is a technique that can convert low-resolution images into high-resolution images, thereby improving the quality of the original image. However, current image super-resolution techniques suffer from inter-domain drift, which affects the super-resolution effect. Summary of the Invention
[0004] This application provides an image processing method, apparatus, electronic device, and storage medium that can improve the super-resolution effect of images.
[0005] In a first aspect, embodiments of this application provide an image processing method, the method comprising:
[0006] Obtain the image to be processed;
[0007] The camera intrinsic parameter features of the image to be processed are extracted, and the camera intrinsic parameter features are used to characterize the resolution performance of the imaging device that acquires the image to be processed;
[0008] The image features of the image to be processed are extracted, and the camera intrinsic features of the image to be processed are fused with the image features to obtain the super-resolution image corresponding to the image to be processed, wherein the resolution of the super-resolution image corresponding to the image to be processed is higher than the resolution of the image to be processed.
[0009] In the above embodiments, camera intrinsic features are introduced to assist in image super-resolution. During the super-resolution process of the image to be processed, by fusing the camera intrinsic features and image features of the image to be processed, not only the features of the image to be processed itself are considered, but also the intrinsic features of the imaging device that acquired the image to be processed. This helps to reduce the spatial domain difference between the obtained super-resolution image and the original image to be processed, thereby helping to solve the inter-domain drift problem and improve the super-resolution effect of the image.
[0010] In some possible implementations, extracting the camera intrinsic features of the image to be processed includes:
[0011] The camera intrinsic features of the image to be processed are extracted using a camera intrinsic representation network;
[0012] The camera intrinsic parameter representation network is trained based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category.
[0013] In the above embodiments, the camera intrinsic parameter representation network trained based on the difference between the resolution category recognition result of the first sample image and the corresponding real resolution category can effectively distinguish images of different resolutions, which is beneficial for quickly and accurately extracting camera intrinsic parameter features in the image to characterize the resolution performance of the imaging device.
[0014] In some possible implementations, the training method for the camera intrinsic representation network includes:
[0015] Obtain a first sample image and its true resolution category, wherein the first sample image includes images acquired by two imaging devices with different resolutions;
[0016] The camera intrinsic feature of the first sample image is extracted by the camera intrinsic feature representation network, and the classification network is used to classify and identify the first sample image based on the camera intrinsic feature of the first sample image to obtain the resolution category identification result of the first sample image.
[0017] Based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category, the parameters of the camera intrinsic parameter representation network and the classification network are adjusted until the first training termination condition is met.
[0018] In the above embodiments, the camera intrinsic parameter representation network trained can extract distinctive camera intrinsic parameter features from images from imaging devices with different resolutions. Using the camera intrinsic parameter features extracted by the camera intrinsic parameter representation network to assist in image super-resolution is beneficial to improving the super-resolution effect of the image.
[0019] In some possible implementations, the step of extracting image features from the image to be processed and fusing the camera intrinsic features of the image to be processed with the image features to obtain a super-resolution image corresponding to the image to be processed includes:
[0020] Using a first perceptual network, image features of the image to be processed are extracted, and the camera intrinsic parameter features are fused with the image features to obtain a super-resolution image corresponding to the image to be processed.
[0021] The first perceptual network is trained based on the difference between the first super-resolution image and the second sample image. The first super-resolution image is obtained by increasing the resolution of the third sample image. The second sample image has the same content as the third sample image, and the resolution of the second sample image is higher than that of the third sample image.
[0022] In the above embodiments, the first perceptual network trained based on the difference between the first super-resolution image obtained by improving the resolution of the third sample image and the second sample image can be used to perform super-resolution on the image to obtain the corresponding super-resolution image, which is beneficial to quickly improve the image resolution.
[0023] In some possible implementations, the training method for the first perceptual network includes:
[0024] Acquire the second sample image, the third sample image, and the camera intrinsic features of the third sample image;
[0025] The image features of the third sample image are extracted through the first perception network, and the camera intrinsic features of the third sample image are fused with the image features to obtain the first super-resolution image.
[0026] Based on the difference between the first super-resolution image and the second sample image, the parameters of the first perceptual network are adjusted until the second training termination condition is met.
[0027] In the above embodiments, the first perceptual network trained can perceive the camera intrinsic features of low-resolution images and integrate the camera intrinsic features into the image features to generate corresponding super-resolution images. This is beneficial for realizing the spatial domain transfer from the super-resolution image to the original low-resolution image, thereby helping to solve the inter-domain drift problem and improve the super-resolution effect of the image.
[0028] In some possible implementations, the step of extracting image features from the image to be processed and fusing the camera intrinsic features of the image to be processed with the image features to obtain a super-resolution image corresponding to the image to be processed includes:
[0029] Using a second sensing network, image features of the image to be processed are extracted, and the camera intrinsic parameter features are fused with the image features to obtain a super-resolution image corresponding to the image to be processed.
[0030] The second perceptual network is trained based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image. The first super-resolution image is obtained by increasing the resolution of the third sample image, and the degraded resolution image is obtained by decreasing the resolution of the first super-resolution image. The second sample image has the same content as the third sample image, and the resolution of the second sample image is higher than that of the third sample image.
[0031] In the above embodiments, the second perceptual network, trained based on the difference between the first super-resolution image (obtained by increasing the resolution of the third sample image) and the second sample image, and the difference between the degraded resolution image (obtained by decreasing the resolution of the first super-resolution image) and the third sample image, can be used to perform super-resolution on images to obtain corresponding super-resolution images, which is beneficial for quickly improving image resolution. Furthermore, since the training process of the second perceptual network considers the difference between the degraded resolution image and the third sample image, this helps to limit the search space of network parameters during training and optimize the training effect of the network model.
[0032] In some possible implementations, the training method for the second perceptual network includes:
[0033] Acquire the second sample image, the third sample image, the camera intrinsic features of the second sample image, and the camera intrinsic features of the third sample image;
[0034] The second sensing network is used to extract the image features of the third sample image, and the camera intrinsic features of the third sample image are fused with the image features to obtain the first super-resolution image.
[0035] By using a recurrent sensing network, the image features of the first super-resolution image are extracted, and the camera intrinsic features of the second sample image are fused with the image features of the first super-resolution image to obtain a degraded resolution image.
[0036] Based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image, the parameters of the second perceptual network and the recurrent perceptual network are adjusted until the third training termination condition is met.
[0037] In the above embodiments, the trained second perceptual network can perceive the camera intrinsic features of low-resolution images and integrate these features into its image features to generate a corresponding super-resolution image. This facilitates spatial domain transfer from the super-resolution image to the original low-resolution image, thereby helping to solve the inter-domain drift problem and improve the super-resolution effect. Furthermore, a recurrent perceptual network is incorporated into the training of the second perceptual network. This network perceives the camera intrinsic features of high-resolution images and integrates them into the image features of the super-resolution image, generating a degraded low-resolution image to aid training. This helps to limit the search space of network parameters during training, solve the ill-posed problem, and optimize the training effect of the network model.
[0038] In some possible implementations, the step of extracting image features from the third sample image and fusing the camera intrinsic features of the third sample image with the image features to obtain the first super-resolution image includes:
[0039] The third sample image is convolved and downsampled to obtain a first feature, the first feature is downsampled to obtain a second feature, and the second feature is downsampled to obtain a third feature;
[0040] The camera intrinsic features of the third sample image are fully connected and reconstructed to obtain the first convolution kernel. The first convolution kernel is used to convolve the third feature to obtain the fourth feature. The fourth feature is superimposed with the third feature to obtain the fifth feature.
[0041] The fifth feature is upsampled to obtain the sixth feature. The sixth feature is fused with the second feature and then upsampled to obtain the seventh feature. The seventh feature is fused with the first feature and then upsampled and convolved to obtain the eighth feature.
[0042] The eighth feature is superimposed on the third sample image to obtain the first super-resolution image.
[0043] In the above embodiments, by fully fusing the camera intrinsic features of the low-resolution image with the image features of the low-resolution image, the fused features can contain richer information, which is beneficial to improving the resolution of the low-resolution image and obtaining the corresponding super-resolution image.
[0044] In some possible implementations, the step of extracting image features from the first super-resolution image and fusing the camera intrinsic features of the second sample image with the image features of the first super-resolution image to obtain a degraded resolution image includes:
[0045] The ninth feature is obtained by convolution and downsampling the first super-resolution image;
[0046] The camera intrinsic features of the second sample image are fully connected and reconstructed to obtain the second convolution kernel. The second convolution kernel is used to convolve the ninth feature to obtain the tenth feature. The tenth feature is superimposed with the ninth feature to obtain the eleventh feature.
[0047] The eleventh feature is obtained by upsampling and convolution.
[0048] The twelfth feature is superimposed on the first super-resolution image to obtain a degraded resolution image.
[0049] In the above embodiments, by fully fusing the camera intrinsic features of the high-resolution image with the image features of the super-resolution image, the fused features can contain richer information, which is beneficial for better reducing the resolution of the super-resolution image and obtaining the corresponding degraded resolution image.
[0050] Secondly, embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0051] The acquisition unit is used to acquire the image to be processed;
[0052] The first processing unit is used to extract camera intrinsic parameter features of the image to be processed, wherein the camera intrinsic parameter features are used to characterize the resolution performance of the imaging device that acquires the image to be processed;
[0053] The second processing unit is used to extract image features of the image to be processed, fuse the camera intrinsic features of the image to be processed with the image features, and obtain a super-resolution image corresponding to the image to be processed, wherein the resolution of the super-resolution image is higher than the resolution of the image to be processed.
[0054] In some possible implementations, the first processing unit is specifically used to: extract camera intrinsic parameter features of the image to be processed using a camera intrinsic parameter representation network; wherein the camera intrinsic parameter representation network is trained based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category.
[0055] In some possible implementations, the apparatus further includes a first training unit for training a camera intrinsic parameter representation network; the first training unit is specifically used for:
[0056] Obtain a first sample image and its true resolution category, wherein the first sample image includes images acquired by two imaging devices with different resolutions;
[0057] The camera intrinsic feature of the first sample image is extracted by the camera intrinsic feature representation network, and the classification network is used to classify and identify the first sample image based on the camera intrinsic feature of the first sample image to obtain the resolution category identification result of the first sample image.
[0058] Based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category, the parameters of the camera intrinsic parameter representation network and the classification network are adjusted until the first training termination condition is met.
[0059] In some possible implementations, the second processing unit is specifically used to: extract image features of the image to be processed using a first perceptual network, fuse the camera intrinsic parameter features with the image features, and obtain a super-resolution image corresponding to the image to be processed;
[0060] The first perceptual network is trained based on the difference between the first super-resolution image and the second sample image. The first super-resolution image is obtained by increasing the resolution of the third sample image. The second sample image has the same content as the third sample image, and the resolution of the second sample image is higher than that of the third sample image.
[0061] In some possible implementations, the apparatus further includes a second training unit for training the first perceptual network; the second training unit is specifically used for:
[0062] Acquire the second sample image, the third sample image, and the camera intrinsic features of the third sample image;
[0063] The image features of the third sample image are extracted through the first perception network, and the camera intrinsic features of the third sample image are fused with the image features to obtain the first super-resolution image.
[0064] Based on the difference between the first super-resolution image and the second sample image, the parameters of the first perceptual network are adjusted until the second training termination condition is met.
[0065] In some possible implementations, the second processing unit is specifically used to: extract image features of the image to be processed using a second sensing network, fuse the camera intrinsic parameter features with the image features, and obtain a super-resolution image corresponding to the image to be processed;
[0066] The second perceptual network is trained based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image. The first super-resolution image is obtained by increasing the resolution of the third sample image, and the degraded resolution image is obtained by decreasing the resolution of the first super-resolution image. The second sample image has the same content as the third sample image, and the resolution of the second sample image is higher than that of the third sample image.
[0067] In some possible implementations, the device further includes a third training unit for training a second perceptual network; the third training unit is specifically used for:
[0068] Acquire the second sample image, the third sample image, the camera intrinsic features of the second sample image, and the camera intrinsic features of the third sample image;
[0069] The second sensing network is used to extract the image features of the third sample image, and the camera intrinsic features of the third sample image are fused with the image features to obtain the first super-resolution image.
[0070] By using a recurrent sensing network, the image features of the first super-resolution image are extracted, and the camera intrinsic features of the second sample image are fused with the image features of the first super-resolution image to obtain a degraded resolution image.
[0071] Based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image, the parameters of the second perceptual network and the recurrent perceptual network are adjusted until the third training termination condition is met.
[0072] In some possible implementations, when the second training unit or the third training unit extracts image features from the third sample image and fuses the camera intrinsic features of the third sample image with the image features to obtain the first super-resolution image, it is specifically used for:
[0073] The third sample image is convolved and downsampled to obtain a first feature, the first feature is downsampled to obtain a second feature, and the second feature is downsampled to obtain a third feature;
[0074] The camera intrinsic features of the third sample image are fully connected and reconstructed to obtain the first convolution kernel. The first convolution kernel is used to convolve the third feature to obtain the fourth feature. The fourth feature is superimposed with the third feature to obtain the fifth feature.
[0075] The fifth feature is upsampled to obtain the sixth feature. The sixth feature is fused with the second feature and then upsampled to obtain the seventh feature. The seventh feature is fused with the first feature and then upsampled and convolved to obtain the eighth feature.
[0076] The eighth feature is superimposed on the third sample image to obtain the first super-resolution image.
[0077] In some possible implementations, when the third training unit extracts image features from the first super-resolution image and fuses the camera intrinsic features of the second sample image with the image features of the first super-resolution image to obtain a degraded resolution image, it is specifically used for:
[0078] The ninth feature is obtained by convolution and downsampling the first super-resolution image;
[0079] The camera intrinsic features of the second sample image are fully connected and reconstructed to obtain the second convolution kernel. The second convolution kernel is used to convolve the ninth feature to obtain the tenth feature. The tenth feature is superimposed with the ninth feature to obtain the eleventh feature.
[0080] The eleventh feature is obtained by upsampling and convolution.
[0081] The twelfth feature is superimposed on the first super-resolution image to obtain a degraded resolution image.
[0082] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein when the processor executes the computer instructions, the electronic device performs the method as described in the first aspect above and any possible implementation thereof.
[0083] Fourthly, embodiments of this application provide an electronic device, including: a processor, a transmitting device, an input device, an output device, and a memory, wherein the memory is used to store computer program code, the computer program code including computer instructions, and when the processor executes the computer instructions, the electronic device performs the method as described in the first aspect above and any possible implementation thereof.
[0084] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method described in the first aspect and any possible implementation thereof.
[0085] Sixthly, embodiments of this application provide a computer program product comprising a computer program or instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect and any possible implementation thereof.
[0086] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description
[0087] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.
[0088] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0089] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of this application;
[0090] Figure 2 This is a flowchart illustrating a training method for a camera intrinsic parameter representation network provided in an embodiment of this application.
[0091] Figure 3 This is a training schematic diagram of an image resolution recognition network provided in an embodiment of this application;
[0092] Figure 4 This is a flowchart illustrating a training method for a first perceptual network provided in an embodiment of this application;
[0093] Figure 5 This is a training schematic diagram of a first perceptual network provided in an embodiment of this application;
[0094] Figure 6 This is a schematic diagram of the structure of a downsampling module, a sensing module, and an upsampling module provided in an embodiment of this application;
[0095] Figure 7 This is a flowchart illustrating a training method for a second perceptual network provided in an embodiment of this application;
[0096] Figure 8 This is a training schematic diagram of a second sensing network provided in an embodiment of this application;
[0097] Figure 9 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;
[0098] Figure 10This is a schematic diagram of another image processing device provided in an embodiment of this application;
[0099] Figure 11 This is a schematic diagram of the hardware structure of an image processing device provided in an embodiment of this application. Detailed Implementation
[0100] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0101] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0102] It should be understood that in this application, "at least one (item)" means one or more, "more than one" means two or more, and "at least two (items)" means two or three or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " indicates that the preceding and following related objects are in an "or" relationship, meaning any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0103] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0104] With the continuous development of imaging technology, it has important applications in industry, agriculture, and medicine. For example, thermal imaging technology is widely used for non-contact human body temperature measurement, and its applications in civilian use are also increasing. However, the materials used to manufacture thermal imaging cameras are very expensive, and the manufacturing process is very complex. Improving the resolution of thermal imaging cameras through hardware upgrades is extremely costly. Therefore, thermal imaging super-resolution technology is an important means to reduce the cost of obtaining high-quality thermal images and make high-quality thermal imaging widely available for civilian use.
[0105] In thermal imaging super-resolution technology, super-resolution networks are used to super-resolution the input low-resolution image and output a corresponding super-resolution image. However, current super-resolution networks often have input and output images that are not in the same spatial domain, resulting in inter-domain drift, which degrades super-resolution performance and affects the super-resolution effect.
[0106] Based on this, embodiments of this application provide an image processing method to solve the inter-domain drift problem in image super-resolution and improve the super-resolution effect of the image.
[0107] The entity executing the image processing method can be an image processing device. For example, the image processing method can be executed by a terminal device, a server, or other processing devices. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the image processing method can be implemented by a processor calling computer-readable instructions stored in memory.
[0108] Please see Figure 1 , Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of this application. The image processing method includes the following steps S101 to S103.
[0109] S101, Obtain the image to be processed.
[0110] An image to be processed refers to an image whose resolution needs to be improved. An image to be processed can be an image acquired by an imaging device with a lower resolution; for example, the imaging device acquiring the image to be processed is a thermal imaging sensor.
[0111] S102, extract the camera intrinsic parameter features of the image to be processed. The camera intrinsic parameter features are used to characterize the resolution performance of the imaging device that acquires the image to be processed.
[0112] The resolution performance of an imaging device is determined by its internal parameters, and the resolution of an image is related to the resolution performance of the imaging device that acquires the image. For example, the higher the resolution of the imaging device, the higher the resolution of the image acquired by the imaging device; the lower the resolution of the imaging device, the lower the resolution of the image acquired by the imaging device.
[0113] The camera intrinsic parameter features of the image to be processed can be understood as the information contained in the image to be processed that is related to the resolution performance of the imaging device that acquired the image. It reflects the intrinsic parameter features of the imaging device that acquired the image.
[0114] S103, extract the image features of the image to be processed, fuse the camera intrinsic parameter features of the image to be processed with the image features to obtain the super-resolution image corresponding to the image to be processed, wherein the resolution of the super-resolution image corresponding to the image to be processed is higher than the resolution of the image to be processed.
[0115] Image features of the image to be processed can be understood as features related to the image content of the image to be processed, reflecting the inherent characteristics of the image itself. The camera intrinsic features of the image to be processed are fused with the image features; that is, camera intrinsic features are incorporated into the image features to assist in image super-resolution, obtaining the super-resolution image corresponding to the image to be processed. Here, image super-resolution refers to increasing the resolution of an image, and the super-resolution image corresponding to the image to be processed is the image after the resolution of the image to be processed has been increased.
[0116] In the above embodiments, camera intrinsic features are introduced to assist in image super-resolution. During the super-resolution process of the image to be processed, by fusing the camera intrinsic features and image features of the image to be processed, not only the features of the image to be processed itself are considered, but also the intrinsic features of the imaging device that acquired the image to be processed. This helps to reduce the spatial domain difference between the obtained super-resolution image and the original image to be processed, thereby helping to solve the inter-domain drift problem and improve the super-resolution effect of the image.
[0117] In one possible implementation, the camera intrinsic parameter features of the image to be processed are extracted. Specifically, the camera intrinsic parameter features of the image to be processed are extracted using a camera intrinsic parameter representation network. The camera intrinsic parameter representation network is trained based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category.
[0118] A camera intrinsic parameter representation network is a network used to extract camera intrinsic parameter features of an image. The image to be processed is input into the camera intrinsic parameter representation network, and the camera intrinsic parameter representation network outputs the camera intrinsic parameter features of the image to be processed.
[0119] Optionally, the camera intrinsic parameter representation network is a sub-network in the image resolution recognition network used for feature extraction. Here, the image resolution recognition network refers to the network used to identify image resolution categories. Specifically, the image resolution recognition network includes the following two sub-networks: the camera intrinsic parameter representation network and the classification network. The processing steps of the image resolution recognition network include: extracting camera intrinsic parameter features of the image using the camera intrinsic parameter representation network, and performing classification and recognition based on the camera intrinsic parameter features of the image using the classification network to obtain the image resolution category recognition result.
[0120] Image resolution recognition networks can be trained through supervised learning. Specifically, they can be trained based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category. The training objective is to make the resolution category recognition result closer to the corresponding true resolution category. It can be understood that training the image resolution recognition network includes training the parameters of the camera intrinsic parameter representation network and the classification network. Once the image resolution recognition network is trained, it means that both the camera intrinsic parameter representation network and the classification network are also trained. The trained camera intrinsic parameter representation network is then used to extract the camera intrinsic parameter features of the image to be processed.
[0121] In the above embodiments, the camera intrinsic parameter representation network trained based on the difference between the resolution category recognition result of the first sample image and the corresponding real resolution category can effectively distinguish images of different resolutions, which is beneficial for quickly and accurately extracting camera intrinsic parameter features in the image to characterize the resolution performance of the imaging device.
[0122] Please see Figure 2 , Figure 2 This is a flowchart illustrating a training method for a camera intrinsic parameter representation network provided in an embodiment of this application. The training method includes the following steps S201 to S203.
[0123] S201, Obtain a first sample image and its true resolution category. The first sample image includes images acquired by two imaging devices with different resolutions.
[0124] The first sample image can be selected from a first sample image library, which includes images acquired by imaging devices with two different resolution performances. The true resolution category of the first sample image includes two categories: low resolution and high resolution. Specifically, the true resolution category of the image acquired by the imaging device with relatively low resolution (referred to as the low-resolution imaging device) is low resolution, and the true resolution category of the image acquired by the imaging device with relatively high resolution (referred to as the high-resolution imaging device) is high resolution.
[0125] S202, the camera intrinsic feature of the first sample image is extracted through the camera intrinsic feature representation network, and the classification network is used to perform classification and recognition based on the camera intrinsic feature of the first sample image to obtain the resolution category recognition result of the first sample image.
[0126] Please see Figure 3 , Figure 3 This is a training diagram of an image resolution recognition network provided in an embodiment of this application. Wherein, C1 represents the camera intrinsic parameter representation network; C2 represents the classification network; A1 represents the first sample image acquired by the low-resolution imaging device (referred to as the low-resolution sample image), whose true resolution category is low resolution; A2 represents the first sample image acquired by the high-resolution imaging device (referred to as the high-resolution sample image), whose true resolution category is high resolution; B1 and B2 represent the camera intrinsic parameter features of the low-resolution sample image and the high-resolution sample image, respectively; T1 and T2 represent the resolution category recognition results of the low-resolution sample image and the high-resolution sample image, respectively; Loss represents the cross-entropy loss.
[0127] Optionally, the camera intrinsic parameter representation network C1 adopts the convolutional layer structure of the ResNet18 network, the classification network C2 adopts the fully connected layer, and the resolution category recognition result adopts the binary classification one-hot encoding result.
[0128] Specifically, a low-resolution sample image A1 is input into a camera intrinsic parameter representation network C1, which outputs camera intrinsic parameter features B1 of the low-resolution sample image A1. The camera intrinsic parameter features B1 of the low-resolution sample image A1 are then input into a classification network C2, which outputs the resolution category recognition result T1 of the low-resolution sample image A1. Similarly, a high-resolution sample image A2 is input into a camera intrinsic parameter representation network C1, which outputs camera intrinsic parameter features B2 of the high-resolution sample image A2. The camera intrinsic parameter features B2 of the high-resolution sample image A2 are then input into a classification network C2, which outputs the resolution category recognition result T2 of the high-resolution sample image A2.
[0129] S203, based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category, adjust the parameters of the camera intrinsic parameter representation network and the classification network until the first training termination condition is met.
[0130] Based on the differences between the resolution category recognition results of low-resolution sample images and the corresponding true resolution category, and the differences between the resolution category recognition results of high-resolution sample images and the corresponding true resolution category, a first loss function is constructed. The value of the first loss function is backpropagated to adjust the parameters of the classification network and the camera intrinsic parameter representation network.
[0131] In one example, the first training termination condition is that the value of the first loss function is less than or equal to a first threshold. If the value of the first loss function is greater than the first threshold, the parameters of the classification network and the camera intrinsic parameter representation network are adjusted according to the value of the first loss function. After parameter adjustment, iterative training continues until the value of the first loss function is less than or equal to the first threshold, at which point training ends.
[0132] In another example, the first training termination condition is that the number of training iterations reaches a first preset number. If the number of training iterations has not reached the first preset number, the parameters of the classification network and the camera intrinsic parameter representation network are adjusted according to the value of the first loss function. After parameter adjustment, iterative training continues until the number of training iterations reaches the first preset number, at which point training ends.
[0133] In the above embodiments, the camera intrinsic parameter representation network trained can extract distinctive camera intrinsic parameter features from images from imaging devices with different resolutions. Using the camera intrinsic parameter features extracted by the camera intrinsic parameter representation network to assist in image super-resolution is beneficial to improving the super-resolution effect of the image.
[0134] In one possible implementation, image features of the image to be processed are extracted, and the camera intrinsic features of the image to be processed are fused with the image features to obtain a super-resolution image corresponding to the image to be processed. Specifically, this can be done by: using a first perceptual network to extract image features of the image to be processed, and fusing the camera intrinsic features with the image features to obtain a super-resolution image corresponding to the image to be processed; wherein, the first perceptual network is trained based on the difference between the first super-resolution image and the second sample image, the first super-resolution image is obtained by increasing the resolution of the third sample image, the second sample image and the third sample image have the same content, and the resolution of the second sample image is higher than the resolution of the third sample image.
[0135] The first perception network is a network used for image super-resolution. The image to be processed and its camera intrinsic features are input into the first perception network, and the first perception network outputs the super-resolution image corresponding to the image to be processed.
[0136] A second sample image and a third sample image constitute a set of sample images. In each set of sample images, the second sample image and the third sample image have the same content, but the resolution of the second sample image is higher than that of the third sample image. Optionally, the second sample image and the third sample image are images acquired by a high-resolution imaging device and a low-resolution imaging device, respectively, for the same target, such that the difference between the second sample image and the third sample image is only due to the resolution difference caused by the different imaging devices.
[0137] The first perceptual network can be trained through supervised learning, specifically based on the difference between a first super-resolution image and a second sample image. The first resolution image can be understood as the high-resolution image generated during network training, and the second sample image as the original high-resolution image. The training objective is to make the generated high-resolution image closer to the original high-resolution image. After the first perceptual network is trained, it is used to improve the resolution of the image to be processed, thus obtaining the corresponding super-resolution image.
[0138] In the above embodiments, the first perceptual network trained based on the difference between the first super-resolution image obtained by improving the resolution of the third sample image and the second sample image can be used to perform super-resolution on the image to obtain the corresponding super-resolution image, which is beneficial to quickly improve the image resolution.
[0139] Please see Figure 4 , Figure 4 This is a flowchart illustrating a training method for a first perceptual network provided in an embodiment of this application. The training method includes the following steps S401 to S403.
[0140] S401, acquire the camera intrinsic features of the second sample image, the third sample image, and the third sample image.
[0141] The camera intrinsic features of the third sample image can be extracted using the camera intrinsic representation network described in the previous embodiment. The third sample image is input into the camera intrinsic representation network, which then outputs the camera intrinsic features of the third sample image.
[0142] S402, through the first perception network, extract the image features of the third sample image, and fuse the camera intrinsic feature of the third sample image with the image features to obtain the first super-resolution image.
[0143] In one possible implementation, image features of a third sample image are extracted, and the camera intrinsic features of the third sample image are fused with the image features to obtain a first super-resolution image. Specifically, this may include: performing convolution and downsampling on the third sample image to obtain a first feature; downsampling the first feature to obtain a second feature; downsampling the second feature to obtain a third feature; performing fully connected and reconstructed processing on the camera intrinsic features of the third sample image to obtain a first convolution kernel; convolving the third feature with the first convolution kernel to obtain a fourth feature; superimposing the fourth feature with the third feature to obtain a fifth feature; upsampling the fifth feature to obtain a sixth feature; fusing the sixth feature with the second feature and then upsampling to obtain a seventh feature; fusing the seventh feature with the first feature and then upsampling and convolving to obtain an eighth feature; and superimposing the eighth feature with the third sample image to obtain the first super-resolution image.
[0144] Among them, the first feature, the second feature, and the third feature can be understood as image features of different scales extracted from the third sample image. The fourth feature and the fifth feature can be understood as preliminary fusion features obtained by incorporating the camera intrinsic features of the third sample image into the image features of the third sample image. The sixth feature, the seventh feature, and the eighth feature can be understood as further fusion features obtained by multi-scale fusion of the preliminary fusion features with the image features of the third sample image.
[0145] In the above embodiments, by fully fusing the camera intrinsic features of the low-resolution image with the image features of the low-resolution image, the fused features can contain richer information, which is beneficial to improving the resolution of the low-resolution image and obtaining the corresponding super-resolution image.
[0146] Please see Figure 5 , Figure 5 This is a training schematic diagram of a first perceptual network provided in an embodiment of this application. Wherein, Y1 represents the second sample image; Y2 represents the third sample image; C1 represents the camera intrinsic parameter representation network; R2 represents the camera intrinsic parameter features of the third sample image Y2; Z1 represents the first super-resolution image; D1 represents the first perceptual network; 1 represents a 3×3 convolutional layer; 2 represents a 5×5 convolutional layer; 3 represents a downsampling module; 4 represents a perceptual module; 5 represents an upsampling module; and 6 represents a feature addition layer. The first perceptual network D1 is a U-shaped network, including the following structure: a 3×3 convolutional layer 1, a 5×5 convolutional layer 2, a downsampling module 3, a perceptual module 4, an upsampling module 5, and a feature addition layer 6, wherein there are three 3×3 convolutional layers 1, three downsampling modules 3, and three upsampling modules.
[0147] Please see Figure 6 , Figure 6 This is a schematic diagram of a downsampling module, a perception module, and an upsampling module provided in an embodiment of this application. In the diagram, Fin represents the features input to the downsampling module 3; R represents the camera intrinsic features; Fout represents the features output by the upsampling module 5; 7 represents the residual module; 8 represents the max pooling layer; 9 represents the fully connected layer; 10 represents the reconstruction layer; a represents the convolutional kernel; 11 represents the depthwise convolutional layer; and 12 represents the sub-pixel convolutional layer. The downsampling module 3 includes the residual module 7 and the max pooling layer 8; the perception module 4 includes the fully connected layer 9, the reconstruction layer 10, the depthwise convolutional layer 11, the feature addition layer 6, and the 3×3 convolutional layer 1; and the upsampling module 5 includes the residual module 7 and the sub-pixel convolutional layer 12.
[0148] Specifically, the third sample image Y2 is input into the first perceptual network D1, and processed sequentially through the first 3×3 convolutional layer 1, the 5×5 convolutional layer 2, and the first downsampling module 3 to obtain the first feature; the first feature is downsampled using the second downsampling module 3 to obtain the second feature; and the second feature is downsampled using the third downsampling module 3 to obtain the third feature.
[0149] The third feature and the camera intrinsic feature R2 of the third sample image Y2 are input into the perception module 4. The perception module 4 integrates the camera intrinsic feature R2 into the third feature. Specifically, the processing includes: using a fully connected layer 9 and a reconstruction layer 10 to convert the camera intrinsic feature R2 of the third sample image Y2 into a convolution kernel; using this convolution kernel and a depth convolution layer 11 to convolve the third feature to obtain the fourth feature; and using a feature addition layer 6 to superimpose the fourth feature with the third feature to obtain the fifth feature.
[0150] The fifth feature is upsampled using the first upsampling module 5 to obtain the sixth feature. The feature fused with the second feature using the second upsampling module 5 is then upsampled to obtain the seventh feature. The feature fused with the first feature is then processed sequentially through the third upsampling module 5, the second 3×3 convolutional layer 1, and the third 3×3 convolutional layer 1 to obtain the eighth feature. The eighth feature is then superimposed on the third sample image Y2 using the feature addition layer 6 to obtain the first super-resolution image Z1.
[0151] It should be noted that the structure of the first-sensory network is not limited to Figure 5 The structure shown, such as the number and size of convolutional layers, and the number of downsampling and upsampling modules, can be configured according to actual needs. Furthermore, the structures of the downsampling, sensing, and upsampling modules are not limited to specific requirements. Figure 6 The structures shown, such as residual modules, max pooling layers, depthwise convolutional layers, and subpixel convolutional layers, can be replaced with other structures that have the same or similar functions.
[0152] S403, based on the difference between the first super-resolution image and the second sample image, adjust the parameters of the first perceptual network until the second training termination condition is met.
[0153] Based on the difference between the first super-resolution image Z1 and the second sample image Y1, a second loss function is constructed. The value of the second loss function is backpropagated to adjust the parameters of the first perceptual network D1. Optionally, the second loss function is determined based on the minimum absolute value deviation (1-norm loss) and structural similarity index (SSIM) between the first super-resolution image Z1 and the second sample image Y1.
[0154] In one example, the second training termination condition is that the value of the second loss function is less than or equal to the second threshold. If the value of the second loss function is greater than the second threshold, the parameters of the first perceptual network are adjusted according to the value of the second loss function. After parameter adjustment, iterative training continues until the value of the second loss function is less than or equal to the second threshold, at which point training ends.
[0155] In another example, the second training termination condition is that the number of training iterations reaches a second preset number. If the number of training iterations has not reached the second preset number, the parameters of the first perceptual network are adjusted according to the value of the second loss function. After parameter adjustment, iterative training continues until the number of training iterations reaches the second preset number, at which point training ends.
[0156] In the above embodiments, the first perceptual network trained can perceive the camera intrinsic features of low-resolution images and integrate the camera intrinsic features into the image features to generate corresponding super-resolution images. This is beneficial for realizing the spatial domain transfer from the super-resolution image to the original low-resolution image, thereby helping to solve the inter-domain drift problem and improve the super-resolution effect of the image.
[0157] In one possible implementation, the step of extracting image features from the image to be processed and fusing the camera intrinsic features of the image to be processed with the image features to obtain a super-resolution image corresponding to the image to be processed can specifically be: using a second perceptual network to extract image features from the image to be processed, fusing the camera intrinsic features with the image features to obtain a super-resolution image corresponding to the image to be processed; wherein, the second perceptual network is trained based on the difference between the first super-resolution image and the second sample image, and the difference between the degraded resolution image and the third sample image, the first super-resolution image is obtained by increasing the resolution of the third sample image, the degraded resolution image is obtained by decreasing the resolution of the first super-resolution image, the second sample image and the third sample image have the same content, and the resolution of the second sample image is higher than the resolution of the third sample image.
[0158] The second sensing network is a network used for image super-resolution. The image to be processed and its camera intrinsic features are input into the second sensing network, and the second sensing network outputs the super-resolution image corresponding to the image to be processed.
[0159] The second perceptual network can be trained through supervised learning, specifically based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image. Here, the first resolution image can be understood as the high-resolution image generated during network training, the second sample image as the original high-resolution image, the degraded resolution image as the low-resolution image generated during network training, and the third sample image as the original low-resolution image. The training objective is to make the generated high-resolution image close to the original high-resolution image, and to make the generated low-resolution image close to the original low-resolution image. After the second perceptual network is trained, it is used to improve the resolution of the image to be processed, thus obtaining the corresponding super-resolution image.
[0160] In the above embodiments, the second perceptual network, trained based on the difference between the first super-resolution image (obtained by increasing the resolution of the third sample image) and the second sample image, and the difference between the degraded resolution image (obtained by decreasing the resolution of the first super-resolution image) and the third sample image, can be used to perform super-resolution on images to obtain corresponding super-resolution images, which is beneficial for quickly improving image resolution. Furthermore, since the training process of the second perceptual network considers the difference between the degraded resolution image and the third sample image, this helps to limit the search space of network parameters during training and optimize the training effect of the network model.
[0161] Please see Figure 7 , Figure 7 This is a flowchart illustrating a training method for a second perceptual network provided in an embodiment of this application. The training method includes the following steps S701 to S704.
[0162] S701, acquire the second sample image, the third sample image, the camera intrinsic feature of the second sample image, and the camera intrinsic feature of the third sample image.
[0163] The camera intrinsic features of the second and third sample images can be extracted using the camera intrinsic representation network described in the previous embodiment. The second sample image is input into the camera intrinsic representation network, which outputs the camera intrinsic features of the second sample image. Similarly, the third sample image is input into the camera intrinsic representation network, which outputs the camera intrinsic features of the third sample image.
[0164] S702 extracts image features from the third sample image through the second perception network, and fuses the camera intrinsic features of the third sample image with the image features to obtain the first super-resolution image.
[0165] For example, the specific processing procedure of the second sensing network can be referred to the description of the first sensing network in the previous embodiment, and will not be repeated here. The difference between the second sensing network and the first sensing network is that a recurrent sensing network is added during the training of the second sensing network.
[0166] A high-resolution image can be degraded into different low-resolution images through various degradation methods, resulting in a many-to-one ill-posed problem. In network model training, this ill-posed problem forces the model to search a large search space for a low-to-high resolution mapping that meets the training objective, making network model training difficult. Therefore, a recurrent perceptual network is incorporated into the training of the second perceptual network. This helps to limit the search space of network parameters during training, resolve the ill-posed problem, and optimize the network model training effect.
[0167] S703 extracts image features from the first super-resolution image through a recurrent sensing network, and fuses the camera intrinsic features of the second sample image with the image features of the first super-resolution image to obtain a degraded resolution image.
[0168] In one possible implementation, image features of a first super-resolution image are extracted, and the camera intrinsic features of a second sample image are fused with the image features of the first super-resolution image to obtain a degraded resolution image. Specifically, this may include: performing convolution and downsampling on the first super-resolution image to obtain a ninth feature; performing fully connected and reconstructed processing on the camera intrinsic features of the second sample image to obtain a second convolution kernel; using the second convolution kernel to convolve the ninth feature to obtain a tenth feature; superimposing the tenth feature with the ninth feature to obtain an eleventh feature; upsampling and convolution on the eleventh feature to obtain a twelfth feature; and superimposing the twelfth feature with the first super-resolution image to obtain a degraded resolution image.
[0169] Among them, the ninth feature can be understood as the image feature extracted from the first super-resolution image, the tenth feature can be understood as the preliminary fusion feature obtained by integrating the camera intrinsic feature of the second sample image into the image feature of the first super-resolution image, and the eleventh and twelfth features can be understood as the further fusion features obtained by fusing the preliminary fusion feature with the image feature of the first super-resolution image.
[0170] In the above embodiments, by fully fusing the camera intrinsic features of the high-resolution image with the image features of the super-resolution image, the fused features can contain richer information, which is beneficial for better reducing the resolution of the super-resolution image and obtaining the corresponding degraded resolution image.
[0171] Please see Figure 8 , Figure 8This is a training schematic diagram of a second perceptual network provided in an embodiment of this application. R1 represents the camera intrinsic features of the second sample image Y1; Z2 represents the degraded resolution image; D2 represents the second perceptual network; and D3 represents the recurrent perceptual network. The structure of the second perceptual network D2 is the same as that of the first perceptual network D1, and will not be described again here. The recurrent perceptual network D3 includes the following structure: a 3×3 convolutional layer 1, a 5×5 convolutional layer 2, a downsampling module 3, a perceptual module 4, an upsampling module 5, and a feature summing layer 6. There are three 3×3 convolutional layers 1, and two downsampling modules 3 and two upsampling modules 5.
[0172] Specifically, the first super-resolution image Z1 is input into the recurrent perceptron D3, and processed sequentially through the first 3×3 convolutional layer 1, the 5×5 convolutional layer 2, and two downsampling modules 3 to obtain the ninth feature. The ninth feature and the camera intrinsic feature R1 of the second sample image Y1 are input into the perceptron module 4. The perceptron module 4 integrates the camera intrinsic feature R1 into the ninth feature, specifically including the following processing: the camera intrinsic feature R1 of the second sample image Y1 is transformed into a convolutional kernel using a fully connected layer 9 and a reconstruction layer 10; the ninth feature is convolved with this convolutional kernel and a depthwise convolutional layer 11 to obtain the tenth feature; the tenth feature is superimposed with the ninth feature using a feature addition layer 6 to obtain the eleventh feature. The eleventh feature is processed sequentially through two upsampling modules 5, the second first 3×3 convolutional layer 1, and the third 3×3 convolutional layer 1 to obtain the twelfth feature. The twelfth feature is superimposed with the first super-resolution image Z1 using the feature addition layer 6 to obtain the degraded resolution image Z2.
[0173] It should be noted that the structure of recurrent sensing networks is not limited to... Figure 8 The structure shown, such as the number and size of convolutional layers, and the number of downsampling and upsampling modules, can be configured according to actual needs. Furthermore, the structures of the downsampling, sensing, and upsampling modules are not limited to specific requirements. Figure 6 The structures shown, such as residual modules, max pooling layers, depthwise convolutional layers, and subpixel convolutional layers, can be replaced with other structures that have the same or similar functions.
[0174] S704, based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image, adjusts the parameters of the second perceptual network and the recurrent perceptual network until the third training termination condition is met.
[0175] Based on the differences between the first super-resolution image Z1 and the second sample image Y1, and the differences between the degraded resolution image Z2 and the third sample image Y2, a third loss function is constructed. The value of the third loss function is backpropagated to adjust the parameters of the recurrent perceptron D3 and the second perceptron D2. Optionally, the third loss function is determined based on the minimum absolute value deviation (1-norm loss) and structural similarity index (SSIM) between the first super-resolution image Z1 and the second sample image Y1, and the minimum absolute value deviation (1-norm loss) between the degraded resolution image Z2 and the third sample image Y2.
[0176] In one example, the third training termination condition is that the value of the third loss function is less than or equal to the third threshold. If the value of the third loss function is greater than the third threshold, the parameters of the recurrent perceptron and the second perceptron are adjusted according to the value of the third loss function. After parameter adjustment, iterative training continues until the value of the third loss function is less than or equal to the third threshold, at which point training ends.
[0177] In another example, the third training termination condition is that the number of iterations reaches a third preset number. If the number of iterations has not reached the third preset number, the parameters of the recurrent perceptron and the second perceptron are adjusted according to the value of the third loss function. After parameter adjustment, iterative training continues until the number of iterations reaches the third preset number, at which point training ends.
[0178] In the above embodiments, the trained second perceptual network can perceive the camera intrinsic features of low-resolution images and integrate these features into the image features to generate a corresponding super-resolution image. This facilitates spatial domain transfer from the super-resolution image to the original low-resolution image, thereby helping to solve the inter-domain drift problem and improve the super-resolution effect. Furthermore, a recurrent perceptual network is incorporated into the training of the second perceptual network. This network perceives the camera intrinsic features of high-resolution images and integrates them into the image features of the super-resolution image, generating a degraded low-resolution image to aid training. This helps to limit the search space of network parameters during training, solve the ill-posed problem, and optimize the training effect of the network model.
[0179] It should be noted that the recurrent perceptron is only used during the training phase of the second perceptron, and is not required during the inference phase. For example, during the inference phase of the second perceptron, a low-resolution image acquired by a low-resolution imaging device and the camera intrinsic features of that low-resolution image are input into the second perceptron, and the second perceptron outputs a super-resolution image corresponding to that low-resolution image.
[0180] In an exemplary application scenario, the above image processing method is used to optimize the quality of thermal images acquired by low-resolution thermal imaging sensors. By performing super-resolution on low-resolution thermal images to improve thermal image quality, the performance of downstream tasks such as thermal imaging face recognition and target detection is improved, while reducing the cost of acquiring high-quality thermal images.
[0181] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0182] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0183] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of an image processing apparatus 900 provided in an embodiment of this application. The image processing apparatus 900 includes: an acquisition unit 901, a first processing unit 902, and a second processing unit 903, wherein:
[0184] Acquisition unit 901 is used to acquire the image to be processed;
[0185] The first processing unit 902 is used to extract camera intrinsic parameter features of the image to be processed. The camera intrinsic parameter features are used to characterize the resolution performance of the imaging device that acquires the image to be processed.
[0186] The second processing unit 903 is used to extract image features of the image to be processed, fuse the camera intrinsic parameter features of the image to be processed with the image features, and obtain a super-resolution image corresponding to the image to be processed, wherein the resolution of the super-resolution image is higher than the resolution of the image to be processed.
[0187] In some possible implementations, the first processing unit 902 is specifically used to: extract camera intrinsic parameter features of the image to be processed using a camera intrinsic parameter representation network; wherein the camera intrinsic parameter representation network is trained based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category.
[0188] For some possible implementations, please refer to Figure 10The device 900 further includes a first training unit 904 for training a camera intrinsic parameter representation network. The first training unit 904 is specifically used for: acquiring a first sample image and its true resolution category, wherein the first sample image includes images acquired by two imaging devices with different resolutions; extracting camera intrinsic parameter features of the first sample image through the camera intrinsic parameter representation network, and performing classification and recognition based on the camera intrinsic parameter features of the first sample image through a classification network to obtain the resolution category recognition result of the first sample image; and adjusting the parameters of the camera intrinsic parameter representation network and the classification network based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category, until the first training termination condition is met.
[0189] In some possible implementations, the second processing unit 903 is specifically used to: extract image features of the image to be processed using a first perceptual network, fuse camera intrinsic features with image features, and obtain a super-resolution image corresponding to the image to be processed; wherein, the first perceptual network is trained based on the difference between the first super-resolution image and the second sample image, the first super-resolution image is obtained by increasing the resolution of the third sample image, the second sample image and the third sample image have the same content, and the resolution of the second sample image is higher than the resolution of the third sample image.
[0190] For some possible implementations, please refer to Figure 10 The device 900 further includes a second training unit 905 for training a first perceptual network. The second training unit 905 is specifically used to: acquire a second sample image, a third sample image, and camera intrinsic features of the third sample image; extract image features of the third sample image through the first perceptual network, fuse the camera intrinsic features of the third sample image with the image features to obtain a first super-resolution image; and adjust the parameters of the first perceptual network based on the difference between the first super-resolution image and the second sample image until the second training termination condition is met.
[0191] In some possible implementations, the second processing unit 903 is specifically used to: extract image features of the image to be processed using a second perceptual network, fuse camera intrinsic features with image features, and obtain a super-resolution image corresponding to the image to be processed; wherein, the second perceptual network is trained based on the difference between the first super-resolution image and the second sample image, and the difference between the degraded resolution image and the third sample image, the first super-resolution image is obtained by increasing the resolution of the third sample image, the degraded resolution image is obtained by decreasing the resolution of the first super-resolution image, the second sample image and the third sample image have the same content, and the resolution of the second sample image is higher than the resolution of the third sample image.
[0192] For some possible implementations, please refer to Figure 10 The device 900 further includes a third training unit 906 for training a second perceptual network. Specifically, the third training unit 906 is used to: acquire a second sample image, a third sample image, camera intrinsic features of the second sample image, and camera intrinsic features of the third sample image; extract image features of the third sample image through the second perceptual network, and fuse the camera intrinsic features of the third sample image with the image features to obtain a first super-resolution image; extract image features of the first super-resolution image through a recurrent perceptual network, and fuse the camera intrinsic features of the second sample image with the image features of the first super-resolution image to obtain a degraded resolution image; and adjust the parameters of the second perceptual network and the recurrent perceptual network based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image, until a third training termination condition is met.
[0193] In some possible implementations, when the second training unit 905 or the third training unit 906 extracts image features from the third sample image and fuses the camera intrinsic features of the third sample image with the image features to obtain the first super-resolution image, it specifically performs the following: convolution and downsampling on the third sample image to obtain a first feature; downsampling on the first feature to obtain a second feature; downsampling on the second feature to obtain a third feature; performing fully connected and reconstructed operations on the camera intrinsic features of the third sample image to obtain a first convolution kernel; convolving the third feature with the first convolution kernel to obtain a fourth feature; superimposing the fourth feature with the third feature to obtain a fifth feature; upsampling the fifth feature to obtain a sixth feature; fusing the sixth feature with the second feature and then upsampling to obtain a seventh feature; fusing the seventh feature with the first feature and then upsampling and convolving to obtain an eighth feature; and superimposing the eighth feature with the third sample image to obtain the first super-resolution image.
[0194] In some possible implementations, when the third training unit 906 extracts image features from the first super-resolution image and fuses the camera intrinsic features of the second sample image with the image features of the first super-resolution image to obtain a degraded resolution image, it specifically performs the following: convolution and downsampling on the first super-resolution image to obtain a ninth feature; fully connecting and reconstructing the camera intrinsic features of the second sample image to obtain a second convolution kernel; convolving the ninth feature with the second convolution kernel to obtain a tenth feature; superimposing the tenth feature with the ninth feature to obtain an eleventh feature; upsampling and convolution on the eleventh feature to obtain a twelfth feature; and superimposing the twelfth feature with the first super-resolution image to obtain a degraded resolution image.
[0195] For specific limitations regarding the image processing apparatus, please refer to the limitations on the image processing method above, which will not be repeated here. Each unit in the aforementioned image processing apparatus can be implemented entirely or partially through software, hardware, or a combination thereof. Each of these units can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each of the above units.
[0196] Please see Figure 11 , Figure 11 A hardware structure diagram of an image processing apparatus provided in this application embodiment may include:
[0197] The processor 1101, memory 1102, and transceiver 1103 are connected via bus 1104. The memory 1102 is used to store instructions, and the processor 1101 is used to execute the instructions stored in the memory 1102 to implement the steps in the above method.
[0198] The processor 1101 executes the instructions stored in the memory 1102 to control the transceiver 1103 to receive and transmit signals, thus completing the steps in the above method. The memory 1102 can be integrated into the processor 1101 or it can be set separately from the processor 1101.
[0199] As one implementation method, the transceiver 1103 can be implemented using a transceiver circuit or a dedicated transceiver chip. The processor 1101 can be implemented using a dedicated processing chip, processing circuit, processor, or general-purpose chip.
[0200] As another implementation method, the image processing apparatus provided in this application embodiment can be implemented using a general-purpose computer. The program code that implements the functions of processor 1101 and transceiver 1103 is stored in memory 1102, and the general-purpose processor implements the functions of processor 1101 and transceiver 1103 by executing the code in memory 1102.
[0201] For the concepts, explanations, detailed descriptions, and other steps related to the technical solutions provided in the embodiments of this application, please refer to the description of the method steps performed by the device in the foregoing method or other embodiments, which will not be repeated here.
[0202] This application also provides an electronic device, including a processor and a memory, wherein the memory is used to store computer program code, the computer program code including computer instructions, and the electronic device performs the method as described in the above method embodiments when the processor executes the computer instructions.
[0203] This application also provides an electronic device, including: a processor, a transmitting device, an input device, an output device, and a memory. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device performs the method as described in the above method embodiments.
[0204] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the method as described in the above method embodiments.
[0205] This application also provides a computer program product, which includes a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the above method embodiments.
[0206] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0207] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0208] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Obtain the image to be processed; The camera intrinsic parameter features of the image to be processed are extracted, and the camera intrinsic parameter features are used to characterize the resolution performance of the imaging device that acquires the image to be processed; Using a second sensing network, image features of the image to be processed are extracted, and the camera intrinsic features of the image to be processed are fused with the image features to obtain a super-resolution image corresponding to the image to be processed, wherein the resolution of the super-resolution image corresponding to the image to be processed is higher than the resolution of the image to be processed. The second perceptual network is trained based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image. The first super-resolution image is obtained by increasing the resolution of the third sample image, and the degraded resolution image is obtained by decreasing the resolution of the first super-resolution image. The second sample image has the same content as the third sample image, and the resolution of the second sample image is higher than that of the third sample image.
2. The method according to claim 1, characterized in that, The extraction of camera intrinsic feature data from the image to be processed includes: The camera intrinsic features of the image to be processed are extracted using a camera intrinsic representation network; The camera intrinsic parameter representation network is trained based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category.
3. The method according to claim 2, characterized in that, The training methods for the camera intrinsic parameter representation network include: Obtain a first sample image and its true resolution category, wherein the first sample image includes images acquired by two imaging devices with different resolutions; The camera intrinsic feature of the first sample image is extracted by the camera intrinsic feature representation network, and the classification network is used to classify and identify the first sample image based on the camera intrinsic feature of the first sample image to obtain the resolution category identification result of the first sample image. Based on the difference between the resolution category recognition result of the first sample image and the corresponding true resolution category, the parameters of the camera intrinsic parameter representation network and the classification network are adjusted until the first training termination condition is met.
4. The method according to any one of claims 1 to 3, characterized in that, The training method for the second perceptual network includes: Acquire the second sample image, the third sample image, the camera intrinsic features of the second sample image, and the camera intrinsic features of the third sample image; The second sensing network is used to extract the image features of the third sample image, and the camera intrinsic features of the third sample image are fused with the image features to obtain the first super-resolution image. By using a recurrent sensing network, the image features of the first super-resolution image are extracted, and the camera intrinsic features of the second sample image are fused with the image features of the first super-resolution image to obtain a degraded resolution image. Based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image, the parameters of the second perceptual network and the recurrent perceptual network are adjusted until the third training termination condition is met.
5. The method according to claim 4, characterized in that, The step of extracting image features from the third sample image and fusing the camera intrinsic features of the third sample image with the image features to obtain a first super-resolution image includes: The third sample image is convolved and downsampled to obtain a first feature, the first feature is downsampled to obtain a second feature, and the second feature is downsampled to obtain a third feature; The camera intrinsic features of the third sample image are fully connected and reconstructed to obtain the first convolution kernel. The first convolution kernel is used to convolve the third feature to obtain the fourth feature. The fourth feature is superimposed with the third feature to obtain the fifth feature. The fifth feature is upsampled to obtain the sixth feature. The sixth feature is fused with the second feature and then upsampled to obtain the seventh feature. The seventh feature is fused with the first feature and then upsampled and convolved to obtain the eighth feature. The eighth feature is superimposed on the third sample image to obtain the first super-resolution image.
6. The method according to claim 4, characterized in that, The step of extracting image features from the first super-resolution image and fusing the camera intrinsic features of the second sample image with the image features of the first super-resolution image to obtain a degraded resolution image includes: The ninth feature is obtained by convolution and downsampling the first super-resolution image; The camera intrinsic features of the second sample image are fully connected and reconstructed to obtain the second convolution kernel. The second convolution kernel is used to convolve the ninth feature to obtain the tenth feature. The tenth feature is superimposed with the ninth feature to obtain the eleventh feature. The eleventh feature is obtained by upsampling and convolution. The twelfth feature is superimposed on the first super-resolution image to obtain a degraded resolution image.
7. An image processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire the image to be processed; The first processing unit is used to extract camera intrinsic parameter features of the image to be processed, wherein the camera intrinsic parameter features are used to characterize the resolution performance of the imaging device that acquires the image to be processed; The second processing unit is used to extract image features of the image to be processed using the second sensing network, and fuse the camera intrinsic parameter features of the image to be processed with the image features to obtain a super-resolution image corresponding to the image to be processed, wherein the resolution of the super-resolution image is higher than the resolution of the image to be processed. The second perceptual network is trained based on the differences between the first super-resolution image and the second sample image, and the differences between the degraded resolution image and the third sample image. The first super-resolution image is obtained by increasing the resolution of the third sample image, and the degraded resolution image is obtained by decreasing the resolution of the first super-resolution image. The second sample image has the same content as the third sample image, and the resolution of the second sample image is higher than that of the third sample image.
8. An electronic device, characterized in that, The method includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the method as described in any one of claims 1 to 6 by executing the instructions stored in the memory.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image quality optimization method and system
CN113689342A