Image Processing Method, Apparatus, Storage Medium, and Electronic Device

The test paper image is processed through the noise removal model based on the HRNet feature matching module, which solves the problems of insufficient noise removal accuracy and loss of image content in the prior art, and realizes high-quality noise removal test paper images, improving the text detection and recognition accuracy and recording efficiency.

CN114841882BActive Publication Date: 2025-06-20BEIJING DINGSHIXINGJIAOYU CONSULTATION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210474094.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-06-20
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

When removing noise and handwriting traces in the test paper images, it is difficult to achieve subpixel-level accuracy, resulting in unsmoothing image edges, and the results of semantic segmentation cannot retain the grayscale color level, resulting in loss of image content and cannot meet the needs of subsequent applications.

Method used

A noise removal model based on the HRNet feature matching module is used to perform supervised training, and the test paper images are processed, noise is removed and image high-frequency information is retained to achieve higher quality noise removal test paper images.

Benefits of technology

By this method, while ensuring the smoothness of the image edge and the color level of the grayscale image can be retained, noise and handwriting traces in the test paper image can be removed, and the text detection and recognition accuracy and recording efficiency can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841882B_ABST
    Figure CN114841882B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method, apparatus, storage medium, and electronic device, belonging to the field of image processing. The method includes: obtaining a test paper image, where the test paper image includes image noise; inputting the test paper image into a pre-trained noise removal model to obtain a target test paper image with the noise removed; where the noise removal model is obtained through supervised training based on an HRNet feature matching module. It can ensure that the obtained test paper image can retain the high-frequency information of the image and obtain a denoised test paper image with better image quality, which is convenient for subsequent use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and specifically, to an image processing method, apparatus, storage medium, and electronic device. Background Art

[0002] In the application of question bank typesetting, the input test paper or teaching aid material images may contain a large number of handwritten characters, images, and marking traces. Moreover, the photographed images will be interfered by factors such as background and lighting. The above noises will all affect the accuracy of the typesetting algorithm. Restoring and reconstructing the photographed images, removing the noises and handwritten traces can improve the accuracy of text detection and recognition, thereby improving the typesetting efficiency.

[0003] In the related art, part of the noise can be removed by semantic segmentation. However, the printed text and images in the test paper require sub-pixel accuracy in the edge area to ensure the smoothness of the image edge. The semantic segmentation model cannot achieve sub-pixel accuracy. Moreover, the illustrations in the test paper or teaching aid materials may contain large areas of gray-scale regions, and the result of semantic segmentation can only be converted into a binary image, unable to retain the gray-scale levels of the gray-scale image, resulting in loss of image content. Finally, the quality of the image after removing the noise cannot meet the application requirements of the subsequent test paper images. Summary of the Invention

[0004] To solve the problems existing in the related art, the present disclosure provides an image processing method, apparatus, storage medium, and electronic device.

[0005] According to a first aspect of the present disclosure, there is provided an image processing method, the method comprising:

[0006] Obtain a test paper image, the test paper image including image noise;

[0007] Input the test paper image into a pre-trained noise removal model to obtain a target test paper image with the noise removed;

[0008] Wherein, the noise removal model is obtained by supervised training based on an HRNet feature matching module.

[0009] Optionally, the training of the noise removal model includes:

[0010] Add image noise to a first test paper image without image noise to obtain a training test paper image;

[0011] Input the training test paper image into the noise removal model to obtain a first generated image;

[0012] Input the first generated image and the first test paper image into the HRNet feature matching module for feature matching to determine a first loss value;

[0013] Update the parameters of the noise removal model according to the first loss value.

[0014] Optionally, the noise removal model includes a global generation module and a local enhancement network, and the local enhancement network includes a first local enhancement module and a second local enhancement module.

[0015] The step of inputting the training test paper image into the noise removal model to obtain a first generated image includes:

[0016] Doubly downsample the training test paper image and then input it into the global generation module to obtain a first feature map; and

[0017] Input the training test paper image into the first local enhancement module to obtain a second feature map;

[0018] Add the first feature map and the second feature map, and then input the result into the second local enhancement module to obtain the first generated image.

[0019] Optionally, the method further includes:

[0020] Input the training test paper image into the global generation module to train the global generation module; and

[0021] When it is determined that a preset condition is satisfied, input the training test paper image into the global generation module and the local enhancement network to train the global generation module and the local enhancement network simultaneously; and

[0022] When it is determined that the training is completed, extract the local enhancement network as the trained noise removal model.

[0023] Optionally, the method further includes:

[0024] Input the first generated image into a first discriminator to obtain a first discrimination result;

[0025] Doubly downsample the first generated image and then input it into a second discriminator to obtain a second discrimination result;

[0026] Quadruply downsample the first generated image and then input it into a third discriminator to obtain a third discrimination result;

[0027] Calculate a second loss value according to the first test paper image, the first discrimination result, the second discrimination result, and the third discrimination result;

[0028] Update the parameters of the noise removal model according to the second loss value.

[0029] Optionally, the HRNet feature matching module is a network with parallel connections from high to low resolution.

[0030] According to a second aspect of the present disclosure, there is provided an image processing apparatus, the apparatus comprising:

[0031] An acquisition module, configured to acquire a test paper image, the test paper image including image noise;

[0032] A denoising module, configured to input the test paper image into a pre-trained noise removal model to obtain a target test paper image with noise removed;

[0033] Wherein, the noise removal model is obtained through supervised training based on the HRNet feature matching module.

[0034] Optionally, the apparatus includes:

[0035] A processing module, configured to add image noise to a first test paper image without image noise to obtain a training test paper image;

[0036] An input module, configured to input the training test paper image into the noise removal model to obtain a first generated image;

[0037] A matching module, configured to input the first generated image and the first test paper image into the HRNet feature matching module for feature matching to determine a first loss value;

[0038] An update module, configured to update the parameters of the noise removal model according to the first loss value.

[0039] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspects of the present disclosure are implemented.

[0040] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:

[0041] A memory, on which a computer program is stored;

[0042] A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of the first aspects of the present disclosure.

[0043] Through the above technical solutions, a noise removal model is obtained through supervised training based on the HRNet feature matching module, and the test paper image is processed by the noise removal model to obtain a test paper image with noise removed, which can ensure that the obtained test paper image can retain the high-frequency information of the image and obtain a denoised test paper image with better image quality, facilitating subsequent use.

[0044] Other features and advantages of the present disclosure will be described in detail in the following detailed implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following detailed implementation, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:

[0046] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment;

[0047] Figure 2 is a flowchart of a method for training a noise removal model shown according to an exemplary embodiment;

[0048] Figure 3 is a schematic diagram of a noise removal model shown according to an exemplary embodiment;

[0049] Figure 4 is a schematic diagram of the training of a noise removal model shown according to an exemplary embodiment;

[0050] Figure 5 is a block diagram of an image processing apparatus shown according to an exemplary embodiment;

[0051] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment;

[0052] Figure 7 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED IMPLEMENTATION

[0053] The following details the specific implementation of the present disclosure in conjunction with the drawings. It should be understood that the specific implementation described herein is only for explaining and understanding the present disclosure, and is not used to limit the present disclosure.

[0054] It should be understood that the steps recorded in the method implementation of the present disclosure can be executed in different orders and / or in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0055] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0056] It should be noted that concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.

[0057] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0058] To enable those skilled in the art to better understand the improvements of the technical solutions provided by the embodiments of this disclosure, an introduction is made to the related technologies:

[0059] In the related technologies, the method for denoising the test paper image obtained by shooting mainly involves semantic segmentation and handwritten / printed text line detection.

[0060] For semantic segmentation, usually, the semantic segmentation algorithm is used to divide the test paper image pixel by pixel into two categories. The printed part of the test paper, including text, symbols, images, etc., is divided into the first category, and other elements to be removed and the background are divided into the second category. Then, only the first category area is retained to complete the restoration of the test paper image.

[0061] However, the method of semantic segmentation has the following problems, which ultimately lead to poor quality of the obtained test paper image: The printed text and images in the test paper require sub-pixel accuracy in the edge area to ensure the smoothness of the image edge. The semantic segmentation model theoretically cannot reach sub-pixel accuracy. The illustrations in the test paper or teaching aids may contain large areas of gray regions, and the result of semantic segmentation can only be converted into a binary image, unable to retain the gray scale of the gray image, resulting in loss of image content.

[0062] For handwritten / printed text line detection, the text detection model is used to detect the handwritten text and printed text in the test paper respectively, and the handwritten text is erased and then the image is repaired.

[0063] This method also has the following problems: In addition to handwritten text, the handwritten traces in the test paper also contain a large number of other non-text elements, such as symbols, images, underlines, etc. Such handwritten traces cannot be accurately located using the text detection model, resulting in the retention of handwritten traces. The handwritten traces and noise may overlap with the text line area, such as teacher correction traces and underlines, circles, etc. during students' reading. For such overlapping areas, the algorithm based on text detection cannot accurately distinguish, making it difficult to erase such traces. It also cannot meet the requirements for the test paper image.

[0064] To improve the quality of the denoised image of the captured image and implement the functions of handwritten trace erasure, image denoising, and image enhancement, the present disclosure provides an image processing method, apparatus, storage medium, and electronic device.

[0065] Figure 1 FIG. 4 is a flowchart of an image processing method shown according to an exemplary embodiment. The execution subject of the method may be an electronic device with information processing capabilities such as a mobile phone, a computer, or a server. The method includes:

[0066] S101. Obtain a test paper image, where the test paper image includes image noise.

[0067] Among them, the image noise may be, for example, handwritten characters, marking traces, light noise, and so on. Taking the server as the execution subject of this method as an example, the test paper image may be sent to the server after being captured by the user's mobile phone so that the server can obtain the test paper image.

[0068] S102. Input the test paper image into a pre-trained noise removal model to obtain a target test paper image with noise removed.

[0069] Among them, the noise removal model is obtained through supervised training based on the HRNet feature matching module.

[0070] Optionally, the HRNet feature matching module is a network with parallel connections from high to low resolution. Those skilled in the art should know that HRNet is also known as HighResolution Net, which can maintain a high-resolution representation throughout the feature matching process. Starting with a high-resolution subnet as the first stage, high-to-low resolution subnets are added one by one to form more stages, and multi-resolution subnets are connected in parallel. Information in the parallel multi-resolution subnets is repeatedly exchanged throughout the process for repeated multi-scale fusion.

[0071] It can be understood that in the application scenario of question bank recording and typesetting, the input test paper or teaching aid material image may contain a large amount of noise such as handwritten characters, images, and marking traces. Moreover, the captured image will be interfered by factors such as the background and light. The above image noise will all affect the accuracy of the typesetting algorithm. Compared with the VGG model generally used in the related art for feature matching, using HRNet as the feature matching model for supervised training can better retain the high-frequency information of the image.

[0072] In the embodiments of the present disclosure, a noise removal model is obtained through supervised training based on an HRNet feature matching module, and the test paper image is processed by the noise removal model to obtain a noise-removed test paper image, which can ensure that the obtained test paper image can retain the high-frequency information of the image and obtain a denoised test paper image with better image quality, facilitating subsequent use.

[0073] Specifically, Figure 2 is a flowchart of a training method for a noise removal model shown according to an exemplary embodiment, as Figure 2 shown, the training of the noise removal model includes:

[0074] S201. Add image noise to a first test paper image without image noise to obtain a training test paper image.

[0075] Among them, the added image noise may include handwritten characters, marking traces, and illumination noise. The way to add image noise may include: after photographing a clean and tidy test paper image, adding handwritten content to the test paper image and then photographing it to obtain a test paper image with added image noise. Or, through a pre-trained noise addition model, add a noise image to the test paper image without noise to obtain a test paper image with added image noise. Specifically, the present disclosure does not limit the specific image noise addition method.

[0076] S202. Input the training test paper image into the noise removal model to obtain a first generated image.

[0077] Among them, the noise removal model may be a CGAN model. The data distribution in the real world is highly complex and presents a multi-modal distribution. In other words, for the probability distribution function of the data, different data subsets are distributed around different peaks. Mode collapse means that the data generated by the generative model cannot well cover the probability distribution that the data should have, but is concentrated on one or several peaks. For example, for the GAN algorithm for generating cat images, it is expected that the algorithm can output all breeds of cats or cats in various poses. When mode collapse occurs, the generator may only be able to generate one or several breeds of cats or only be able to generate several specific poses of cats.

[0078] Those skilled in the art should know that CGAN is an improvement based on the GAN model, and additional information is added to the generative network of GAN to improve the generation effect. In the original image generation algorithm based on CGAN, the input of the generator is usually an image and noise with a specified distribution. The main purpose of adding randomly distributed noise to the input in the CGAN algorithm is to make the generated image have diversity and solve the problem of mode collapse of the generated results.

[0079] However, considering that in the implementation of the present disclosure, when two identical photographed test paper images are input, the expected output should be two test paper images with handwritten traces erased and all printed parts completely retained. In other words, the distribution of the desired output data should be completely concentrated rather than divergent. Therefore, in some possible implementation manners, only the training test paper images are input into the noise removal model in step S202 for training, without adding any random guiding information, such as random noise. To ensure that the output of the noise removal model is more stable.

[0080] S203. Input the first generated image and the first test paper image into the HRNet feature matching module for feature matching to determine the first loss value.

[0081] S204. Update the parameters of the noise removal model according to the first loss value.

[0082] With the above solution, by adding image noise to the first test paper image without image noise as the training set and inputting it into the untrained noise removal model, and inputting the output value and the original value of the noise removal model into the HRNet feature matching module for feature matching, and then calculating the loss value, the trained noise removal model can ensure the high quality of the output image.

[0083] In some optional implementations, Figure 3 is a schematic diagram of a noise removal model shown according to an exemplary embodiment, as Figure 3 shown, the noise removal model includes a global generation module G1 and a local enhancement network G2, and the local enhancement network includes a first local enhancement module G2a and a second local enhancement module G2b.

[0084] The inputting the training test paper image into the noise removal model to obtain the first generated image includes:

[0085] Performing two-fold downsampling on the training test paper image and then inputting it into the global generation module G1 to obtain a first feature map; and

[0086] Inputting the training test paper image into the first local enhancement module G1a to obtain a second feature map;

[0087] Adding the first feature map and the second feature map and then inputting the result into the second local enhancement module G2a to obtain the first generated image.

[0088] Specifically, the noise removal model during the training process may include two independent models, G1 (global generator) and G2 (local enhancement network). Figure 3A, B, and C in it represent different types of convolutional network operators. Among them, A represents convolution plus pooling, B represents several layers of ReSBlock, C represents deconvolution, and the rectangle represents the feature map obtained by the model. During the inference process, the input image is sent into G2, and at the same time, the image is downsampled by 2 times and sent into G1. The feature map obtained by G1 is added to the feature map obtained by the first half of G2, that is, G2a, and then sent into the subsequent part of G2, that is, G2b, to obtain the final output image.

[0089] By adopting the above scheme, through designing the global generator and the local enhancement network, the noise removal model can output a test paper image with a higher resolution based on the input noisy test paper image, improving the quality of the image after image denoising.

[0090] Furthermore, based on the Figure 3 noise removal model shown as follows, the method further includes:

[0091] Inputting the training test paper image into the global generation module G1 to train the global generation module G1; and,

[0092] When it is determined that the preset condition is satisfied, inputting the training test paper image into the global generation module G1 and the local enhancement network G2 to train the global generation module G1 and the local enhancement network G2 simultaneously; and,

[0093] When it is determined that the training is completed, extracting the local enhancement network G2 as the trained noise removal model.

[0094] Specifically, training the global generation module G1 includes inputting the training test paper image into G1 to obtain a third generated image, taking this generated image as the output of the noise removal model and inputting it into the above-mentioned HRNet feature matching module to calculate the loss value, and adjusting the parameters of G1 according to the loss value. The same applies to training the global generation module G1 and the local enhancement network G2 simultaneously, which will not be elaborated here.

[0095] In addition, the above-mentioned preset condition can be that the loss function converges, or the number of training iterations reaches the target number, and the present disclosure does not limit this. In addition, determining that the training is completed can be determining that the number of iterations of simultaneous training of G1 and G2 reaches the target number, or satisfying other conditions, and the present disclosure does not limit this either.

[0096] With the above solution, during the inference process, only the output of the G2 network is used as the final output. During the training process, the G1 can be trained first, and after several rounds, the G1 and G2 networks can be trained simultaneously. First, the global generator is trained, then the local enhancement network is trained, and then the parameters of all networks are fine-tuned as a whole. After the training is completed, since the embodiments of the present disclosure do not require global image generation and only need to process local information in the image, therefore, only the local enhancement network G2 can be extracted as the noise removal model and the global generator G1 can be removed, which can effectively reduce the occupation of computing resources.

[0097] In still some alternative embodiments, the method further includes:

[0098] Input the first generated image into the first discriminator to obtain a first discrimination result;

[0099] Downsample the first generated image by a factor of two and then input it into the second discriminator to obtain a second discrimination result;

[0100] Downsample the first generated image by a factor of four and then input it into the third discriminator to obtain a third discrimination result;

[0101] Calculate a second loss value according to the first test paper image, the first discrimination result, the second discrimination result, and the third discrimination result;

[0102] Update the parameters of the noise removal model according to the second loss value.

[0103] It can be understood that generally, only when the discriminator has a large receptive field can it accurately distinguish between real images and synthetic images at a higher resolution. However, to achieve a large receptive field, the network must have more layers or use a larger convolutional kernel. Both of these methods will increase the model capacity and make the network more prone to overfitting, and require more memory and video memory resources. To solve this problem, multiple discriminators are used to handle images of different scales. These discriminators have the same network structure, and the only difference is the size of the input image and the output feature map. These discriminators are only used during the training process to provide an adversarial loss function for the generation network. Among them, the loss functions for calculating the first loss value, the second loss value, and the third loss value can be as follows:

[0104]

[0105] Among them,

[0106] Among them, D represents the discriminator, G represents the generator, x represents the target image, s represents the input image, G(s) represents the image generated by the generator, and log(D()) can be approximately regarded as the likelihood probability output by the discriminator. The above formula simply describes the optimization logic of the algorithm. When optimizing the discriminator, make L_GAN as large as possible, which means making the likelihood probability of the discriminator for the target image higher and the likelihood probability for the generated image lower. When optimizing the generator, make the value as small as possible.

[0107] In addition, the first discriminator, the second discriminator, and the third discriminator can be pre-designed discriminators with fixed parameters, or can be iterated simultaneously according to the output result of the noise removal model and the loss value as the noise removal model iterates. The present disclosure does not limit this.

[0108] Adopting the above solution, by designing discriminators of three different scales to discriminate the images output by the noise removal model, and then obtaining the loss value of the output image to adjust the parameters of the noise removal model, so that the noise removal model can accurately generate test paper images with noise removed at different resolution scales, improving the robustness of the noise removal model.

[0109] To enable those skilled in the art to better understand the technical solutions provided by the present disclosure, the present disclosure also provides a schematic diagram of the training of a noise removal model shown as Figure 4 shown. As Figure 4 shown, in the training process of the noise removal model, it includes a generation module 410, a discrimination module 420, and a feature matching module 430. Among them, the generation module 410 includes a global generation module 411 and a local enhancement network 412, the discrimination module 420 includes a first discriminator 421, a second discriminator 422, and a third discriminator 423, and the feature matching module 430 includes an HRNet feature matching module 431.

[0110] Based on Figure 4 the schematic diagram shown, image noise is added to the first test paper image without image noise to obtain a training test paper image, and the training test paper image is input into the generation module 410 to obtain an output image. Specifically, whether to input the training test paper image into the global generation module 411 and / or the local enhancement network 412 can be determined according to the progress of training, which will not be described in detail here.

[0111] Next, the output image is input into the feature matching module 430 and the discrimination module 420 respectively. Specifically, the original output image, the image after two - fold downsampling, and the image after four - fold downsampling are input into the first discriminator 421, the second discriminator 422, and the third discriminator 423 respectively. And, the loss is calculated through the feature matching module 430 and the discrimination module 420, and the global generation module 411 and / or the local enhancement network 412 in the generation module 410 are iteratively updated according to the loss. Finally, the trained local enhancement network 412 is obtained, and the local enhancement network 412 is extracted as the noise removal model. Among them, the first test paper image can also be input into the discrimination module 420. For the convenience of viewing, it is not shown in Figure 4 it.

[0112] Figure 5 is a schematic diagram of an image processing device 50 shown according to an exemplary embodiment. As Figure 5 shown, the device 50 includes:

[0113] An acquisition module 51, configured to acquire a test paper image, where the test paper image includes image noise;

[0114] A denoising module 52, configured to input the test paper image into a pre - trained noise removal model to obtain a target test paper image with noise removed;

[0115] Among them, the noise removal model is obtained through supervised training based on the HRNet feature matching module.

[0116] Optionally, the device 50 further includes:

[0117] A processing module, configured to add image noise to a first test paper image without image noise to obtain a training test paper image;

[0118] An input module, configured to input the training test paper image into the noise removal model to obtain a first generated image;

[0119] A matching module, configured to input the first generated image and the first test paper image into the HRNet feature matching module for feature matching to determine a first loss value;

[0120] An update module, configured to update the parameters of the noise removal model according to the first loss value.

[0121] Optionally, the noise removal model includes a global generation module and a local enhancement network, and the local enhancement network includes a first local enhancement module and a second local enhancement module.

[0122] The input module is specifically configured to:

[0123] Downsample the training test paper image by a factor of two and input it into the global generation module to obtain a first feature map; and,

[0124] Input the training test paper image into the first local enhancement module to obtain a second feature map;

[0125] Add the first feature map and the second feature map, and then input the result into the second local enhancement module to obtain the first generated image.

[0126] Optionally, the device 50 is specifically configured to:

[0127] Input the training test paper image into the global generation module to train the global generation module; and,

[0128] When it is determined that a preset condition is satisfied, input the training test paper image into the global generation module and the local enhancement network to train the global generation module and the local enhancement network simultaneously; and,

[0129] When it is determined that the training is completed, extract the local enhancement network as the trained noise removal model.

[0130] Optionally, the device 50 is specifically configured to:

[0131] Input the first generated image into the first discriminator to obtain a first discrimination result;

[0132] Downsample the first generated image by a factor of two and input it into the second discriminator to obtain a second discrimination result;

[0133] Downsample the first generated image by a factor of four and input it into the third discriminator to obtain a third discrimination result;

[0134] Calculate a second loss value according to the first test paper image, the first discrimination result, the second discrimination result, and the third discrimination result;

[0135] Update the parameters of the noise removal model according to the second loss value.

[0136] Optionally, the HRNet feature matching module is a network with parallel connections from high to low resolution.

[0137] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0138] Figure 6 It is a block diagram of an electronic device 600 shown according to an exemplary embodiment. As Figure 6As shown, the electronic device 600 may include: a processor 601 and a memory 602. The electronic device 600 may also include one or more of a multimedia component 603, an input / output (I / O) interface 604, and a communication component 605.

[0139] Among them, the processor 601 is used to control the overall operation of the electronic device 600 to complete all or part of the steps in the above image processing method. The memory 602 is used to store various types of data to support the operation of the electronic device 600. These data may include, for example, instructions for any application or method operating on the electronic device 600, and application-related data, such as the first test paper image, sample test paper image, and so on. The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 603 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 602 or transmitted through the communication component 605. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 604 provides an interface between the processor 601 and other interface modules, and the above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 605 is used for wired or wireless communication between the electronic device 600 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G, etc., or a combination of one or more of them is not limited here. Accordingly, the communication component 605 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0140] In one exemplary embodiment, the electronic device 600 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned image processing method.

[0141] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned image processing method are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 602 including program instructions, and the above-mentioned program instructions may be executed by the processor 601 of the electronic device 600 to complete the above-mentioned image processing method.

[0142] Figure 7 FIG. 700 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. For example, the electronic device 700 may be provided as a server. Referring to Figure 7 , the electronic device 700 includes a processor 722, the number of which may be one or more, and a memory 732 for storing computer programs executable by the processor 722. The computer programs stored in the memory 732 may include one or more modules each corresponding to a set of instructions. In addition, the processor 722 may be configured to execute the computer program to execute the above-mentioned image processing method.

[0143] In addition, the electronic device 700 may further include a power supply component 726 and a communication component 750. The power supply component 726 may be configured to perform power management of the electronic device 700, and the communication component 750 may be configured to implement communication of the electronic device 700, for example, wired or wireless communication. In addition, the electronic device 700 may further include an input / output (I / O) interface 758. The electronic device 700 may operate based on an operating system stored in the memory 732, such as Windows Server TM , Mac OSX TM , Unix TM , Linux TM and so on.

[0144] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-described image processing method are implemented. For example, the non-transitory computer-readable storage medium may be the above-described memory 732 including program instructions, and the above program instructions may be executed by the processor 722 of the electronic device 700 to complete the above-described image processing method.

[0145] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program capable of being executed by a programmable device, and the computer program has a code portion for executing the above-described image processing method when executed by the programmable device.

[0146] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0147] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure does not separately describe various possible combination manners.

[0148] Furthermore, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining a test paper image, where the test paper image includes image noise; Inputting the test paper image into a pre-trained noise removal model to obtain a target test paper image with noise removed; Wherein, the noise removal model is obtained through supervised training based on an HRNet feature matching module; The training of the noise removal model includes: Adding image noise to a first test paper image without image noise to obtain a training test paper image; Inputting the training test paper image into the noise removal model to obtain a first generated image; Inputting the first generated image and the first test paper image into the HRNet feature matching module for feature matching to determine a first loss value; Updating the parameters of the noise removal model according to the first loss value; The noise removal model includes a global generation module and a local enhancement network, and the local enhancement network includes a first local enhancement module and a second local enhancement module. The step of inputting the training test paper image into the noise removal model to obtain a first generated image includes: Doubling the downsampling of the training test paper image and inputting it into the global generation module to obtain a first feature map; and Inputting the training test paper image into the first local enhancement module to obtain a second feature map; Adding the first feature map and the second feature map and then inputting the result into the second local enhancement module to obtain the first generated image.

2. The method according to claim 1, characterized in that, The method further includes: Inputting the training test paper image into the global generation module to train the global generation module; and When it is determined that a preset condition is satisfied, inputting the training test paper image into the global generation module and the local enhancement network to train the global generation module and the local enhancement network simultaneously; and When it is determined that the training is completed, extracting the local enhancement network as the trained noise removal model.

3. The method according to claim 1, characterized in that, The method further includes: Inputting the first generated image into a first discriminator to obtain a first discrimination result; Doubling the downsampling of the first generated image and inputting it into a second discriminator to obtain a second discrimination result; Quadrupling the downsampling of the first generated image and inputting it into a third discriminator to obtain a third discrimination result; Calculating a second loss value according to the first test paper image, the first discrimination result, the second discrimination result, and the third discrimination result; Updating the parameters of the noise removal model according to the second loss value.

4. The method according to any one of claims 1-3, characterized in that, The HRNet feature matching module is a network with parallel connections from high to low resolution.

5. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire a test paper image, where the test paper image includes image noise; A denoising module, configured to input the test paper image into a pre-trained noise removal model to obtain a target test paper image with noise removed; Wherein, the noise removal model is obtained through supervised training based on an HRNet feature matching module; The device further includes: A processing module, configured to add image noise to a first test paper image without image noise to obtain a training test paper image; An input module, configured to input the training test paper image into the noise removal model to obtain a first generated image; A matching module, configured to input the first generated image and the first test paper image into the HRNet feature matching module for feature matching to determine a first loss value; An updating module, configured to update the parameters of the noise removal model according to the first loss value; The noise removal model includes a global generation module and a local enhancement network, and the local enhancement network includes a first local enhancement module and a second local enhancement module. The input module is specifically configured to: Perform two-fold downsampling on the training test paper image and input it into the global generation module to obtain a first feature map; and Input the training test paper image into the first local enhancement module to obtain a second feature map; Add the first feature map and the second feature map and input the result into the second local enhancement module to obtain the first generated image.

6. A non - transitory computer - readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-4.

7. An electronic device, characterized in that, Including: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Image restoration method based on multi-scale generative adversarial network model

    CN112541864A

  • Ancient book Chinese character image denoising method based on progressive generative adversarial

    CN113792743A

  • Scene text recognition method based on HRNet coding and double-branch decoding

    CN114140786A