Shadow restoration method for high-resolution document image

By building a low-resolution image correction model and a high-resolution refinement model, and using a dual-channel collaborative conversion framework for rapid expansion, the problem of bottleneck in the shadow repair speed of high-resolution document images in the existing technology is solved, and efficient and high-speed shadow repair effect is achieved.

CN120013815APending Publication Date: 2025-05-16FUJIAN POLICE ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411906577.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing document image shadow repair algorithms have severe bottlenecks in processing high-resolution document images, making it difficult to ensure the quality of repair and real-time processing requirements at the same time. Especially in complex backgrounds or multiple shadows, the processing time is extended, affecting user experience and work efficiency.

Method used

A shadow repair method for high-resolution document images is proposed. By constructing a low-resolution image correction model, high-resolution shadow mixing model and high-resolution refinement model, the dual-channel collaborative conversion framework is used to rapidly scale from low resolution to high resolution, reducing the computational amount and improving processing efficiency.

Benefits of technology

It significantly improves the efficiency of shadow removal, reduces hardware resource requirements, realizes high-quality shadow repair results, and can process large-scale high-resolution document images in a short time, improving user experience and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013815A_ABST
    Figure CN120013815A_ABST
Patent Text Reader

Abstract

The invention relates to a shadow restoration method for a high-resolution document image, and the method comprises the following steps: collecting a high-resolution document image with a shadow, and constructing a shadow image data set; constructing a low-resolution image correction model, down-sampling the shadow image data set, inputting the down-sampled shadow image data set into a background estimation network to extract a shadow prediction thermodynamic diagram of each low-resolution shadow image, and inputting the low-resolution shadow image and the corresponding shadow prediction thermodynamic diagram into a shadow removal network to obtain a low-resolution shadow removal image; constructing a high-resolution shadow hybrid model, and inputting the low-resolution shadow image and the corresponding low-resolution shadow-removed image into the model to obtain a high-resolution shadow-removed image; and constructing a high-resolution refinement model, and inputting the low-resolution shadow removal image corresponding to each shadow image, the shadow prediction thermodynamic diagram, the last layer of feature map of the decoder of the shadow removal network and the high-resolution shadow removal image into the model to obtain a high-resolution document image shadow repair image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a shadow repair method for high-resolution document images, and belongs to the technical field of computer vision. Background Art

[0002] Document images are widely used in daily life and work. Whether it is textbooks, newspapers or various bills, they are usually saved in the form of electronic documents for digital document archiving or online message transmission. With the popularity of smartphones and their high-performance cameras, more and more people use mobile phones instead of scanners to digitize documents. However, when the light source is blocked, shadows may appear in the captured document image; the low brightness of the shadow area reduces the quality and readability of the document image, making the content difficult to recognize and affecting the user experience; in addition, these shadows may cover part of the text, causing great trouble to the subsequent text recognition task; therefore, document image shadow removal is an important image processing task, which is particularly critical to ensure image quality and user experience;

[0003] Existing document image shadow restoration algorithms generally face significant speed bottlenecks when processing high-resolution document images. These methods require detailed analysis of each pixel, resulting in a significant increase in computational complexity, especially in the case of complex backgrounds or multiple shadows, where the processing time is particularly prolonged. Therefore, when processing large-scale high-resolution document images, the overall efficiency is extremely low, which poses a severe challenge to application scenarios that require fast turnaround (such as online document scanning services, instant file sharing platforms, and high-speed document processing systems). These applications usually need to complete the entire process from image acquisition to output in a very short time to provide a smooth user experience. However, current shadow restoration technologies are difficult to simultaneously guarantee restoration quality and real-time processing requirements, which not only affects work efficiency but may also reduce user satisfaction.

[0004] In addition, the refinement network in the existing document image shadow restoration algorithm may find it difficult to fully preserve image details due to insufficient guidance information. Although attention mechanisms and multi-scale analysis are used to enhance the recognition and restoration of shadow areas, some tiny but important structural information may be lost in the process of mapping low-level features to high-level features, such as thin lines of text or subtle texture changes. This information loss will make the final restored image less visually clear than the original image, especially when zoomed in, the lack of details is more obvious. Summary of the invention

[0005] In order to solve the above problems in the prior art, the present invention proposes a shadow restoration method for high-resolution document images.

[0006] The technical solution of the present invention is as follows:

[0007] In one aspect, the present invention provides a shadow repair method for high-resolution document images, comprising the following steps:

[0008] Collect high-resolution document images with shadows, and construct a shadow image dataset after preprocessing the high-resolution document images with shadows;

[0009] Constructing a low-resolution image correction model, wherein the low-resolution image correction model includes a background estimation network and a shadow removal network;

[0010] The shadow image dataset is downsampled to obtain a low-resolution shadow image dataset, and is input into the background estimation network to extract the shadow prediction heat map of each low-resolution shadow image. Each low-resolution shadow image and its corresponding shadow prediction heat map are then input into the de-shadowing network to obtain a low-resolution de-shadowed image.

[0011] A high-resolution shadow mixing model is constructed, and each low-resolution shadow image and its corresponding low-resolution de-shadowed image are input into the high-resolution shadow mixing model to obtain a high-resolution de-shadowed image;

[0012] A high-resolution refinement model is constructed, and the low-resolution de-shadowed image corresponding to each shadow image, the shadow prediction heat map, the last layer feature map of the decoder of the de-shadowing network, and the high-resolution de-shadowed image are input into the high-resolution refinement model to obtain a high-resolution document image shadow restoration image.

[0013] As a preferred embodiment of the present invention, the high-resolution document image with shadows is randomly cropped, and then the cropped image with uniform size is normalized to convert the image data into a standard normal distribution.

[0014] As a preferred embodiment of the present invention, the background estimation network includes several convolutional layers, a global average pooling layer and a fully connected prediction layer;

[0015] The loss function of the background estimation network is The specific formula is as follows:

[0016]

[0017] in: represents the predicted background color value of the low-resolution shadow image; b gt Represents the true background color value of the shadow image;

[0018] The prediction results of the background estimation network are back-propagated to obtain the gradient of the last convolutional layer with respect to the prediction results and calculate the weights. The shadow prediction heat map is obtained by weighted summing the weights with the feature map output by the last convolutional layer.

[0019] As a preferred embodiment of the present invention, the shadow removal network is constructed based on the U-Net network and is composed of a plurality of encoders, a plurality of decoders, a skip connection layer between the encoders and the decoders, and an image fusion layer. The image fusion layer is composed of a first fusion convolutional layer D I And the second fused convolutional layer D M constitute;

[0020] After the low-resolution shadow image and its corresponding shadow prediction heat map are input into the de-shadowing network, the features output by the encoder are then input into the image fusion layer, and the image fusion layer outputs a low-resolution de-shadowed image, as shown in the following formula:

[0021]

[0022] in: represents a low-resolution de-shadowed image; x d Represents the characteristics of the encoder output.

[0023] As a preferred embodiment of the present invention, the high-resolution shadow blending model includes an inverse blending module and a shadow blending module;

[0024] The low-resolution shadow image and its corresponding low-resolution de-shadowed image are passed through the inverse blending module to obtain a low-resolution blended layer, as shown in the following formula:

[0025]

[0026] Among them: B l Represents a low-resolution mixed layer; I l Represents a low-resolution shadow image;

[0027] The high-frequency components of the shadow image are extracted as additional inputs to the inverse blending module to obtain a high-resolution blending layer, as shown in the following formula:

[0028] B h =φ2(h(φ1(Cat(Up(B l ),H))))+Up(B l )

[0029] Among them: B h represents a high-resolution mixed layer; φ1 represents the first convolutional layer; φ2 represents the second convolutional layer; h(·) represents the LeakyReLU function; Up(·) represents upsampling; Cat(·) represents the concatenation of feature maps in the channel dimension;

[0030] Input the high-resolution blending layer into the shadow blending module to obtain a high-resolution de-shadowed image, as shown in the following formula:

[0031]

[0032] in: I represents the high-resolution de-shadowed image; h Represents a shadow image.

[0033] As a preferred embodiment of the present invention, the high-resolution refinement model includes several convolutional layers and an image fusion layer;

[0034] After upsampling the low-resolution de-shadowed image, the image is spliced ​​with the high-resolution de-shadowed image in the channel latitude to obtain an initial spliced ​​image;

[0035] The shadow prediction heat map and the last layer feature map of the decoder of the shadow removal network are upsampled and then spliced ​​with the initial spliced ​​image again to obtain the input spliced ​​image;

[0036] The input spliced ​​image is input into the high-resolution refinement model, and the feature x obtained after the input spliced ​​image passes through the convolution layer c Then input the image fusion layer, and the image fusion layer outputs a high-resolution document image shadow repair image, as shown in the following formula:

[0037]

[0038] in: Represents a high-resolution document image shadow inpainted image.

[0039] As a preferred embodiment of the present invention, the generation loss is calculated for the generation results of the low-resolution image correction model, the high-resolution shadow mixing model and the high-resolution refinement model to obtain the overall generation loss The specific formula is as follows:

[0040]

[0041] in: represents the generation loss of the low-resolution image rectification model; Represents the actual low-resolution shadow-free image; Represents the generation loss of the high-resolution shadow blending model; Indicates the actual shadow-free image; represents the generation loss of the high-resolution refined model;

[0042] At the same time, the shadow edge gradient loss is calculated for the generation results of the low-resolution image correction model, the high-resolution shadow mixing model, and the high-resolution refinement model to obtain the overall shadow edge gradient loss The specific formula is as follows:

[0043]

[0044] in: Represents the shadow edge gradient loss of the low-resolution image correction model; Shadow edge gradient loss representing a high-resolution shadow blending model; represents the shadow edge gradient loss of the high-resolution refinement model; M b Indicates the edge of the shadow; Represents gradient calculation;

[0045] The overall generation loss and the overall shadow edge gradient loss are added together to obtain the overall loss. Based on the overall loss, the gradients of each parameter in the model are calculated by the back propagation method, and the parameters are updated using the stochastic gradient descent method. The operation is repeated until the overall loss converges and stabilizes, and the training is terminated.

[0046] As a preferred embodiment of the present invention, the shadow edge calculation step is:

[0047] Calculate the shadow mask M of the image, as shown below:

[0048]

[0049] Where: I i I represents the input shadow image; o represents the input de-shadowed image; R, G, B represent the red channel value, green channel value, and blue channel value of the image respectively; N represents the normalization function, as shown in the following formula:

[0050]

[0051] Where: I max Represents the maximum value of the input image; I min Represents the minimum value of the input image;

[0052] The shadow mask is expanded and eroded respectively to obtain the corresponding expanded mask and eroded mask, and the shadow edge is obtained by subtracting the expanded mask and the eroded mask.

[0053] On the other hand, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any embodiment of the present invention is implemented.

[0054] In yet another aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the present invention.

[0055] The present invention has the following beneficial effects:

[0056] 1. The present invention realizes efficient shadow removal through three parts: a low-resolution correction network, a high-resolution shadow blending network and a high-resolution refinement network, constructs a two-way collaborative conversion framework, and uses a hybrid network to quickly expand from low resolution to high resolution; performing repair on a low-resolution image can significantly reduce the amount of calculation and reduce the demand for hardware resources, thereby greatly reducing costs; through a rapid expansion mechanism, high-resolution repair results can be quickly generated, greatly improving processing efficiency; the lightweight shadow blending layer design reduces the occupation of computing resources and achieves efficient resource utilization.

[0057] 2. The present invention introduces a background estimation network to generate a shadow prediction heat map, which is used as a guide for local restoration of low-resolution images, thereby enhancing the model's ability to handle shadows under different backgrounds and improving the restoration quality; the network can fully consider the global background and local texture features, provide powerful shadow restoration guidance, and ensure that the restored image has a more natural visual effect.

[0058] 3. Based on the high-resolution shadow mixing network, the present invention further introduces a high-resolution refinement network, combines low-resolution and high-resolution de-shadowed images, and compensates local information with the decoder feature map to produce robust and high-quality shadow restoration results. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flow chart of the method of the present invention;

[0060] Figure 2 This is a structural diagram of the model of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] It should be understood that the step numbers used in this document are only for convenience of description and are not intended to limit the order in which the steps are executed.

[0063] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0064] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0065] The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.

[0066] Embodiment 1:

[0067] See also Figure 1 , a shadow repair method for high-resolution document images, comprising the following steps:

[0068] Collect high-resolution document images with shadows, and construct a shadow image dataset after preprocessing the high-resolution document images with shadows.

[0069] In this embodiment, the document shadow image dataset FSDSRD is used. For the images in the dataset, the dataset is divided into a training set and a test set according to a certain ratio. The training set includes 12780 images and the test set includes 1420 images. Data enhancement is performed on the images in the training set to increase the number of samples in the dataset, including random flipping of images, random cropping of images, photometric distortion, etc.

[0070] Constructing a low-resolution image correction model, wherein the low-resolution image correction model includes a background estimation network and a shadow removal network;

[0071] The shadow image dataset is downsampled to obtain a low-resolution shadow image dataset, and is input into the background estimation network to extract the shadow prediction heat map of each low-resolution shadow image. Each low-resolution shadow image and its corresponding shadow prediction heat map are then input into the de-shadowing network to obtain a low-resolution de-shadowed image.

[0072] A high-resolution shadow mixing model is constructed, and each low-resolution shadow image and its corresponding low-resolution de-shadowed image are input into the high-resolution shadow mixing model to obtain a high-resolution de-shadowed image;

[0073] A high-resolution refinement model is constructed, and the low-resolution de-shadowed image corresponding to each shadow image, the shadow prediction heat map, the last layer feature map of the decoder of the de-shadowing network, and the high-resolution de-shadowed image are input into the high-resolution refinement model to obtain a high-resolution document image shadow restoration image.

[0074] As a preferred implementation of this embodiment, the images in the shadow data set are cropped to 512×512 pixels, and then the cropped images of uniform size are normalized to convert the image data into a standard normal distribution; in order to ensure that the size and position of the label correspond to the high-resolution document shadow image, the same operation is performed on the label while enhancing and preprocessing the image data in each step.

[0075] As a preferred implementation of this embodiment, the background estimation network includes four 3×3 convolutional layers, a global average pooling layer and a fully connected prediction layer;

[0076] The loss function of the background estimation network is The specific formula is as follows:

[0077]

[0078] in: represents the predicted background color value of the low-resolution shadow image; b gt Represents the true background color value of the shadow image;

[0079] The prediction results of the background estimation network are back-propagated to obtain the gradient of the last convolutional layer with respect to the prediction results and calculate the weights. The shadow prediction heat map is obtained by weighted summation of the weights and the feature map output by the last convolutional layer. The last convolutional layer contains rich spatial and semantic information, so that the shadow prediction heat map obtained in this way can effectively capture the relationship between the shadow-free background and the shadowed area in the document image.

[0080] As a preferred implementation of this embodiment, the shadow removal network is constructed based on the U-Net network and consists of several encoders, several decoders, a skip connection layer between the encoder and the decoder, and an image fusion layer. The image fusion layer consists of a 1×1 convolution layer D with a step size of 3. I And a 1×1 convolutional layer D with a stride of 1 M constitute;

[0081] Using low-resolution document shadow images as input to the de-shadowing network can reduce computational complexity and memory burden;

[0082] After the low-resolution shadow image and its corresponding shadow prediction heat map are input into the de-shadowing network, the features output by the encoder are then input into the image fusion layer, and the image fusion layer outputs a low-resolution de-shadowed image, as shown in the following formula:

[0083]

[0084] in: represents a low-resolution de-shadowed image; x dRepresents the characteristics of the encoder output; D I The feature x d Adjust the generated result to the dimension of H×W×3, and D M The feature x d Adjust to a soft attention mask map with a dimension of H×W×1; the image fusion layer combines the generated result with the original image I l Combined, they can produce low-resolution deshading results with more natural restoration effects.

[0085] As a preferred implementation of this embodiment, the high-resolution shadow blending model includes an inverse blending module and a shadow blending module;

[0086] The low-resolution shadow image and its corresponding low-resolution de-shadowed image are passed through the inverse blending module to obtain a low-resolution blended layer, as shown in the following formula:

[0087]

[0088] Among them: B l Represents a low-resolution mixed layer; I l Represents a low-resolution shadow image;

[0089] Because B l It is generated from the low-resolution correction result. It cannot retain the transformation information of the high-frequency components. In order to apply the low-resolution mixed layer on the high-resolution image, an effective strategy is needed to compensate for the information loss caused by downsampling. Therefore, the high-frequency components of the shadow image are extracted as additional inputs to the inverse mixing module to obtain the high-resolution mixed layer, as shown in the following formula:

[0090] B h =φ2(h(φ1(Cat(Up(B l ),H))))+Up(B l )

[0091] Among them: B h represents a high-resolution mixed layer; φ1 represents the first convolutional layer; φ2 represents the second convolutional layer; h(·) represents the LeakyReLU function; Up(·) represents bilinear interpolation upsampling; Cat(·) represents feature map concatenation in the channel dimension;

[0092] High Resolution Blending Layer B h High-resolution detail conversion information is preserved;

[0093] Input the high-resolution blending layer into the shadow blending module to obtain a high-resolution de-shadowed image, as shown in the following formula:

[0094]

[0095] in: I represents the high-resolution de-shadowed image; h represents a shadow image;

[0096] The inverse blending module and the shadow blending module act as intermediates to achieve rapid expansion from low-resolution results to high-resolution, ensuring the scalability and detail fidelity of the network.

[0097] As a preferred implementation of this embodiment, the high-resolution refinement model includes two convolutional layers (the convolution kernel size is 3×3 and the step size is 1) and an image fusion layer;

[0098] After upsampling the low-resolution de-shadowed image, the image is spliced ​​with the high-resolution de-shadowed image in the channel latitude to obtain an initial spliced ​​image;

[0099] The shadow prediction heat map and the last layer feature map of the decoder of the shadow removal network are upsampled and then spliced ​​with the initial spliced ​​image again to obtain the input spliced ​​image;

[0100] The input spliced ​​image is input into the high-resolution refinement model, and the feature x obtained after the input spliced ​​image passes through the convolution layer c Then input the image fusion layer, and the image fusion layer outputs a high-resolution document image shadow repair image, as shown in the following formula:

[0101]

[0102] in: Represents a high-resolution document image shadow inpainted image.

[0103] As a preferred implementation of this embodiment, the generation loss is calculated for the generation results of the low-resolution image correction model, the high-resolution shadow mixing model and the high-resolution refinement model to obtain the overall generation loss The specific formula is as follows:

[0104]

[0105] in: represents the generation loss of the low-resolution image rectification model; Represents the actual low-resolution shadow-free image; Represents the generation loss of the high-resolution shadow blending model; Indicates the actual shadow-free image; represents the generation loss of the high-resolution refined model;

[0106] At the same time, the shadow edge gradient loss is calculated for the generation results of the low-resolution image correction model, the high-resolution shadow mixing model, and the high-resolution refinement model to obtain the overall shadow edge gradient loss The specific formula is as follows:

[0107]

[0108] in: Represents the shadow edge gradient loss of the low-resolution image correction model; Shadow edge gradient loss representing a high-resolution shadow blending model; represents the shadow edge gradient loss of the high-resolution refinement model; M b Indicates the edge of the shadow; Represents gradient calculation;

[0109] The overall generation loss and the overall shadow edge gradient loss are added together to obtain the overall loss. Based on the overall loss, the gradients of each parameter in the model are calculated by the back propagation method, and the parameters are updated using the stochastic gradient descent method. The operation is repeated until the overall loss converges and stabilizes, and the training is terminated.

[0110] As a preferred implementation of this embodiment, the shadow edge calculation step is:

[0111] Calculate the shadow mask M of the image, as shown below:

[0112]

[0113] Where: I i I represents the input shadow image; o represents the input de-shadowed image; R, G, B represent the red channel value, green channel value, and blue channel value of the image respectively; N represents the normalization function, as shown in the following formula:

[0114]

[0115] Where: I max Represents the maximum value of the input image; I min Represents the minimum value of the input image;

[0116] The shadow mask is expanded and eroded respectively to obtain the corresponding expanded mask and eroded mask, and the shadow edge is obtained by subtracting the expanded mask and the eroded mask.

[0117] Embodiment 2:

[0118] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any embodiment of the present invention when executing the program.

[0119] Embodiment three:

[0120] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in any embodiment of the present invention is implemented.

[0121] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.

[0122] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0123] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0124] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.

[0125] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A shadow restoration method for high-resolution document images, characterized in that: The following steps are involved: Collect high-resolution document images with shadows, and construct a shadow image dataset after preprocessing the high-resolution document images with shadows; Constructing a low-resolution image correction model, wherein the low-resolution image correction model includes a background estimation network and a shadow removal network; The shadow image dataset is downsampled to obtain a low-resolution shadow image dataset, and is input into the background estimation network to extract the shadow prediction heat map of each low-resolution shadow image. Each low-resolution shadow image and its corresponding shadow prediction heat map are then input into the de-shadowing network to obtain a low-resolution de-shadowed image. A high-resolution shadow mixing model is constructed, and each low-resolution shadow image and its corresponding low-resolution de-shadowed image are input into the high-resolution shadow mixing model to obtain a high-resolution de-shadowed image; A high-resolution refinement model is constructed, and the low-resolution de-shadowed image corresponding to each shadow image, the shadow prediction heat map, the last layer feature map of the decoder of the de-shadowing network, and the high-resolution de-shadowed image are input into the high-resolution refinement model to obtain a high-resolution document image shadow restoration image.

2. The shadow repair method for high-resolution document images according to claim 1, characterized in that: The high-resolution document images with shadows are randomly cropped, and then the cropped images with uniform size are normalized to transform the image data into a standard normal distribution.

3. The shadow repair method for high-resolution document images according to claim 1, characterized in that: The background estimation network includes several convolutional layers, a global average pooling layer and a fully connected prediction layer; The loss function of the background estimation network is The specific formula is as follows: in: represents the predicted background color value of the low-resolution shadow image; b gt Represents the true background color value of the shadow image; The prediction results of the background estimation network are back-propagated to obtain the gradient of the last convolutional layer with respect to the prediction results and calculate the weights. The shadow prediction heat map is obtained by weighted summing the weights with the feature map output by the last convolutional layer.

4. The shadow repair method for high-resolution document images according to claim 3, characterized in that: The de-shadowing network is constructed based on the U-Net network and consists of several encoders, several decoders, a skip connection layer between the encoders and decoders, and an image fusion layer. The image fusion layer consists of the first fusion convolutional layer D I And the second fused convolutional layer D M constitute; After the low-resolution shadow image and its corresponding shadow prediction heat map are input into the de-shadowing network, the features output by the encoder are then input into the image fusion layer, and the image fusion layer outputs a low-resolution de-shadowed image, as shown in the following formula: in: represents a low-resolution de-shadowed image; x d Represents the characteristics of the encoder output.

5. The shadow repair method for high-resolution document images according to claim 4, characterized in that: The high-resolution shadow blending model includes an inverse blending module and a shadow blending module; The low-resolution shadow image and its corresponding low-resolution de-shadowed image are passed through the inverse blending module to obtain a low-resolution blended layer, as shown in the following formula: Among them: B l Represents a low-resolution mixed layer; I l Represents a low-resolution shadow image; The high-frequency components of the shadow image are extracted as additional inputs to the inverse blending module to obtain a high-resolution blending layer, as shown in the following formula: B h =φ2(h(φ1(Cat(Up(B l ),H))))+Up(B l ) Among them: B h represents a high-resolution mixed layer; φ1 represents the first convolutional layer; φ2 represents the second convolutional layer; h(·) represents the LeakyReLU function; Up(·) represents upsampling; Cat(·) represents the concatenation of feature maps in the channel dimension; Input the high-resolution blending layer into the shadow blending module to obtain a high-resolution de-shadowed image, as shown in the following formula: in: I represents the high-resolution de-shadowed image; h Represents a shadow image.

6. The shadow repair method for high-resolution document images according to claim 5, characterized in that: The high-resolution refinement model includes several convolutional layers and an image fusion layer; After upsampling the low-resolution de-shadowed image, the image is spliced ​​with the high-resolution de-shadowed image in the channel latitude to obtain an initial spliced ​​image; The shadow prediction heat map and the last layer feature map of the decoder of the shadow removal network are upsampled and then spliced ​​with the initial spliced ​​image again to obtain the input spliced ​​image; The input spliced ​​image is input into the high-resolution refinement model, and the feature x obtained after the input spliced ​​image passes through the convolution layer c Then input the image fusion layer, and the image fusion layer outputs a high-resolution document image shadow repair image, as shown in the following formula: in: Represents a high-resolution document image shadow inpainted image.

7. The shadow repair method for high-resolution document images according to claim 6, characterized in that: At the same time, the generation loss is calculated for the generation results of the low-resolution image correction model, the high-resolution shadow mixing model, and the high-resolution refinement model to obtain the overall generation loss The specific formula is as follows: in: represents the generation loss of the low-resolution image rectification model; Represents the actual low-resolution shadow-free image; Represents the generation loss of the high-resolution shadow blending model; Indicates the actual shadow-free image; represents the generation loss of the high-resolution refined model; At the same time, the shadow edge gradient loss is calculated for the generation results of the low-resolution image correction model, the high-resolution shadow mixing model, and the high-resolution refinement model to obtain the overall shadow edge gradient loss The specific formula is as follows: in: Represents the shadow edge gradient loss of the low-resolution image correction model; Shadow edge gradient loss representing a high-resolution shadow blending model; represents the shadow edge gradient loss of the high-resolution refinement model; M b Indicates the edge of the shadow; Represents gradient calculation; The overall generation loss and the overall shadow edge gradient loss are added together to obtain the overall loss. Based on the overall loss, the gradients of each parameter in the model are calculated by the back propagation method, and the parameters are updated using the stochastic gradient descent method. The operation is repeated until the overall loss converges and stabilizes, and the training is terminated.

8. The shadow restoration method for high-resolution document images according to claim 7, characterized in that: The shadow edge calculation steps are: Calculate the shadow mask M of the image, as shown below: Where: I i I represents the input shadow image; o represents the input de-shadowed image; R, G, B represent the red channel value, green channel value, and blue channel value of the image respectively; N represents the normalization function, as shown in the following formula: Where: I max Represents the maximum value of the input image; I min Represents the minimum value of the input image; The shadow mask is expanded and eroded respectively to obtain the corresponding expanded mask and eroded mask, and the shadow edge is obtained by subtracting the expanded mask and the eroded mask.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.