Document image shadow removal method, device, computer equipment and storage medium
Through the multi-scale shadow flow matching framework and adaptive adjustment of contrast information, the problems of high computational requirements and poor results in existing technologies are solved, and efficient shadow removal and document detail preservation in complex shadow scenes are achieved.
Patent Information
- Application Number
- CN202411922087.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing document image processing technologies rely on external auxiliary tools or complex network designs during the shadow removal process, resulting in high computational requirements and poor performance in complex and changing shadow scenes, and are unable to effectively preserve the integrity and tonal consistency of document content.
The shadow removal process is divided into stages of different shadow scales through a multi-scale shadow flow matching framework. The contrast information of the document image is used for adaptive adjustment and conditional probability path learning to reduce computational redundancy and achieve efficient and automatic shadow detection and positioning.
In complex and changing shadow scenes such as uneven lighting or strong projection areas, it accurately removes shadows and retains document details, reducing dependence on external information and improving shadow removal effects.
Smart Images

Figure CN119762404B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of document image processing, and in particular to a document image shadow removal method, apparatus, computer equipment, and storage medium. Background Art
[0002] With the increasing demand for document digitization, document image processing technology plays a vital role in modern information systems. The successful application of tasks such as OCR (Optical Character Recognition), table recognition, and document layout reconstruction has placed higher demands on the quality of document images. However, in the actual process of photographing or scanning documents, shadows often appear in document images due to reasons such as light source occlusion and the complexity of ambient lighting. These shadows can obscure key information content such as text, table lines, and background details in the document, directly leading to a decrease in the accuracy of information extraction and affecting the performance of subsequent tasks. Therefore, document shadow removal has become a challenging and important research direction in low-level vision tasks.
[0003] Document shadows have characteristics that are different from shadows in natural scenes. Natural scene shadow removal typically deals with complex and diverse environmental backgrounds, such as lighting changes, complex textures, or light and shadow interactions between objects. Document images, on the other hand, typically have a high information density, especially structured information such as text, tables, and layouts. The presence of shadows not only reduces document readability but can also lead to information loss or misidentification. Therefore, document shadow removal requires not only visually removing shadows but also maintaining the integrity of the document content and the consistency of tones. Currently, research on document shadow removal mainly includes the following categories:
[0004] First, shadow mask-based methods: This method generates a shadow mask to guide shadow removal, such as using semantic segmentation or specific region marking. However, this type of method relies on high-quality shadow masks, and generating the mask itself requires additional annotation or computational overhead.
[0005] 2. Deep learning methods: In recent years, with the development of deep learning technology, some methods have attempted to use convolutional neural networks (CNN) or Transformer-based models to improve the removal effect by increasing network depth and introducing attention mechanisms. However, this has also increased the complexity of network design and computational overhead.
[0006] 3. Data enhancement-based method: This method improves the shadow removal effect by constructing a large-scale labeled dataset or frequency domain feature decomposition, such as the SD7K dataset. It requires a large amount of high-quality data and may have insufficient generalization capabilities in different shadow scenarios.
[0007] The inventors discovered in their research that document shadows exhibit significant contrast differences between shadowed and non-shadowed areas. However, most existing methods fail to fully utilize the inherent information in the document shadow image itself, relying instead on external tools (such as shadow masks) or complex network designs. Complex network structures and multi-stage training processes increase computational requirements, placing excessive demands on training and inference resources. Furthermore, existing methods are effective for certain types of shadows but are prone to failure in complex and variable shadow scenarios, such as those with uneven lighting or areas with strong shadow projections. Summary of the Invention
[0008] In view of this, the present application provides a document image shadow removal method, apparatus, computer equipment and storage medium, aiming to solve at least one of the above-mentioned technical problems in the prior art to a certain extent.
[0009] In order to solve the above problems, this application provides the following technical solutions:
[0010] A document image shadow removal method, comprising:
[0011] Using image processing technology to adjust the contrast of the document shadow image to be processed, and obtain a contrast heat map that highlights the shadow area;
[0012] A multi-scale shadow flow matching framework is used to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales. In each stage, a conditional probability path between the data and noise in the current stage is learned through a flow matching model, and shadow removal is performed on the current stage according to the conditional probability path.
[0013] Based on the shadow removal result of the current stage, a next-scale prediction mechanism is used to predict the shadow scale of the next stage, and the contrast heat map is used to guide the shadow removal in the next stage;
[0014] When all stages are processed, the final document shadow-free image is output.
[0015] The technical solution adopted by the embodiment of the present application further includes: after using the image processing technology to adjust the contrast of the document shadow image to be processed and obtaining a contrast heat map highlighting the shadow area, it also includes:
[0016] The contrast heat map is input into an adjustment network, and the contrast heat map is adaptively adjusted by the adjustment network to generate an adaptive shadow scale heat map.
[0017] The technical solution adopted by the embodiment of the present application further includes: before using the multi-scale shadow flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales, it also includes:
[0018] The document shadow image is preprocessed; the preprocessing specifically includes: performing multiple downsampling processes on the document shadow image to reduce the resolution of the document shadow image.
[0019] The technical solution adopted in the embodiment of the present application also includes: learning the conditional probability path between data and noise in the current stage through the flow matching model, and performing shadow removal on the current stage according to the conditional probability path, specifically:
[0020] Interpolate the flow between the data and the compressed low-resolution shadow representation through the flow matching model, and divide the time step of the flow matching model into K resolution windows, each resolution window is half of the previous resolution window; represents the interpolated version, the stream is represented as:
[0021]
[0022] Among them, t represents the time, Down(·,2 k ) represents the downsampling function in the k-th resolution window, is the sampled random noise and x1 is a data sample; for the k-th resolution window [t min,k ,t max,k ], let t′=(tt min,k ) / (t max,k -t min,k ) represents the rescaled time step, then the flow follows:
[0023]
[0024] Among them, Up(·) and Down(·) represent the upsampling function and downsampling function respectively, and Represent the data values at the last and first moments in the kth window respectively.
[0025] The technical solution adopted in the embodiment of the present application also includes: the conditional probability path is defined as follows:
[0026]
[0027] Among them, t max,k and t min,k Respectively represent the last and first moments under the k-th window, Indicates that it obeys the normal distribution;
[0028] At the kth window resolution, the stream is matched to the model v t Return to the conditional vector field And unify shadow removal and image enhancement through flow matching objectives:
[0029]
[0030] The technical solution adopted in the embodiment of the present application also includes: the calculation formula of the next scale prediction mechanism is:
[0031]
[0032] Where Δt represents the offset time, Δt=t max,t -t min,t , v t Represents a flow matching model.
[0033] The technical solution adopted in the embodiment of the present application also includes: based on the shadow removal result of the current stage, using the next scale prediction mechanism to predict the shadow scale of the next stage, and using the contrast heat map to guide the shadow removal in the next stage, specifically:
[0034] The shadow removal result of the current stage is input into the contrast perception module of the next scale prediction mechanism. The contrast perception module adaptively perceives the shadow scale of the next stage and uses the shadow scale heat map as conditional information to guide the shadow scale prediction of the next stage:
[0035]
[0036] Among them, v θ represents the v-prediction network, c is the extracted contrast-based shadow scale representation, conditioned on the flow matching model via cross attention, and x inp is the document shadow image, is the target image predicted at time t, is the shadow scale prediction result of the previous stage.
[0037] Another technical solution adopted in the embodiment of the present application is: a document image shadow removal device, comprising:
[0038] Contrast adjustment module: used to adjust the contrast of the document shadow image to be processed using image processing technology to obtain a contrast heat map that highlights the shadow area;
[0039] Flow matching module: used to use a multi-scale shadow flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales. In each stage, the flow matching model is used to learn the conditional probability path between the data and noise in the current stage, and shadow removal is performed on the current stage according to the conditional probability path;
[0040] Scale prediction module: used to predict the shadow scale of the next stage based on the shadow removal result of the current stage using the next scale prediction mechanism, use the contrast heat map to guide the shadow removal of the next stage, and output the final document shadow-free image after all stages are processed.
[0041] Compared with the prior art, the beneficial effects produced by the embodiments of the present application are as follows: the document image shadow removal method, device, computer equipment and storage medium of the embodiments of the present application effectively locate the shadow area by utilizing the contrast information in the document image, and use it as the core condition for shadow removal. By designing a multi-scale shadow flow matching framework, the shadow removal process of the document shadow image is divided into multiple stages representing different shadow scales, and the document shadow image is subjected to staged shadow removal according to the shadow scale. In the shadow removal process, contrast information is used as conditional information to achieve efficient and automatic shadow detection and positioning, which greatly reduces computational redundancy. The embodiments of the present application reduce dependence on external information, do not require additional annotation or external auxiliary data, and can accurately remove shadows and retain detailed information in the original document in complex and changeable shadow scenes such as uneven lighting or strong projection areas, greatly improving the shadow removal effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flow chart of a document image shadow removal method according to an embodiment of the present application;
[0043] Figure 2 is a schematic diagram of image contrast correlation according to an embodiment of the present application, wherein (a) is a document shadow image, and (b) is a contrast heat map extracted from the document shadow image;
[0044] Figure 3 This is the shadow scale heat map adjusted by adjusting the network in the embodiment of the present application;
[0045] Figure 4 This is a schematic diagram of a multi-scale shadow flow matching framework according to an embodiment of the present application;
[0046] Figure 5 This is a schematic structural diagram of a document image shadow removal device according to an embodiment of the present application;
[0047] Figure 6 This is a schematic diagram of the computer device structure according to an embodiment of the present application;
[0048] Figure 7 A schematic diagram of the structure of the storage medium of an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0050] The terms "first," "second," and "third" in this application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of this application, "multiple" means at least two, for example, two, three, etc., unless otherwise specifically defined. All directional indications in the embodiments of this application (such as up, down, left, right, front, back...) are only used to explain the relative positional relationship, movement, etc. between the components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications also change accordingly. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or computer device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products, or computer devices.
[0051] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0052] Specifically, see Figure 1 , is a flow chart of a document image shadow removal method according to an embodiment of the present application. The document image shadow removal method according to an embodiment of the present application comprises the following steps:
[0053] S100: Acquire a document shadow image to be processed;
[0054] S110: Using image processing technology to adjust the contrast of the document shadow image to generate a contrast heat map that highlights the shadow area;
[0055] In this step, since there is a significant difference in contrast between the shadow area and the non-shadow area in the document shadow image, in order to more reasonably represent the shadow scale, the embodiment of the present application starts with contrast analysis, and uses the significant difference in contrast between the shadow area and the non-shadow area to adjust the contrast of the document shadow image through image processing technology to generate a contrast heat map c with a clear contrast. h , which can not only effectively locate the shadow area, but also provide rich guidance information for the shadow removal process without the need for additional annotation or external auxiliary data. Figure 2 Figure 2 shows a schematic diagram of image contrast in an embodiment of the present application, where (a) is a document shadow image and (b) is a contrast heat map extracted from the document shadow image. In the embodiments of the present application, image processing techniques include, but are not limited to, image digitization or image enhancement and restoration.
[0056] S120: Inputting the contrast heat map into the adjustment network, adaptively adjusting the contrast heat map through the adjustment network to generate an adaptive shadow-scale heat map;
[0057] In this step, since different images have different responses to contrast adjustment, the embodiment of the present application converts the contrast heat map c generated by the image processing technology into h Input into the adjustment network for adaptive adjustment to adapt to images under different lighting conditions and generate an adaptive shadow scale heat map c = a θ (c h ). Figure 3 The figure shows the shadow scale heatmap adjusted by the network in the embodiment of this application. The shadow scale heatmap can adapt to the contrast differences between different images, capturing the intensity and location characteristics of the shadow area, which can serve as conditional information to guide shadow removal. Compared with static masks, adaptive dynamic adjustment is more adaptable and accurate.
[0058] S130: Preprocessing the document shadow image and using a multi-shadow-scale flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales;
[0059] In this step, preprocessing includes: performing multiple downsampling processes on the document shadow image to reduce the resolution of the document shadow image and reduce the computational complexity of subsequent processing; then, dividing the shadow removal process of the document shadow image into multiple stages (Stage k, Stage k+1,...), so that each stage corresponds to a different shadow scale.
[0060] S140: In each stage, a conditional probability path between the data and noise in the current stage is learned through a flow matching model, and shadow removal is performed on the document shadow image in the current stage according to the conditional probability path;
[0061] In this step, flow matching is similar to the diffusion model, which aims to learn a velocity field v t , random noise is converted into Mapping to data samples x1~q:
[0062]
[0063] The flow matching framework directly regresses the velocity v t To the conditional vector field u t (·|x1), which provides a simplified training objective for flow generation models without simulation:
[0064]
[0065] Among them, u t (·|x1) uniquely determines a conditional probability path leading to data sample x1. An effective choice of conditional probability path is the linear interpolation of data and noise:
[0066]
[0067] And u(x t |x1)=x1-x0.
[0068] Flow matching can be flexibly extended to interpolation between distributions other than the standard Gaussian distribution.
[0069] Furthermore, under the multi-scale shadow flow matching framework, the document shadow image in the early stage shows strong shadows and severe degradation. Therefore, the early stage only needs to focus on global illumination adjustment and removal of severe shadows at a lower resolution to reduce computational costs. As the shadow removal progresses, the shadow of the image gradually becomes lighter. The mid-term stage requires local area restoration at a medium resolution, while the late stage requires fine detail enhancement and text restoration at a high resolution. In order to reduce redundant calculations in the early stages, the embodiment of the present application interpolates the flow between the data and the compressed low-resolution shadow representation through a flow matching model. Specifically, the time step of the flow matching model is divided into K resolution windows, each resolution window is half of the previous resolution window; represents the interpolated version, the stream can be expressed as:
[0070]
[0071] Where t represents time, Down(·,2 k ) represents the downsampling function in the k-th resolution window, is the sampled random noise and x1 is a data sample. For the k-th resolution window [t min,k ,t max,k ], let t′=(tt min,k ) / (t max,k -t min,k ) represents the rescaled time step, then its flow follows:
[0072]
[0073] Among them, Up(·) and Down(·) represent the upsampling function and downsampling function respectively. and denote the data values at the last and first moments in the kth window, respectively. Thus, only the last stage performs shadow removal at full resolution, while most other stages perform shadow removal at low resolution. With uniform stage division, the multi-scale shadow flow matching framework reduces the computation time to approximately 1 / K, thus reducing computational requirements.
[0074] Furthermore, if Figure 4 The figure shows a schematic diagram of the multi-scale shadow flow matching framework of an embodiment of the present application. In the process of constructing the multi-scale shadow flow matching framework, the main goal is to uniformly model the different stages. In order to unify the goals of shadow removal and image enhancement, the embodiment of the present application constructs a conditional probability path by interpolating between different resolutions and shadow level conditions, starting from a high-noise, high-shadow version at low resolution, and gradually generating clearer and more detailed results at high resolution. Specifically, the conditional probability path is defined as follows:
[0075]
[0076]
[0077] In the above formula, t max,k and t min,k Respectively represent the last and first moments under the k-th window, Indicates that it obeys the normal distribution, the former is the mean and the latter is the variance.
[0078] Then, at the kth window resolution, the flow is matched to the model v t Return to the conditional vector field And the shadow removal and image enhancement are unified through the following flow matching targets:
[0079]
[0080] S150: Based on the shadow removal result of the current stage, the shadow scale of the next stage is predicted through the next scale prediction mechanism, and the shadow scale heat map is used as conditional information to guide the shadow removal of the next stage;
[0081] In this step, a next-scale prediction mechanism is constructed by combining the generative capabilities of the diffusion model and the efficiency of the multi-scale framework. Shadow removal tasks of different intensities are distributed to different stages. The shadow removal results of the current stage are used to gradually predict the shadow scale of the next stage. The shadow scale heat map is used as conditional information to guide the shadow removal of the next stage, gradually generating high-quality shadow-free images of the document.
[0082] Specifically, the calculation formula of the next scale prediction mechanism is:
[0083]
[0084] Where Δt represents the offset time, Δt=t max,t -t min,t , v t Represents a flow matching model.
[0085] This process combines the iterative generation and enhancement capabilities of the diffusion model with the next-scale prediction advantages of the autoregressive model to form an efficient training and inference framework, which improves the shadow removal effect while reducing the amount of computation.
[0086] By using the shadow removal result of the current stage as the quantification of the shadow scale of each stage, it is input into the contrast perception module of the next scale prediction mechanism. The contrast perception module adaptively perceives the shadow scale of the next stage, and uses the shadow scale heat map as conditional information to guide the shadow scale prediction of the next stage:
[0087]
[0088]
[0089]
[0090] Among them, v θ represents the v-prediction network, c is the extracted contrast-based shadow scale representation, conditioned on the flow matching model via cross attention, and x inp is the input document shadow image, is the target image predicted at time t. is the shadow scale prediction result of the previous stage, which serves as the conditional information of the current stage. Therefore, at t=0, xinp As x inp and Then Connect on the channel.
[0091] S160: After all stages are processed, the final document shadow-free image is output.
[0092] Based on the above, the document image shadow removal method of the embodiment of the present application effectively locates the shadow area by utilizing the contrast information in the document image, and uses it as the core condition for shadow removal. By designing a multi-scale shadow flow matching framework, the shadow removal process of the document shadow image is divided into multiple stages representing different shadow scales, and the document shadow image is subjected to staged shadow removal according to the shadow scale. In the shadow removal process, contrast information is used as conditional information to achieve efficient and automatic shadow detection and positioning, greatly reducing computational redundancy. The embodiment of the present application reduces dependence on external information, does not require additional annotation or external auxiliary data, and can accurately remove shadows and retain detailed information in the original document in complex and changeable shadow scenes such as uneven lighting or strong projection areas, greatly improving the shadow removal effect.
[0093] See also Figure 5 , is a schematic diagram of the structure of a document image shadow removal device according to an embodiment of the present application. The document image shadow removal device 40 according to an embodiment of the present application comprises:
[0094] Contrast adjustment module 41: used to adjust the contrast of the document shadow image to be processed by using image processing technology, and obtain a contrast heat map highlighting the shadow area;
[0095] Flow matching module 42: configured to use a multi-scale shadow flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales. In each stage, a conditional probability path between the data and noise in the current stage is learned through a flow matching model, and shadow removal is performed on the current stage based on the conditional probability path.
[0096] Scale prediction module 43: used to predict the shadow scale of the next stage based on the shadow removal result of the current stage using the next scale prediction mechanism, use the contrast heat map to guide the shadow removal of the next stage, and output the final document shadow-free image after all stages are processed.
[0097] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0098] The device provided in the embodiment of the present application can be applied in the aforementioned method embodiment. For details, please refer to the description of the aforementioned method embodiment, which will not be repeated here.
[0099] See also Figure 6 , is a schematic diagram of the computer device structure of an embodiment of the present application. The computer device 50 includes:
[0100] A memory 51 storing executable program instructions;
[0101] a processor 52 connected to the memory 51;
[0102] The processor 52 is used to call the executable program instructions stored in the memory 51 and perform the following steps: using image processing technology to adjust the contrast of the document shadow image to be processed, and obtain a contrast heat map of the highlighted shadow area; using a multi-scale shadow flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales, in each stage, using the flow matching model to learn the conditional probability path between the data and noise in the current stage, and performing shadow removal on the current stage according to the conditional probability path; based on the shadow removal result of the current stage, using the next scale prediction mechanism to predict the shadow scale of the next stage, and using the contrast heat map to guide the shadow removal of the next stage; when all stages are processed, outputting the final document shadow-free image.
[0103] The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip having signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.
[0104] See also Figure 7, is a schematic diagram of the structure of the storage medium of an embodiment of the present application. The storage medium of the embodiment of the present application stores program instructions 61 capable of implementing the following steps: using image processing technology to adjust the contrast of the document shadow image to be processed, and obtaining a contrast heat map of the highlighted shadow area; using a multi-scale shadow flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales, in each stage, using a flow matching model to learn the conditional probability path between the data and noise in the current stage, and performing shadow removal on the current stage according to the conditional probability path; based on the shadow removal result of the current stage, using a next scale prediction mechanism to predict the shadow scale of the next stage, and using the contrast heat map to guide the shadow removal of the next stage; when all stages are processed, outputting the final document shadow-free image.
[0105] Among them, the program instructions 61 can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network computer device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application. The aforementioned storage medium includes: various media that can store program instructions, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or a terminal computer device such as a computer, a server, a mobile phone, or a tablet. Among them, the server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0106] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0107] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the content of the description and drawings of this application, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A document image shadow removal method, characterized in that: include: Using image processing technology to adjust the contrast of the document shadow image to be processed, and obtain a contrast heat map that highlights the shadow area; A multi-scale shadow flow matching framework is used to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales. In each stage, a conditional probability path between the data and noise in the current stage is learned through a flow matching model, and shadow removal is performed on the current stage according to the conditional probability path. Based on the shadow removal result of the current stage, a next-scale prediction mechanism is used to predict the shadow scale of the next stage, and the contrast heat map is used to guide the shadow removal in the next stage; When all stages are processed, the final document shadow-free image is output.
2. The document image shadow removal method according to claim 1, wherein: After using image processing technology to adjust the contrast of the document shadow image to be processed and obtaining a contrast heat map highlighting the shadow area, the method further includes: The contrast heat map is input into an adjustment network, and the contrast heat map is adaptively adjusted by the adjustment network to generate an adaptive shadow scale heat map.
3. The document image shadow removal method according to claim 2, characterized in that: Before evenly dividing the shadow removal process of the document shadow image into K stages representing different shadow scales using the multi-scale shadow flow matching framework, the method further includes: The document shadow image is preprocessed; the preprocessing specifically includes: performing multiple downsampling processes on the document shadow image to reduce the resolution of the document shadow image.
4. The document image shadow removal method according to claim 3, wherein: The conditional probability path between the data and noise in the current stage is learned by the flow matching model, and shadow removal is performed on the current stage according to the conditional probability path, specifically: Interpolate the flow between the data and the compressed low-resolution shadow representation through the flow matching model, and divide the time step of the flow matching model into K resolution windows, each resolution window is half of the previous resolution window; represents the interpolated version, the stream is represented as: Where t represents time, Down(·,2 k ) represents the downsampling function in the k-th resolution window, is the sampled random noise and x1 is a data sample; for the K-th resolution window [t min,k ,t max,k ], let t′=(tt min,k ) / (t max,k -t min,k ) represents the rescaled time step, then the flow follows: Among them, Up(·) and Down(·) represent the upsampling function and downsampling function respectively. and Represent the data values at the last and first moments in the kth window respectively.
5. The document image shadow removal method according to claim 4, characterized in that: The conditional probability path is defined as follows: Among them, t max,k and t min,k Respectively represent the last and first moments under the k-th window, Indicates that it obeys the normal distribution; At the kth window resolution, the stream is matched to the model v t Return to the conditional vector field And unify shadow removal and image enhancement through flow matching objectives:
6. The document image shadow removal method according to any one of claims 1 to 5, characterized in that: The calculation formula of the next scale prediction mechanism is: Where Δt represents the offset time, Δt=t max,t -t min,t , v t Represents a flow matching model.
7. The document image shadow removal method according to claim 6, characterized in that: Based on the shadow removal result of the current stage, the shadow scale of the next stage is predicted using the next scale prediction mechanism, and the contrast heat map is used to guide the shadow removal in the next stage, specifically: The shadow removal result of the current stage is input into the contrast perception module of the next scale prediction mechanism. The contrast perception module adaptively perceives the shadow scale of the next stage and uses the shadow scale heat map as conditional information to guide the shadow scale prediction of the next stage: Among them, v θ represents the v-prediction network, c is the extracted contrast-based shadow scale representation, conditioned on the flow matching model via cross attention, and x inp is the document shadow image, is the target image predicted at time t, is the shadow scale prediction result of the previous stage.
8. A document image shadow removal device, characterized in that: include: Contrast adjustment module: used to adjust the contrast of the document shadow image to be processed using image processing technology, and obtain a contrast heat map that highlights the shadow area; Flow matching module: used to use a multi-scale shadow flow matching framework to evenly divide the shadow removal process of the document shadow image into K stages representing different shadow scales. In each stage, the flow matching model is used to learn the conditional probability path between the data and noise in the current stage, and shadow removal is performed on the current stage according to the conditional probability path; Scale prediction module: used to predict the shadow scale of the next stage based on the shadow removal result of the current stage using the next scale prediction mechanism, use the contrast heat map to guide the shadow removal of the next stage, and output the final document shadow-free image after all stages are processed.
9. A computer device, characterized in that: The computer device includes a processor and a memory coupled to the processor, wherein: The memory stores program instructions for implementing the document image shadow removal method according to any one of claims 1 to 7; The processor is configured to execute the program instructions stored in the memory to control the robot to perform a document image shadow removal method.
10. A storage medium, characterized in that: Program instructions executable by a processor are stored, and the program instructions are used to execute the document image shadow removal method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Floor tile image shadow removal method and device, computer equipment and storage medium
CN115546073A
Method for removing shadow of document image and related device
CN119107255A