Method for eliminating mirror surface highlight of text image, medium and program product
Through the use of semantic pixel combined with adaptive filtering network, the problem of difficulty in removing large-area highlights in text images is solved, and a greater receptive field and detail recovery is achieved, improving the accuracy and effect of highlight removal.
Patent Information
- Application Number
- CN202510290552.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, when processing text images containing rich text information and complex textures, it is difficult to effectively remove large-area specular highlights, and problems of inconsistent tones, missing details or distortion of color are often encountered.
Semantic pixel joint adaptive filtering network is adopted to predict branches and adaptive filtering branches through kernels, and realize semantic level filtering and pixel-level expansion filtering, enhancing the receptive field and restoring semantic information.
It effectively solves the problem of semantic information recovery under large-area highlights, overcomes the challenge of missing details at different scales, and achieves a more accurate highlight removal effect.
Smart Images

Figure CN120219199A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method, medium and program product for eliminating specular highlights in text images. Background Art
[0002] Specular highlights are extremely common in daily life, especially in the process of photography, where unwanted highlights are likely to appear. For example, when taking an ID card photo, the highlight area may obscure important information. In addition, in computer vision tasks such as image segmentation, text detection, and object detection, images with highlights also pose challenges. The generation of specular highlights is the result of the combined action of various factors, such as lighting conditions, object geometry, and surface texture. Therefore, removing specular highlights has become an important and challenging task in computer vision.
[0003] Traditional highlight removal methods are mostly based on the analysis of the physical and statistical characteristics of images, and use various technical means such as color space analysis, optimization, clustering, filtering, and lighting estimation. However, these methods have poor effects when dealing with complex images, often causing problems such as hue changes, incomplete removal of highlights, and black spots. The main reason is that these methods cannot capture high-level semantic information and do not fully utilize the effective information in weak highlight areas and non-highlight areas.
[0004] Highlight removal methods based on deep learning have achieved certain results in medical images, natural scene images, and specific object images. However, when dealing with text images containing rich text information and complex textures, the effects of existing methods are often unsatisfactory. Most of these methods focus on the detail restoration of highlight areas and their adjacent areas. Although they have good effects on removing small-area highlights, when dealing with large-area highlights, problems such as inconsistent hue, missing details, or color distortion usually occur. Removing large-area highlights on text images requires a larger receptive field to understand the structure and semantic information of the entire image, and then through local detail restoration, more accurate highlight removal can be achieved. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a method, medium and program product for eliminating specular highlights in text images, which can effectively solve the problem of semantic information restoration under large-area highlights.
[0006] To achieve the above technical objectives and reach the above technical effects, the present invention is realized through the following technical solutions:
[0007] In a first aspect, the present invention provides a method for eliminating specular highlights in text images, including:
[0008] Input the obtained high - light image into a pre - trained semantic pixel joint adaptive filtering network, so that the semantic pixel joint adaptive filtering network outputs a non - high - light image;
[0009] The semantic pixel joint adaptive filtering network includes a kernel prediction branch and an adaptive filtering branch; the adaptive filtering branch includes a first down - sampling unit, a semantic adaptive filtering unit, a first residual convolution unit, a first up - sampling unit, and a pixel dilation filtering unit arranged in sequence; the kernel prediction branch includes a second down - sampling unit, a second residual convolution unit, and a semantic - level kernel generation unit arranged in sequence, and a second up - sampling unit connected to the second down - sampling unit; the output of the semantic - level kernel generation unit is connected to the input of the semantic adaptive filtering unit; the output of the second up - sampling unit is connected to the pixel dilation filtering unit; the output of the first down - sampling unit is also connected to the input of the second down - sampling unit.
[0010] Combined with the first aspect, optionally, the training method of the semantic pixel joint adaptive filtering network includes:
[0011] Process the obtained non - high - light image to generate a high - light image;
[0012] Input the high - light image into a pre - constructed semantic pixel joint adaptive filtering network to obtain a non - high - light image;
[0013] Based on the obtained non - high - light image and the acquired non - high - light image, combined with a pre - designed loss function, train the semantic pixel joint adaptive filtering network to obtain a trained semantic pixel joint adaptive filtering network.
[0014] Combined with the first aspect, optionally, the first down - sampling unit and the second down - sampling unit have the same structure, both including a first down - sampling layer, a second down - sampling layer, and a third down - sampling layer arranged in sequence, which are used to extract low - level features, intermediate features, and high - level features in turn; the output of the third down - sampling layer in the first down - sampling unit is also connected to the input of the third down - sampling layer in the second down - sampling unit;
[0015] The first up - sampling unit includes a first up - sampling layer, a second up - sampling layer, and a third up - sampling layer arranged in sequence; the second up - sampling unit includes a fourth up - sampling layer, a fifth up - sampling layer, and a sixth up - sampling layer arranged in sequence; channel - spatial dual attention mechanism modules are provided in the fourth up - sampling layer, the fifth up - sampling layer, and the sixth up - sampling layer.
[0016] Combined with the first aspect, optionally, the second residual convolution unit generates a semantic - level feature map based on the received feature map ;
[0017] The semantic level kernel generation unit receives the semantic level feature map and converts the semantic level feature map into a prediction kernel through outer product operation . .
[0018] Combined with the first aspect, optionally, the semantic adaptive filtering unit filters the received feature map based on the received prediction kernel , and the filtering formula used is:
[0019] ;
[0020] where is the feature point at (x, y) on the filtered feature map, * represents the filtering operation, is the block centered at (x, y) with size in the feature map before filtering, is the coordinate position on the H*W plane of the feature map, represents the prediction kernel and the block centered at (x, y) with size in it.
[0021] Combined with the first aspect, optionally, the second upsampling unit performs upsampling processing on the received feature map to obtain a prediction kernel at the image level .
[0022] Combined with the first aspect, optionally, the filtering formula used by the pixel dilation filtering module is:
[0023] ;
[0024] where represents the feature point on the filtered feature map corresponding to different dilation factors , and represent the coordinate positions on the H*W plane of the feature map; is a coefficient, and its range is from to , represents the block at in the prediction kernel , is the size of the block at , represents the feature point on the feature map before filtering.
[0025] Combined with the first aspect, optionally, the mathematical expression of the pre-designed loss function is:
[0026] ;
[0027] Wherein:
[0028] ;
[0029] ;
[0030] ;
[0031] In the formula, is the total loss, is the mean absolute error loss, is the relativistic mean adversarial loss, is the perceptual loss, is the style loss, , , and are all weight parameters, is the sigmoid function, is the binary cross entropy; D is the discriminator; represents the feature map of the th layer of the semantic pixel joint adaptive filtering network, , , are respectively the number of channels, height and width of the feature map of the th layer of the semantic pixel joint adaptive filtering network, is the image after highlight removal; is the Gram matrix, which is used to calculate the correlation between the channels of the feature map, is the obtained highlight-free image; for the generator: y' and y are set to (1, 0); for the discriminator: y' and y are set to (0, 1).
[0032] In a second aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for eliminating specular highlights in text images according to any one of the first aspect is implemented.
[0033] In a third aspect, the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the method for eliminating specular highlights in text images according to any one of the first aspect is implemented.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] Through semantic-level filtering, the present invention achieves a larger receptive field and effectively solves the problem of semantic information restoration under large-area specular highlights; through pixel-level dilation filtering, it overcomes the challenge of missing details at different scales. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0037] Figure 1 is the overall flowchart of a method for eliminating specular highlights in text images according to an embodiment of the present invention;
[0038] Figure 2 is the structural diagram of a semantic-pixel joint adaptive filtering network according to an embodiment of the present invention.
[0039] Figure 3 is the structural schematic diagram of a semantic adaptive filtering module according to an embodiment of the present invention.
[0040] Figure 4 is the structural schematic diagram of a pixel dilation filtering module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0042] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0043] Embodiment 1
[0044] An embodiment of the present invention provides a method for eliminating specular highlights in a text image, including:
[0045] Input the obtained highlight image into a pre-trained semantic pixel joint adaptive filtering network, so that the semantic pixel joint adaptive filtering network outputs a highlight-free image;
[0046] Wherein, the semantic pixel joint adaptive filtering network includes a kernel prediction branch and an adaptive filtering branch; the adaptive filtering branch includes a first downsampling unit, a semantic adaptive filtering unit, a first residual convolution unit, a first upsampling unit, and a pixel dilation filtering unit arranged in sequence; the kernel prediction branch includes a second downsampling unit, a second residual convolution unit, and a semantic-level kernel generation unit arranged in sequence, and a second upsampling unit connected to the second downsampling unit; the output of the semantic-level kernel generation unit is connected to the input of the semantic adaptive filtering unit; the output of the second upsampling unit is connected to the pixel dilation filtering unit; the output of the first downsampling unit is also connected to the input of the second downsampling unit.
[0047] In a specific implementation manner of the embodiment of the present invention, the training method of the semantic pixel joint adaptive filtering network includes:
[0048] Process the obtained highlight-free image to generate a highlight image;
[0049] Input the highlight image into a pre-constructed semantic pixel joint adaptive filtering network to obtain a highlight-free image;
[0050] Based on the obtained highlight-free image and the obtained highlight-free image, combined with a pre-designed loss function, train the semantic pixel joint adaptive filtering network to obtain a trained semantic pixel joint adaptive filtering network.
[0051] In a specific implementation manner of the embodiment of the present invention, the first downsampling unit and the second downsampling unit have the same structure, and both include a first downsampling layer, a second downsampling layer, and a third downsampling layer arranged in sequence, which are used to extract low-level features, intermediate features, and high-level features in sequence; the output of the third downsampling layer in the first downsampling unit is also connected to the input of the third downsampling layer in the second downsampling unit;
[0052] The first upsampling unit includes a first upsampling layer, a second upsampling layer, and a third upsampling layer arranged in sequence; the second upsampling unit includes a fourth upsampling layer, a fifth upsampling layer, and a sixth upsampling layer arranged in sequence; channel-spatial dual attention mechanism modules are provided in the fourth upsampling layer, the fifth upsampling layer, and the sixth upsampling layer.
[0053] In a specific implementation manner of the embodiment of the present invention, the second residual convolution unit generates a semantic-level feature map based on the received feature map ;
[0054] The semantic-level kernel generation unit receives the semantic-level feature map , and converts the semantic-level feature map into a prediction kernel .
[0055] In a specific implementation manner of the embodiment of the present invention, the semantic adaptive filtering unit filters the received feature map based on the received prediction kernel , and the adopted filtering formula is:
[0056] ;
[0057] Wherein, is the feature point at (x, y) on the filtered feature map, * represents the filtering operation, is the block in the feature map centered at (x, y) with a size of , is the coordinate position on the H*W plane in the feature map, represents the prediction kernel centered at (x, y) with a size of .
[0058] In a specific implementation manner of the embodiment of the present invention, the second upsampling unit performs upsampling processing on the received feature map to obtain an image-level prediction kernel .
[0059] In a specific implementation manner of the embodiment of the present invention, as Figure 4 shown, the filtering formula adopted by the pixel dilation filtering module is:
[0060] ;
[0061] Wherein, represents the feature point on the filtered feature map corresponding to different dilation factors , and represent the coordinate positions on the H*W plane in the feature map; is a coefficient, and its range is from to , represents the block at in the prediction kernel , is the size of the block at , Denote the feature points on the feature map before filtering.
[0062] In a specific implementation manner of the embodiment of the present invention, the mathematical expression of the pre-designed loss function is:
[0063] ;
[0064] Where:
[0065] ;
[0066]
[0067]
[0068] In the formula, is the total loss, is the mean absolute error loss, is the relativistic mean adversarial loss, is the perceptual loss, is the style loss, , , and are all weight parameters, is the sigmoid function, is the binary cross-entropy; D is the discriminator; Denote the feature map of the th layer of the semantic pixel joint adaptive filtering network, , , are respectively the number of channels, height and width of the feature map of the th layer of the semantic pixel joint adaptive filtering network, is the image after highlight removal; is the Gram matrix, used to calculate the correlation between the channels of the feature map, is the obtained image without highlight; for the generator: y' and y are set to (1, 0); for the discriminator: y' and y are set to (0, 1).
[0069] Next, a specific implementation manner is used to detail the method for eliminating the specular highlight in the text image in the embodiment of the present invention.
[0070] As Figure 1 shown, the specific implementation steps are as follows:
[0071] Step 1, an image without highlights is obtained by camera shooting or online downloading. Then, the lighting system of Unity3D and a custom shader are used to simulate the highlight effects under various lighting conditions, thereby constructing a dataset and generating highlight images.
[0072] Taking the highlight simulation on a bank card as an example, a directional light source is used to simulate sunlight, and the color of the light source is set to a warm tone, approaching the ambient lighting effect in the real scene. Adjust the intensity of the light source as needed to achieve the required brightness. The shadow type and intensity are set to soft shadows with medium intensity to ensure smooth shadow edges and enhance the realism of the scene. The indirect lighting intensity is set between 1.0 and 2.0 to ensure that reflections and indirect lighting play an important role in the highlights and overall lighting effects on the object surface. The rendering mode and culling mask are configured with high priority, and the culling mask is adjusted to ensure that the lighting only affects the layer where the bank card is located, thereby reducing unnecessary performance overhead and ensuring the accuracy of specular highlights.
[0073] In the bank card shader, the following property settings are as follows:
[0074] MainTex: Represents the main texture of the bank card, providing the basic appearance and visual details of the card surface.
[0075] Smoothness: Set to 0.6, defining the smoothness of the bank card surface and affecting the glossiness of the surface.
[0076] Metallic: Set to 0.8, indicating that the metallic printed part on the card has highly metallic reflection characteristics.
[0077] OcclusionStrength: Set to 1, enhancing the intensity of ambient occlusion and simulating the shadow effects in crevices and detail areas.
[0078] Emissive: Used to define the self-illuminating color, enabling certain parts on the bank card (such as security marks or decorative elements) to emit light even under low light conditions.
[0079] ReflectionTex: Cubemap texture is adopted for reflection probes to enhance the specular reflection effect on the card surface and improve the authenticity of highlights.
[0080] Anisotropy: Set to 0.5, used to simulate anisotropic reflection and show directional reflection effects such as stripes or brush strokes on the card surface.
[0081] Step 2, construct a semantic pixel joint adaptive filtering network, and the structure of the semantic pixel joint adaptive filtering network is as Figure 2As shown. The semantic pixel joint adaptive filtering network includes two branches: a kernel prediction branch and an adaptive filtering branch. In the first half of these two branches, semantic-level kernels are predicted and filtered; in the second half, pixel-level kernels are predicted and filtered. There are interactive links between the two branches, providing information to each other. The adaptive filtering branch provides multi-level features for the kernel prediction branch, while the kernel prediction branch predicts the dynamic kernels required by the adaptive filtering branch. The highlight image is respectively input into
[0082] As Figure 2 shown, the adaptive filtering branch includes a first downsampling unit, a semantic adaptive filtering unit, a first residual convolution unit, a first upsampling unit, and a pixel dilation filtering unit arranged in sequence;
[0083] The first layer is Conv + BN + ReLU, where Conv represents the convolution operation, BN represents the normalization operation, and ReLu represents the activation function used. The main function of this layer is downsampling + low-level feature extraction (edges, textures, etc.), initially separating the brightness features of the highlight and non-highlight regions. The features change from [512, 512, 3] -> [256, 256, 64].
[0084] The second layer is Pool+Conv+BN+ReLU, and its main function is downsampling + intermediate feature extraction. Adding the pooling operation will retain the significant features (such as edges, textures) within the local area, suppress noise and small perturbations, and improve the robustness of the model to small changes in the input. The features change from [256, 256, 64] -> [128, 128, 128].
[0085] The third layer is Pool+Conv+BN+ReLU, which further compresses the spatial dimension while extracting high-level features. At this time, the features represent deep semantic-level information. The features change from [128, 128, 128] -> [64, 64, 256].
[0086] The first layer, the second layer, and the third layer together constitute the first downsampling unit.
[0087] The fourth layer: The received feature map is filtered through the designed semantic-level adaptive filtering module to remove highlights and restore semantic information. The features change from [64, 64, 256] -> [64, 64, 256].
[0088] The fifth layer is Residual Block×2: The filtered feature map is processed by the Residual Block to enhance the modeling ability of the differences between the highlight area and the non-highlight area, while retaining the basic features and avoiding the vanishing gradient problem caused by multi-layer processing. The features change from [64, 64, 256] -> [64, 64, 256].
[0089] The sixth layer is DeConv+BN+ReLU: The upsampling initially restores the resolution, and combines with the intermediate features of the encoder to restore the local brightness distribution. Among them, DeConv is the transposed convolution operation. The features change from [64, 64, 256] -> [128, 128, 128].
[0090] The seventh layer is DeConv+BN+ReLU: Continuing the upsampling to refine the transition area by combining with the low-level details. The features change from [128, 128, 128] -> [256, 256, 64].
[0091] The eighth layer is DeConv+BN+ReLU: Continuing the upsampling to restore the original size and output the highlight-free feature with the complete texture color. The features change from [256, 256, 64] -> [512, 512, 3].
[0092] The ninth layer: Adjust and restore the missing details in the image through the designed pixel-level dilation filtering module. The features change from [512, 512, 3] -> [512, 512, 3].
[0093] The kernel prediction branch includes a second downsampling unit, a second residual convolution unit, and a semantic-level kernel generation unit arranged in sequence, and a second upsampling unit and a pixel-level kernel generation unit arranged in sequence;
[0094] The kernel prediction branch is consistent with the adaptive filtering branch in the downsampling stage (semantic-level kernel prediction). However, in the upsampling stage (pixel-level kernel prediction), the CBMA (Channel-Spatial Dual Attention Mechanism) module is introduced to refine the feature map by applying the attention mechanism, emphasizing the key features and suppressing the irrelevant features in the channel and spatial dimensions. This is crucial for generating accurate dynamic kernels, ensuring precise detail restoration, and eliminating highlight artifacts to generate more natural images. Finally, dilation filtering is performed through the pixel-level dilation filtering module to adjust and restore the details in the image.
[0095] Step 3, design the semantic adaptive filtering module
[0096] In the specific data processing process, the highlight image is used as the input and downsampled to the depth feature size of [64, 64, 256] through the kernel prediction branch. The semantic-level feature map of 3×3 is generated through Residual Block+Conv [64, 64, 6, 256], and the semantic-level adaptive filtering module combines the concept of separable kernel estimation to further reduce the number of parameters. The features are transformed into the prediction kernel through the outer product operation [64, 64, 6, 256] into the prediction kernel [64, 64, 9,256]:
[0097]
[0098] Among them, and come from the prediction kernel. Specifically, a 1*1 feature vector is taken on the channel and then the first half is taken as , and the second half is taken as , as shown in Figure 3 . represents the outer product operation is the prediction kernel (i.e., the semantic-level kernel) of size s at (x, y) in the feature map. This method can reduce the number of parameters from s 2 to 2s. Specifically, when filtering at (x, y), the [x,y,25,c] prediction kernel at (x, y) on the same channel of the prediction kernel [64, 64, 25,256] is selected to filter this place:
[0099]
[0100] Among them, is the feature point at (x, y) in the filtered feature map, and * represents the filtering operation is the block of size s centered at (x, y) in the feature map before filtering
[0101] Step 4, design the pixel dilation filtering module
[0102] Although the highlight can be eliminated and the semantic information can be restored after semantic filtering, a lot of detailed information will be lost, and the lack of various details usually shows different sizes. Therefore, by designing pixel-level dilation filtering, the deep features are upsampled to the image level for dilation filtering, and the detailed loss of different scales can be well restored while avoiding artifacts. The pixel dilation filtering is defined as follows
[0103] The filtering formula adopted by the pixel dilation filtering module is
[0104] ;
[0105] Among them, represents the feature points on the filtered feature map corresponding to different dilation factors and and represent the coordinate positions on the H*W plane in the feature map; is a coefficient, and its range is from to , represents the block at in the prediction kernel is the size of the block at and represents the feature points on the feature map before filtering.
[0106] Specifically, after the features filtered by semantics are processed by the second upsampling unit, a feature map F[512, 512, 3] at the image level is obtained. Subsequently, the image-level kernel [512, 512, 9, 3] is dynamically obtained through the kernel prediction branch to perform dilation filtering on it, and the dilation factors are selected as l = 1, 2, 3. Three filtered features F1[512, 512, 3], F2[512, 512, 3], and F3[512, 512, 3] are obtained. Finally, a Conv operation is used to simply fuse these three features F1, F2, and F3 to obtain the final highlight-free image [512, 512, 3].
[0107] Step 5, train the network by designing a loss function. The loss function is as follows:
[0108] 1. Relativistic average adversarial loss, which not only considers the scores of the discriminator for the generated image and the real image, but also considers their relative differences, and is defined as follows:
[0109] ;
[0110] Among them, is the sigmoid function, and BCE(*) is the binary cross-entropy. For the generator, (y', y) is set to (1, 0), and for the discriminator, it is set to (0, 1). D is the discriminator. This relative scoring strategy can better capture subtle differences and promote the generation of higher-quality images.
[0111] 2. Perceptual loss, which captures the high-level features of the image and ensures that the estimated highlight-free image retains important content and structural information, and is defined as follows:
[0112] ;
[0113] Among them, represents the feature of the i-th layer of the pre-trained VGG-19 network, and Ci, Hi, and Wi are the dimensions of the feature.
[0114] 3. Style loss, which calculates the difference between the Gram matrices of the generated image and the target image, effectively captures and preserves textures, and is defined as follows:
[0115] ;
[0116] Among them, represents the feature of the i-th layer of the same pre-trained network used for perceptual loss, is the Gram matrix, which is used to calculate the correlation between the channels of the feature map. This loss helps to restore the texture details lost due to highlights, making the image without highlights more visually natural and coherent.
[0117] The total loss is defined as follows:
[0118]
[0119] Among them the loss is also called the mean absolute error, , , and are weight parameters. In our experiments, we set = 1, = 0.1, = 0.1 and = 0.5.
[0120] Example 2
[0121] Based on the same inventive concept as in Example 1, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for eliminating specular highlights in text images described in any one of Example 1.
[0122] Example 3
[0123] Based on the same inventive concept as in Example 1, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the method for eliminating specular highlights in text images described in any one of Example 1.
[0124] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0125] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0128] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present invention and without departing from the spirit and scope protected by the present invention's claims, can still make many forms, and all of these fall within the protection scope of the present invention.
[0129] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for eliminating specular highlights of a text image, characterized in that: include: Inputting the acquired highlight image into a pre-trained semantic pixel joint adaptive filtering network, so that the semantic pixel joint adaptive filtering network outputs a highlight-free image; The semantic pixel joint adaptive filtering network includes a kernel prediction branch and an adaptive filtering branch; the adaptive filtering branch includes a first downsampling unit, a semantic adaptive filtering unit, a first residual convolution unit, a first upsampling unit and a pixel dilation filtering unit arranged in sequence; the kernel prediction branch includes a second downsampling unit, a second residual convolution unit and a semantic level kernel generation unit arranged in sequence, and a second upsampling unit connected to the second downsampling unit; the output of the semantic level kernel generation unit is connected to the input of the semantic adaptive filtering unit; the output of the second upsampling unit is connected to the pixel dilation filtering unit; the output of the first downsampling unit is also connected to the input of the second downsampling unit.
2. The method for eliminating specular highlights of a text image according to claim 1, characterized in that: The training method of the semantic pixel joint adaptive filtering network includes: Processing the acquired non-highlight image to generate a highlight image; Inputting the highlight image into a pre-built semantic pixel joint adaptive filtering network to obtain a highlight-free image; Based on the obtained non-highlight image and the acquired non-highlight image, combined with a pre-designed loss function, the semantic pixel joint adaptive filtering network is trained to obtain a trained semantic pixel joint adaptive filtering network.
3. The method for eliminating specular highlights of a text image according to claim 1, characterized in that: The first downsampling unit and the second downsampling unit have the same structure, and both include a first downsampling layer, a second downsampling layer and a third downsampling layer which are arranged in sequence, and the three are used to extract low-level features, intermediate features and high-level features in sequence; The output of the third downsampling layer in the first downsampling unit is also connected to the input of the third downsampling layer in the second downsampling unit; The first upsampling unit includes a first upsampling layer, a second upsampling layer and a third upsampling layer arranged in sequence; the second upsampling unit includes a fourth upsampling layer, a fifth upsampling layer and a sixth upsampling layer arranged in sequence; the fourth upsampling layer, the fifth upsampling layer and the sixth upsampling layer are all provided with a channel-spatial dual attention mechanism module.
4. The method for eliminating specular highlights of a text image according to claim 1, characterized in that: The second residual convolution unit generates a semantic level feature map based on the received feature map ; The semantic level kernel generating unit receives the semantic level feature map , and the semantic level feature map is transformed into Convert to prediction kernel .
5. The method for eliminating specular highlights of a text image according to claim 4, characterized in that: The semantic adaptive filtering unit is based on the received prediction kernel The received feature map is filtered using the following filtering formula: ; in, is the feature point at (x, y) on the feature map after filtering, * indicates the filtering operation, is the feature map before filtering, centered at (x, y) and of size The block, is the coordinate position on the H*W plane in the feature map, Represents the prediction kernel The center of the image is (x, y) and the size is of blocks.
6. The method for eliminating specular highlights of text images according to claim 1, characterized in that: The second upsampling unit performs upsampling processing on the received feature map to obtain an image-level prediction kernel .
7. The method for eliminating specular highlights of text images according to claim 6, characterized in that: The filtering formula used by the pixel expansion filtering module is: ; in, Indicates different expansion factors The corresponding feature points on the filtered feature map, and Represents the coordinate position on the H*W plane in the feature map; is a coefficient ranging from arrive , Represents the prediction kernel middle The block at for The size of the block at Represents the feature points on the feature map before filtering.
8. The method for eliminating specular highlights of text images according to claim 1, characterized in that: The mathematical expression of the pre-designed loss function is: ; in: ; ; ; In the formula, is the total loss, is the mean absolute error loss, is the relativistic average adversarial loss, is the perceived loss, is the style loss, , , and are weight parameters, is the sigmoid function, is the binary cross entropy; D is the discriminator; Represents the semantic pixel joint adaptive filtering network The feature map of the layer, , , They are the semantic pixel joint adaptive filtering network The number of channels, height and width of the feature map of the layer, This is the image after highlight removal; is the Gram matrix, which is used to calculate the correlation between feature map channels. is the obtained image without highlight; for the generator: y' and y are set to (1, 0); for the discriminator: y' and y are set to (0, 1).
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for eliminating specular highlights of a text image as described in any one of claims 1 to 8 is implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for eliminating specular highlights of a text image according to any one of claims 1 to 8 is implemented.