Visible watermark removing method and device for image, equipment and medium

By combining the depth segmentation model and image diffusion model, segmentation and denoising watermarks, the problem of poor watermark removal in the prior art is solved, and a more efficient and better watermark removal effect is achieved.

CN119963435APending Publication Date: 2025-05-09BEIJING SINOVOICE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411970825.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art has poor results in removing visible watermarks, especially when the watermark area is small, the color is light, and the transparency is high, the extraction effect is poor, and the training of GAN or diffusion models requires huge samples and poor results.

Method used

Using a method of combining the depth segmentation model with the image diffusion model, the watermark is first segmented through the depth segmentation model to generate a watermark mask image, and then input it into the image diffusion model as a control condition to perform watermark denoising processing.

Benefits of technology

Decomposing the watermark removal into two independent optimization stages reduces the requirements for the number of training samples and improves the effect and quality of watermark removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963435A_ABST
    Figure CN119963435A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device and equipment for removing a visible watermark of an image and a medium, and belongs to the field of visible watermark removal, and the method comprises the steps: carrying out the angle correction of an image with a watermark, and obtaining a first image; performing segmentation processing on the first image through an image segmentation model to obtain watermark image data; inputting the first image and the watermark image data into a trained image diffusion model for processing to obtain a watermark mask image; inputting the obtained watermark mask image as a control condition into an encoder of an image diffusion model; and performing watermark de-noising processing on the first image based on the control condition through the encoder to obtain a watermark-free image. A watermark is segmented through an image segmentation model, and then a watermark segmentation instance is used as an input control condition to assist an image diffusion model to carry out watermark denoising to generate a watermark-removed image. According to the method, the watermark removal is decomposed into two stages, and the two stages can be independently optimized, so that the requirement on the number of training samples is greatly reduced, and the watermark removal effect is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visible watermark removal, and in particular to a method, device, equipment and medium for removing visible watermarks from images. Background Art

[0002] In the field of visible watermark removal, watermarks are usually overlaid on the document content in the form of semi-transparent text, pictures or patterns, which may interfere with the reading experience, especially when the watermark is large or overlaid on key content.

[0003] The existing technical solution obtains the rectangular target frame of the watermark by extracting the rectangular area. Due to the characteristics of the watermark being small in area, light in color, and high in transparency, the extraction effect of the existing technical solution is poor.

[0004] Existing technologies also include using GAN or diffusion models to directly perform end-to-end watermark removal. This solution requires a large number of samples for training and has no explicit watermark region to assist diffusion, resulting in poor training results. Summary of the invention

[0005] To solve the above problems, the present application provides a method, device, equipment and medium for removing visible watermarks from images.

[0006] A first aspect of an embodiment of the present application provides a method for removing a visible watermark of an image, comprising: Performing angle correction on the watermarked image to obtain a first image; Performing segmentation processing on the first image by using a segmentation module of a trained image segmentation model to obtain watermark image data; Inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents a watermark mask image; Inputting the processing result as a control condition into the encoder of the trained image diffusion model; The encoder performs watermark denoising processing on the first image based on the control condition to obtain a watermark-free image.

[0007] Optionally, the performing segmentation processing on the first image by a segmentation module of the trained image segmentation model to obtain watermark image data specifically includes: Correcting the first image by a correction module of a trained image segmentation model to obtain a corrected first image; wherein the segmentation module includes: a detection branch and a segmentation branch; Detecting the watermark position of the corrected first image by using the detection branch of the trained image segmentation model to obtain a watermark position result; Generating mask data of the corrected first image through the segmentation branch of the trained image segmentation model; A watermark instance mask is obtained according to the watermark position result and the mask data, and the watermark instance mask is used as watermark image data.

[0008] Optionally, the step of inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result specifically includes: Inputting the watermark image data into the encoder in the trained image diffusion model for encoding to obtain a watermark mask in a latent space; Encoding the first image through the encoder in the trained image diffusion model to obtain a first image encoding result in a latent space; The watermark mask and the first image encoding result are processed by a conditional generation module in the trained image diffusion model to obtain a processing result.

[0009] Optionally, performing watermark denoising on the first image based on the control condition by the encoder to obtain a watermark-free image specifically includes: Encoding the first image by the encoder to obtain a first image encoding result in a latent space; Detecting the encoding result of the first image by the encoder to obtain a watermark feature of the first image; Performing watermark removal processing on the first image encoding result according to the watermark feature and the control condition to obtain a watermark-free image encoding; The watermark-free image code is decoded by a decoder in the trained image diffusion model to generate a watermark-free image.

[0010] Optionally, performing watermark removal processing on the first image encoding result by using the watermark feature and the control condition to obtain a watermark-free image encoding specifically includes: The control condition is used as a watermark removal guide, and the first image encoding result is subjected to an i-th watermark removal process by using a preset denoising algorithm and the watermark feature to obtain an i-th watermark-free image encoding; i≥1; The i-th watermark-free image code is used as the i+1-th watermark removal processing input, and the i+1-th watermark removal processing is performed in combination with the watermark feature and the control condition until the watermark removal in the first image coding result is completed to obtain the watermark-free image code.

[0011] Optionally, performing angle correction on the watermarked image to obtain the first image specifically includes: The watermarked image is angle-corrected by a preset method to obtain a first image; wherein the preset method includes: a Hough transform method, an edge detection method or a calculation horizontal projection method.

[0012] Optionally, it also includes: Acquire a watermark image data set; the watermark image data set includes: a plurality of watermark sample images; Acquire a watermark-free image dataset; the watermark-free image dataset includes: a plurality of watermark-free sample images; Adjusting at least one of the size, position and transparency of the watermark sample image to obtain an adjusted watermark sample image; Adding the adjusted watermark sample image to a non-watermark sample image, and adding at least one of size information, position information and transparency information of the adjusted watermark sample image to the annotation information of the non-watermark sample image, to obtain a watermarked sample image; Integrate multiple watermarked sample images to obtain a watermarked image dataset; The watermarked image data set is used to train an image segmentation model and an image diffusion model respectively to obtain a trained image segmentation model and a trained image diffusion model.

[0013] A second aspect of the embodiment of the present application provides a device for removing visible watermarks from an image, comprising: a correction module, a watermark extraction module, a condition control module, and a watermark removal module; The correction module is used to perform angle correction on the watermarked image to obtain a first image; The watermark extraction module is used to perform segmentation processing on the first image through the segmentation module of the trained image segmentation model to obtain watermark image data; The condition control module is used to input the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents the watermark mask image, and the processing result is input into the encoder of the trained image diffusion model as a control condition; The watermark removal module is used to perform watermark denoising on the first image based on the control condition through the encoder to obtain a watermark-free image.

[0014] A third aspect of an embodiment of the present application provides an electronic device, including a memory and a processor, wherein: The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in a method for removing a visible watermark from an image in any of the above solutions.

[0015] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is used to store a computer-readable program or instruction, which, when executed by a processor, can implement the steps of a method for removing a visible watermark from an image in any of the above-mentioned schemes.

[0016] The technical solution provided by the embodiment of the present application is applied. The solution of the present invention is based on a visible watermark removal method that combines a deep segmentation model with an image diffusion model. In the first stage, the watermark in the document image is segmented into watermark image data using a deep segmentation model. A watermark mask image is generated based on the document image and the watermark image data. In the second stage, the watermark mask image is used as an input control condition to assist the image diffusion model in watermark denoising and generate a de-watermarked image. In this way, watermark removal is decomposed into two stages, and the two stages can be optimized independently, which greatly reduces the number of training samples required, and the effect of watermark removal is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0018] Figure 1 is a flowchart of a method for removing visible watermarks from an image provided in an embodiment of the present invention; Figure 2 is a flowchart of the process steps of removing visible watermarks from document images and restoring the original images in an embodiment of the present invention; Figure 3 is a flowchart of the steps of performing segmentation processing on a first image by an image segmentation model provided in an embodiment of the present invention to obtain watermark image data; Figure 4 is a flowchart of the steps of the processing process of the image diffusion model provided in an embodiment of the present invention; Figure 5 Schematic diagram of the structure of the U-Net network provided in an embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of each network module in the U-Net network provided in an embodiment of the present invention; Figure 7 is a flowchart of the steps of the image diffusion model watermark removal process provided in an embodiment of the present invention; Figure 8 It is a flowchart of the steps of the UNet model watermark removal process provided in an embodiment of the present invention; Fig. 9is a flowchart of steps for obtaining watermark-free image encoding provided in an embodiment of the present invention; Fig.10 is a schematic diagram of an angle correction process of a document image provided in an embodiment of the present invention; Fig.11 is a flowchart of the steps of the model training process provided in an embodiment of the present invention; Fig.12 It is a structural block diagram of a device for removing visible watermark of an image provided in an embodiment of the present invention; Fig.13 It is a hardware structure block diagram of an electronic device provided in each embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] Reference Figure 1 As shown, a flowchart of a method for removing visible watermarks from an image is shown. The method can be applied to the field of visible watermark removal, such as Figure 1 As shown, the following steps may be specifically included: Step S101, performing angle correction on the watermarked image to obtain a first image; For example, document correction may be to correct an input document image with a tilt to a frontal image.

[0021] In one scenario, the visible watermark may be a visible watermark on a natural scene image or a document image, such as a company logo, a text watermark, a dot watermark, etc.

[0022] Step S102: segment the first image through the segmentation module of the trained image segmentation model to obtain watermark image data; the image segmentation model can be a yolo model or a DeepLab series segmentation model.

[0023] For example, the image segmentation of the segmentation module of the yolo model can be to perform pixel segmentation on the target of interest in the image, and the pixel value belonging to this target is 1, and the pixel value not belonging to this target is 0, so as to obtain a watermark instance; the present invention belongs to the common semantic segmentation problem, mainly referring to the binarization of visible watermark images.

[0024] Step S103, inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents a watermark mask image; In one example, a controlnet module that is used to control a pre-trained large diffusion model and can support additional input conditions can be used to process the first image and watermark image data to obtain a processing result, i.e., a watermark mask image, which serves as a control condition to guide the image diffusion model to perform de-watermarking processing.

[0025] Step S104, inputting the processing result as a control condition into the encoder of the trained image diffusion model; In this embodiment, image segmentation is performed according to the YOLO model to obtain a watermark instance, i.e., watermark image data; the first image and the watermark instance are input into the controlnet module for processing to generate a watermark mask image, and the watermark mask image description is input into the image diffusion model as a control condition.

[0026] Step S105: Perform watermark denoising processing on the first image based on the control condition by the encoder to obtain a watermark-free image.

[0027] In this embodiment, the watermarked image is angle-corrected to obtain a first image; the first image is segmented by a segmentation module of a trained image segmentation model to obtain watermarked image data; the first image and the watermarked image data are input into a trained image diffusion model encoder for processing to obtain first image latent variables and watermark segmentation latent variables, respectively.

[0028] The processing results of the first image latent variables and the watermark segmentation latent variables are input into the core module Unet of the image diffusion model. The watermark latent variables are used as control conditions to perform watermark denoising on the first image latent variables to obtain the latent variables of the watermark-free image. Finally, the latent variables of the watermark-free image are decoded by the decoder of the image diffusion model to obtain the final watermark-free image.

[0029] In some examples, a segmentation module of an image segmentation model performs segmentation processing on the first image to obtain the watermark image data; Inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a watermark mask image; The segmentation module of the image segmentation model performs segmentation processing on the first image to obtain a watermark feature text description; the watermark feature text description may include a vector of feature information such as size, position, transparency, etc. of the watermark; Using the watermark mask image as a control condition; The encoder of the image diffusion model is based on the control condition and takes the text description as a watermark removal guide, and performs watermark denoising on the first image through the DDIM denoising algorithm to obtain a watermark-free image.

[0030] In some examples, the cross-attention layer of the encoder of the image diffusion model uses text description information to guide the global direction of image generation. In this way, the model can pay attention to both local details and global context when restoring the image, thereby improving the quality of the restored image.

[0031] In this embodiment, a watermark condition control mechanism is introduced, which uses the previously extracted watermark information as an auxiliary signal to guide the diffusion model to restore the image more accurately.

[0032] For example, Figure 2 FIG. 1 shows a process of removing a visible watermark from a document image and restoring the original image, which may specifically include the following steps: The first step is to correct the input watermarked document image to eliminate the text tilt caused by image rotation. The second step is to use the deep segmentation network yolo to segment the watermark pixels of the rectified image and extract the watermark image; In the third step, controlnet uses the extracted watermark image as the input control condition assisted diffusion model UNet to perform watermark denoising and generate the de-watermarked image.

[0033] This solution can remove visible watermarks from document images and restore the original image. The watermark area must be removed cleanly, the non-watermark area must remain the same as the original image, and the original image of the area blocked by the watermark must be restored without affecting the reading experience.

[0034] The technical solution of the embodiment of the present application is a visible watermark removal method based on the combination of the deep segmentation model yoLO and the UNet diffusion model. In the first stage, the deep segmentation model is used to segment the watermark in the document image to obtain an instance of the visible watermark. In the second stage, the instance of the explicit watermark segmented in the first stage is used as an input control condition to assist the UNet diffusion model in watermark denoising and generate a de-watermarked image. In this way, the watermark removal is decomposed into two stages, and the two stages can be optimized independently, which greatly reduces the number of training samples required, and the effect of watermark removal is better. Among them, the visible watermark refers to a visible watermark, such as Figure 8 The CCTV watermark in the figure is opposite to the implicit watermark, which is marked in the data. The scheme of the present invention is used to process this watermark, and the scheme of the present application is applied to the removal of the display watermark.

[0035] In one example, referring to Figure 3As shown, a flowchart of the steps of performing segmentation processing on the first image by the image segmentation model to obtain watermark image data is shown, which may specifically include the following steps: Step S301, correcting the first image by a correction module of a trained image segmentation model to obtain a corrected first image; wherein the segmentation module includes: a detection branch and a segmentation branch; Step S302, detecting the watermark position of the corrected first image by using the detection branch of the trained image segmentation model to obtain a watermark position result; Step S303, generating mask data of the corrected first image through the segmentation branch of the trained image segmentation model; Step S304: obtaining a watermark instance mask according to the watermark position result and the mask data, and using the watermark instance mask as watermark image data.

[0036] In one scenario, the visual significance of the watermark in the image is very low, usually manifested by small area, light color, high transparency and other characteristics. Therefore, the difference between the image with watermark and the image without watermark is often very small, resulting in low differentiation. The traditional approach is to simply extract the rectangular target box of the watermark, but we found that if only the rectangular area is extracted, this simple and violent image processing technology will leave obvious traces of removal in the image after the watermark is removed, and in most cases, satisfactory results cannot be obtained. This solution extracts the watermark instance and separates it from the original image, so the extraction method is more refined.

[0037] The single-stage YOLACT processes the detection and segmentation parts in parallel. The detection branch is responsible for locating the instance, while the segmentation branch is responsible for generating the corresponding mask. These two parts generate the final instance mask through a linear combination. The advantage of this method is that it decomposes the instance segmentation task into two complementary parts, which not only improves the efficiency of segmentation, but also ensures a high segmentation quality.

[0038] The technical solution of the embodiment of the present application extracts watermarks through an image segmentation model and uses it as an auxiliary control to restore the image. It can provide stable performance when processing complex scenes and can effectively cope with common challenges in watermark extraction, ensuring that the processing speed is improved while maintaining the segmentation accuracy.

[0039] The image segmentation model extracts watermarks not by traditional rectangular area extraction, but by extracting watermark entities, which facilitates the subsequent removal of more complex watermarks.

[0040] In one example, referring to Figure 4 As shown, a flowchart of the steps of the processing process of the image diffusion model is shown, which may specifically include the following steps: Step S401, inputting the watermark image data into the encoder in the trained image diffusion model for encoding to obtain a watermark mask in a latent space; Step S402, encoding the first image through the encoder in the trained image diffusion model to obtain a first image encoding result in a latent space; Step S403, the watermark mask and the first image encoding result are processed by the conditional generation module in the trained image diffusion model to obtain a processing result. The conditional generation module can be a conditional control network ControlNet, which is used to extract the watermark as an auxiliary control signal to guide the model to focus on the area where the watermark is located for restoration.

[0041] For example, through the Latent Diffusion diffusion model, using the U-Net network, such as Figure 5 As shown in the figure, a schematic diagram of the structure of the U-Net network is shown. Its network structure is "U" shaped, including a contraction path (encoder) and a symmetrical expansion path (decoder), as well as jump connections. The encoder gradually reduces the spatial dimension of the image while increasing the number of feature channels to capture the contextual information of the image content; the decoder gradually restores the spatial dimension and detail information of the image. The key jump connection directly connects the high-resolution features of the encoder to the corresponding layers of the decoder, allowing the network to combine position information and rich features in each decoding step, thereby effectively utilizing the local and overall information of the image for accurate pixel-level prediction.

[0042] In the U-Net network structure, the input layer (Input Image Tile): the input is an image block of size 572x572; the encoder part (downsampling): through a series of convolutional layers (conv 3x3, ReLU) and maximum pooling layers (maxpool 2x2) to gradually reduce the spatial dimension of the image, while increasing the number of feature channels; the size of the feature map gradually decreases, for example, from 572x572 to 28x28, while the number of channels increases from 1 to 1024; Bottleneck layer: At the bottom of the network, the feature map has the smallest size and the largest number of channels. This is usually the part with the richest feature extraction. Decoder part (up-sampling): The spatial dimensions of the image are gradually restored through up-convolutional layers (up-conv2x2) and convolutional layers (conv3x3, ReLU); during the up-sampling process, the size of the feature map gradually increases, for example, from 28x28 to 388x388; Output Segmentation Map: Finally, a 1x1 convolutional layer is used to convert the feature map into the required number of output channels. Here, the output is a segmentation map of size 388x388 with 2 channels, representing two different categories. Skip Connections: There may be skip connections between the encoder and decoder (not explicitly shown in the figure), which help to recover image details in the decoder; Among them, color coding, different colors are used in the figure to represent different operations: the blue arrow conv 3x3 represents 3x3 convolution and ReLU activation; the gray arrow copy and crop represents copy and crop operations; the red arrow max pool2x2 represents 2x2 maximum pooling; the green arrow up-conv represents 2x2 up convolution; the dark green arrow conv1x1 represents 1x1 convolution; In addition to the U-Net network, in each small module, such as Figure 6 As shown in the figure, a schematic diagram of the structure of the network module is shown, and an attention mechanism is added, including a self-attention layer (Self-Attention), a cross-attention layer (Cross-Attention) and a feedforward network (ReLU). The self-attention layer captures the relationship between local features by calculating the interaction between input features, while the cross-attention layer uses text information to guide the global direction of image generation. The cross-attention layer uses conditional input C to calculate the attention to the input x. In this way, the model can pay attention to local details and global context at the same time when restoring the image, thereby improving the quality of the restored image.

[0043] A conditional control network ControlNet is introduced. First, the input watermark instance is preprocessed as a control condition, and then the watermarked image and the watermark control condition are encoded into the latent space respectively. Then, the convolutional layer of the U-Net network is cloned to extract the features related to the control condition. By fusing the conditional feature map with the intermediate feature map of the input generative model, a Mask is added and the information of the watermark control condition is injected into the generation process. This fusion process enables the generative model to adjust the image restoration behavior according to the watermark control condition, thereby ensuring that the restoration result is closely related to the input image and the watermark control condition.

[0044] The extracted watermark is used as an auxiliary control signal to guide the model to focus on the watermark area for restoration. The technical solution of the embodiment of the present application uses the extracted watermark as an auxiliary control signal through the controlnet module to guide the image diffusion model to focus on the area where the watermark is located for restoration.

[0045] In one example, referring to Figure 7 As shown, a flowchart of the process of image diffusion model watermark removal is shown, which may specifically include the following steps: Step S701, encoding the first image by the encoder to obtain a first image encoding result in a latent space; For example, Latent Diffusion Models are a class of generative models based on diffusion processes that simulate the evolution of data distribution in a latent space. These models generate data by gradually adding noise to the latent space and then learning to reverse this process. In the context of image generation, latent diffusion models can be used to generate high-quality images and can control specific properties of the generation process.

[0046] They operate in latent space, rather than directly in pixel space. This means that they first encode the data into a compact latent representation and then apply a diffusion process in the latent space.

[0047] In this embodiment, the encoder is a neural network that converts raw data (such as an image) into a representation in a latent space. This latent representation is usually of lower dimensionality and is able to capture the main features of the data; Step S702, detecting the encoding result of the first image by the encoder to obtain the watermark feature of the first image; Step S703, performing watermark removal processing on the first image encoding result according to the watermark feature and the control condition to obtain a watermark-free image encoding; In this embodiment, the watermark features are identified by the encoder. Once the watermark is detected, the model can de-watermark it in the latent space to obtain a watermark-free image encoding.

[0048] Step S704, the decoder in the trained image diffusion model decodes the watermark-free image code to generate a watermark-free image. In this embodiment, the decoder is the inverse process of the encoder, which converts the representation in the latent space back to the original data space. In image generation, the decoder is responsible for converting the latent representation back to a visualized image.

[0049] The technical solution of the embodiment of the present application is to introduce watermark control conditions into the image diffusion model during the watermark removal process, thereby enhancing the ability of the image diffusion model to focus on watermark features, thereby improving the quality of the restored image.

[0050] The technical solution of the embodiment of the present application introduces a watermark condition control module, uses the corrected image encoding as the main image feature space, and uses the extracted watermark instance encoding as the auxiliary image feature space to input into the ControlNet and SD main networks for diffusion denoising respectively. In one example, referring to Figure 8 As shown in the figure, it shows a schematic diagram of the image diffusion model watermark removal process. The image in the upper left corner is the input image to be processed including the watermark. The input image is converted into a latent space representation after encoding, which is recorded as Stable Diffusion (UNet) is a stable diffusion model based on the UNet architecture. It uses the processing results of controlnet as the control condition to guide the stable diffusion model of the UNet architecture to perform watermark removal. The image diffusion model uses the DDIM denoising algorithm to reduce the sampling steps in the process of generating watermark-free images, thereby speeding up the generation speed. The watermark image processing result is output after being processed by the image diffusion model, denoted as . As the input of the next watermark processing, Time is the number of cycles to generate a watermark-free image. After multiple cycles of processing, the watermark-free image is obtained. The whole process is an iterative process. Start by processing with Stable Diffusion and ControlNet to generate ,Then It can be used as the input for the next step and continue to iterate until a satisfactory dewatermarking result is achieved. By combining the powerful generation capability of Stable Diffusion and the control capability of ControlNet, precise image editing can be achieved.

[0051] In the example, ControlNet is a control network used to guide the image generation process according to a specific condition or mask (Mask). For example, the CCTV watermark condition generates control features (Cf) through the encoder (E), which are used to affect the output of Stable Diffusion. Mask (Cf) represents a mask that specifies the areas in the image that need to be processed or retained. The mask can be a binary image, where the white area represents the part that needs to be processed and the black area represents the part that is retained.

[0052] In one example, referring to Fig. 9 As shown, a flowchart of the steps of performing watermark removal processing on the first image encoding result by using the watermark feature and the control condition to obtain a watermark-free image encoding is shown, which may specifically include the following steps: Step S901, using the control condition as a watermark removal guide, performing the i-th watermark removal process on the first image encoding result through a preset denoising algorithm and the watermark feature to obtain the i-th watermark-free image encoding; i≥1; wherein the denoising algorithm can be a DDIM denoising algorithm.

[0053] Step S902, taking the i-th watermark-free image code as the i+1-th watermark removal processing input, and performing the i+1-th watermark removal processing in combination with the watermark feature and the control condition, until the watermark removal in the first image coding result is completed, and obtaining the watermark-free image code.

[0054] The technical solution of the embodiment of the present application increases the processing speed by reducing the number of denoising steps N through a preset denoising algorithm DDIM, and finally improves the image restoration quality through gradual denoising.

[0055] In one example, the process of performing angle correction on a watermarked image may specifically include the following steps: The watermarked image is angle-corrected by a preset method to obtain a first image; wherein the preset method includes: a Hough transform method, an edge detection method or a calculation horizontal projection method.

[0056] In one scenario, the captured document images are often tilted. Correcting the tilt angle can align the document content horizontally, reduce the difficulty of subsequent watermark extraction and image restoration, and reduce the computing pressure of the subsequent YOLO model.

[0057] For example, refer to Fig.10 As shown, the angle correction process of the document image is shown. The tilt angle can be estimated by detecting straight lines or edges in the document, and then the image is rotated to correct the tilt to obtain a corrected image. Common methods include Hough transform, edge detection and calculating horizontal projection.

[0058] In the technical solution of the embodiment of the present application, since the captured document image is often tilted, correcting the tilt angle can horizontally align the document content, reduce the difficulty of subsequent watermark extraction and image restoration, and alleviate the pressure of model calculation.

[0059] In one example, the process of constructing a watermark image dataset may specifically include the following steps: Acquire a watermark image data set; the watermark image data set includes: a plurality of watermark sample images; Acquire a watermark-free image dataset; the watermark-free image dataset includes: a plurality of watermark-free sample images; Adjusting at least one of the size, position and transparency of the watermark sample image to obtain an adjusted watermark sample image; In one scenario, we use existing image processing technology to add the collected watermarks to the original image with random size, position, transparency and shadow; and perform enhancement processing such as scaling, rotation and affine transformation on the watermarked image. This process simulates the uncertainty of watermarks in the real world and ensures that the images in the database can provide sufficient challenges for the algorithm model.

[0060] Adding the adjusted watermark sample image to a non-watermark sample image, and adding at least one of size information, position information and transparency information of the adjusted watermark sample image to the annotation information of the non-watermark sample image, to obtain a watermarked sample image; For example, the watermark sample image is adjusted in size, position and / or transparency to obtain an adjusted watermark sample image, the adjusted watermark sample image is added to the watermark-free sample image, and the size, position and / or transparency are added to the annotation information of the watermark-free sample image, and the adjusted watermark sample image is added.

[0061] As another example, the watermark sample image is resized and positioned to obtain an adjusted watermark sample image, the adjusted watermark sample image is added to the watermark-free sample image, and size information and position information are added to the annotation information of the watermark-free sample image, and the adjusted watermark sample image is added.

[0062] As another example, the watermark sample image is adjusted in size and transparency to obtain an adjusted watermark sample image, the adjusted watermark sample image is added to the watermark-free sample image, and size information and transparency information are added to the annotation information of the watermark-free sample image, and the adjusted watermark sample image is added.

[0063] As another example, the watermark sample image is adjusted in position and transparency to obtain an adjusted watermark sample image, the adjusted watermark sample image is added to the watermark-free sample image, and the position and transparency are added to the annotation information of the watermark-free sample image, and the adjusted watermark sample image is added.

[0064] Integrate multiple watermarked sample images to obtain a watermarked image dataset; The watermarked image data set is used to train an image segmentation model and an image diffusion model respectively to obtain a trained image segmentation model and a trained image diffusion model.

[0065] In one example, the watermark dataset construction process may include the following steps: When performing tasks such as watermark extraction or image restoration, building a comprehensive and representative watermark image watermark dataset is crucial for model training and optimization. We build the dataset for this system according to the following process: Watermark example collection and generation: We first set out to collect and generate diverse watermark images to ensure that the database covers a variety of possible watermark types. In addition to including basic geometric shapes such as points, lines, and surfaces, we also added complex designs such as artistic fonts and graphic combinations.

[0066] Selection of public original image data: In order to ensure the universality and verifiability of data images, we selected multiple public datasets as the source of watermark-free images. These datasets contain images of various scenes and objects, ensuring that our watermark image database can represent a variety of different application environments.

[0067] Data enhancement: In order to improve the generalization ability of the algorithm model, we pay special attention to collecting complex watermarks that may appear in the real world. Using advanced image processing technology, we add the collected watermarks to the original image with random size, position, transparency and shadow; and perform enhancement processing such as scaling, rotation and affine transformation on the watermarked image. This process simulates the uncertainty of watermarks in the real world and ensures that the images in the database can provide sufficient challenges for the algorithm model.

[0068] Data annotation: During the watermarking process, we recorded the exact location and attributes of each watermark. This information, together with the original watermark-free image, is crucial for the supervised learning phase of subsequent model training.

[0069] Data division: 80% of the watermark dataset is divided into a training set, and the remaining 20% ​​is divided into a test set. In order to meet the needs of real-life scenarios where machines need to automatically detect and remove watermarks that have never been seen, we ensure that the watermarks in the training set do not appear in the test set, which can well simulate the usage scenarios in real life.

[0070] Through this rigorous data preparation and dataset construction process, we ensure the quality and diversity of the watermark image database, laying a solid foundation for developing efficient and reliable watermark-related algorithms.

[0071] The technical solution of the embodiment of the present application constructs a special watermark data set, improves the existing watermark extraction technology, and optimizes the image restoration process.

[0072] The technical solution of the embodiment of the present application ensures the quality and diversity of the watermark image database, and lays a solid foundation for developing efficient and reliable watermark related algorithms.

[0073] For example, Fig.11 As shown, the training process of the image segmentation model and the image diffusion model for removing visible watermarks is shown, which may specifically include the following steps: Input the original image without watermark and the image with watermark, and pass the watermarked image through the YOLO segmentation model to obtain the pure watermark image; The watermarked image is then passed through the YOLO segmentation model to adjust the direction and output the watermarked image with adjusted angle; The angle-adjusted watermarked image is encoded and input into Unet and controlnet respectively, where N represents the number of cycles required to remove the watermark once; The pure watermark image is encoded and input into the control net; Controlnet uses the input watermarked image and pure watermark image as control conditions and outputs them to Unet, allowing Unet to only focus on the watermark area for watermark removal. This cycle needs to be repeated many times to remove the watermark. The number of cycles is N in the figure. For example, start removing the watermark frame, then remove the watermark little by little, and use the original image as a reference condition to determine whether it has been removed completely. Finally, decode it and output the watermark-free image.

[0074] First, we built a dataset specifically for watermark processing. This dataset contains not only various types of watermark samples, but also a variety of background images. By combining it with a public image library, we performed image processing operations and generated a series of watermarked and data-enhanced images to simulate various situations in actual application scenarios.

[0075] Furthermore, we built a watermark extraction module. In this module, we made improvements based on the YOLO framework. The improvements include: using ACT instance segmentation technology to achieve accurate positioning of the watermark. In addition, in order to deal with the problem of non-positive input images, the improvements also include: adding an ADJ correction output mechanism, which can automatically correct the directional deviation of the image to ensure the accurate extraction of the watermark. Through the above modules, we obtained the corrected input image and the extracted watermark instance, and input the corrected input image into the U-Net module and controlnet module in the image diffusion model for processing. The extracted watermark instance is preprocessed and encoded and input into the controlnet module for processing. During the training process, these two parts are supervised by the original image Original Img and the watermark image WaterMark respectively.

[0076] In the watermark removal-image restoration module, we made improvements based on the Stable Diffusion diffusion model. The improvements include: introducing a watermark condition control mechanism, which uses the watermark information extracted previously as an auxiliary signal to guide the diffusion model to restore the image more accurately. In view of the fact that the present invention is mainly aimed at the removal of targets such as watermarks, we optimized the diffusion steps of the model and lightweighted the U-Net module. The specific improvements do not include: replacing the original DDPM denoising algorithm that focuses on image diversity with the DDIM denoising algorithm that focuses more on image restoration quality and efficiency. Since we only remove the watermark as a goal, we reduce the number of denoising steps N to increase the processing speed while ensuring the quality remains unchanged. Finally, after step-by-step denoising, the original image after the watermark is removed can be obtained after decoding. These improvements have greatly improved the processing efficiency without affecting the removal effect. During the training process, the output results are supervised by the original image.

[0077] This solution adopts a two-stage method. In the first stage, the deep segmentation model is used to explicitly segment the watermark in the document image. In the second stage, the watermark denoising is performed using the controlnet-based assisted diffusion model with explicit watermark segmentation as input control condition to generate the de-watermarked image. This simplifies the complex problem and decomposes the watermark removal into two stages. The two stages can be optimized independently, which greatly reduces the number of training samples required and achieves better watermark removal results.

[0078] Reference Fig.12 FIG. 1 is a block diagram showing a structure of a device 1200 for removing visible watermarks from an image according to an embodiment of the present invention. The device may specifically include the following modules: A correction module 1201 is used to perform angle correction on the watermarked image to obtain a first image; A watermark extraction module 1202 is used to perform segmentation processing on the first image through a segmentation module of a trained image segmentation model to obtain watermark image data; The condition control module 1203 is used to input the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents the watermark mask image, and the processing result is input into the encoder of the trained image diffusion model as a control condition; The watermark removal module 1204 is configured to perform watermark denoising processing on the first image based on the control condition through the encoder to obtain a watermark-free image.

[0079] The apparatus for removing visible watermarks from images provided in the above embodiment can implement the technical solution described in the above embodiment of the method for removing visible watermarks from images. The specific implementation principles of the above modules or units can be found in the corresponding contents of the above embodiment of the method for removing visible watermarks from images, which will not be repeated here.

[0080] like Fig.13 As shown, the present invention also provides an electronic device 1300. The electronic device 1300 includes a processor 1301, a memory 1302 and a display 1303. Fig.13 Only some components of the electronic device 1300 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0081] In some embodiments, the memory 1302 may be an internal storage unit of the electronic device 1300, such as a hard disk or memory of the electronic device 1300. In other embodiments, the memory 1302 may also be an external storage device of the electronic device 1300, such as a plug-in hard disk, a smart memory card (SmartMediaCard, SMC), a secure digital (SecureDigital, SD) card, a flash card (FlashCard), etc. equipped on the electronic device 1300.

[0082] Furthermore, the memory 1302 may include both an internal storage unit of the electronic device 1300 and an external storage device. The memory 1302 is used to store application software installed in the electronic device 1300 and various data.

[0083] In some embodiments, the processor 1301 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 1302, such as the visible watermark removal method for an image in the present invention.

[0084] In some embodiments, the display 1303 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 1303 is used to display information on the electronic device 1300 and to display a visual user interface. The components 1301-1303 of the electronic device 1300 communicate with each other via a system bus.

[0085] In some embodiments of the present invention, when the processor 1301 executes the visible watermark removal program of the image in the memory 1302, the following steps may be implemented: Performing angle correction on the watermarked image to obtain a first image; Performing segmentation processing on the first image by using a segmentation module of a trained image segmentation model to obtain watermark image data; Inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents a watermark mask image; Inputting the processing result as a control condition into the encoder of the trained image diffusion model; The encoder performs watermark denoising processing on the first image based on the control condition to obtain a watermark-free image.

[0086] It should be understood that: when the processor 1301 executes the visible watermark removal program of the image in the memory 1302, in addition to the above functions, other functions may also be implemented, and the details may refer to the description of the corresponding method embodiment above.

[0087] Furthermore, the embodiment of the present invention does not specifically limit the type of the electronic device 1300 mentioned, and the electronic device 1300 may be a portable electronic device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of portable electronic devices include but are not limited to portable electronic devices equipped with IOS, Android, Microsoft or other operating systems. The above-mentioned portable electronic devices may also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, the electronic device 1000 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).

[0088] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the visible watermark removal method for an image provided by the above methods, the method comprising: Performing angle correction on the watermarked image to obtain a first image; Performing segmentation processing on the first image by using a segmentation module of a trained image segmentation model to obtain watermark image data; Inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents a watermark mask image; Inputting the processing result as a control condition into the encoder of the trained image diffusion model; The encoder performs watermark denoising processing on the first image based on the control condition to obtain a watermark-free image.

[0089] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.

[0090] The above is a detailed introduction to the visible watermark removal method, device, equipment and medium for an image provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A method for removing visible watermarks from an image, characterized in that: include: Performing angle correction on the watermarked image to obtain a first image; Performing segmentation processing on the first image by using a segmentation module of a trained image segmentation model to obtain watermark image data; Inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents a watermark mask image; Inputting the processing result as a control condition into the encoder of the trained image diffusion model; The encoder performs watermark denoising processing on the first image based on the control condition to obtain a watermark-free image.

2. The method for removing visible watermarks from an image according to claim 1, characterized in that: The segmentation module of the trained image segmentation model is used to segment the first image to obtain watermark image data, which specifically includes: Correcting the first image by a correction module of a trained image segmentation model to obtain a corrected first image; wherein the segmentation module includes: a detection branch and a segmentation branch; Detecting the watermark position of the corrected first image by using the detection branch of the trained image segmentation model to obtain a watermark position result; Generating mask data of the corrected first image through the segmentation branch of the trained image segmentation model; A watermark instance mask is obtained according to the watermark position result and the mask data, and the watermark instance mask is used as watermark image data.

3. A method for removing visible watermarks from an image according to claim 1 or 2, characterized in that: The inputting the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result specifically includes: Inputting the watermark image data into the encoder in the trained image diffusion model for encoding to obtain a watermark mask in a latent space; Encoding the first image through the encoder in the trained image diffusion model to obtain a first image encoding result in a latent space; The watermark mask and the first image encoding result are processed by a conditional generation module in the trained image diffusion model to obtain a processing result.

4. The method for removing visible watermarks from an image according to claim 1, characterized in that: The step of performing watermark denoising on the first image based on the control condition by the encoder to obtain a watermark-free image specifically includes: Encoding the first image by the encoder to obtain a first image encoding result in a latent space; Detecting the encoding result of the first image by the encoder to obtain a watermark feature of the first image; Performing watermark removal processing on the first image encoding result according to the watermark feature and the control condition to obtain a watermark-free image encoding; The watermark-free image code is decoded by a decoder in the trained image diffusion model to generate a watermark-free image.

5. The method for removing visible watermarks from an image according to claim 4, characterized in that: The dewatering process is performed on the first image encoding result by using the watermark feature and the control condition to obtain the watermark-free image encoding, specifically comprising: The control condition is used as a watermark removal guide, and the first image encoding result is subjected to an i-th watermark removal process by using a preset denoising algorithm and the watermark feature to obtain an i-th watermark-free image encoding; i≥1; The i-th watermark-free image code is used as the i+1-th watermark removal processing input, and the i+1-th watermark removal processing is performed in combination with the watermark feature and the control condition until the watermark removal in the first image coding result is completed to obtain the watermark-free image code.

6. The method for removing visible watermarks from an image according to claim 1, characterized in that: The performing angle correction on the watermarked image to obtain the first image specifically includes: The watermarked image is angle-corrected by a preset method to obtain a first image; wherein the preset method includes: a Hough transform method, an edge detection method or a calculation horizontal projection method.

7. A method for removing visible watermarks from an image according to claim 1, 2 or 4, characterized in that: Also includes: Get watermark image dataset; The watermark image data set includes: a plurality of watermark sample images; Acquire a watermark-free image dataset; the watermark-free image dataset includes: a plurality of watermark-free sample images; Adjusting at least one of the size, position and transparency of the watermark sample image to obtain an adjusted watermark sample image; Adding the adjusted watermark sample image to a non-watermark sample image, and adding at least one of size information, position information and transparency information of the adjusted watermark sample image to the annotation information of the non-watermark sample image, to obtain a watermarked sample image; Integrate multiple watermarked sample images to obtain a watermarked image dataset; The watermarked image data set is used to train an image segmentation model and an image diffusion model respectively to obtain a trained image segmentation model and a trained image diffusion model.

8. A device for removing visible watermarks from an image, characterized in that: include: Correction module, watermark extraction module, condition control module and watermark removal module; The correction module is used to perform angle correction on the watermarked image to obtain a first image; The watermark extraction module is used to perform segmentation processing on the first image through the segmentation module of the trained image segmentation model to obtain watermark image data; The condition control module is used to input the first image and the watermark image data into the trained image diffusion model for processing to obtain a processing result; wherein the processing result represents the watermark mask image, and the processing result is input into the encoder of the trained image diffusion model as a control condition; The watermark removal module is used to perform watermark denoising on the first image based on the control condition through the encoder to obtain a watermark-free image.

9. An electronic device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps of the method for removing visible watermarks from an image as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps of the method for removing visible watermarks from an image as described in any one of claims 1 to 7.