Image watermark removing method and device based on deep learning
By introducing deep learning and YOLOv5 models into image watermarking technology, combining threshold segmentation and image recognition technology, the problem of high time-consuming watermark recognition in the existing technology is solved, and a more efficient and accurate watermark removal effect is achieved.
Patent Information
- Application Number
- CN202510230192.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the image watermarking method is time-consuming and the model is usually non-lightweight, making it difficult to effectively identify and remove watermarks.
The image watermarking method based on deep learning is adopted, and the watermark area frame detection is performed on the watermark image with the trained YOLOv5 model. Combined with threshold segmentation and image recognition technology, the watermark is gradually removed to generate an image with the watermark removal.
The accuracy and efficiency of watermark area box recognition are improved. The model used is designed with lightweight design, and does not require large space storage, which significantly reduces the time for watermark removal.
Smart Images

Figure CN120182318A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular, to an image watermark removal method based on deep learning and an image watermark removal device based on deep learning. Background Art
[0002] The existing image watermark removal methods mainly include the following several types: 1. Covering method: Use pictures, texts, mosaic effects or other watermarks to directly cover the watermark area of the original image.
[0003] 2. Cropping method: Directly crop the video watermark area, and only retain the clean image area. This method is only applicable to videos with upper and lower black frames or videos where the watermark pattern is located at the edge of the video.
[0004] The existing watermark removal methods first need to identify the watermark, but the models used for existing watermark identification are usually not lightweight models and are time-consuming.
[0005] Therefore, it is hoped that there is a technical solution to solve or at least alleviate the above deficiencies of the existing technology. Summary of the Invention
[0006] The purpose of the present invention is to provide an image watermark removal method based on deep learning to at least solve one of the above technical problems.
[0007] The present invention provides the following solutions: According to one aspect of the present invention, there is provided an image watermark removal method based on deep learning, and the image watermark removal method based on deep learning includes: Obtain an image to be watermark-removed; Obtain a trained YOLOv5 model; Input the image to be watermark-removed into the YOLOv5 model to obtain a watermark area box; Perform threshold segmentation on the watermark area box to obtain a segmented watermark area; Remove the watermark in the segmented watermark area to obtain an image with the watermark removed.
[0008] Optionally, after performing threshold segmentation on the watermark area box to obtain a segmented watermark area and before removing the watermark in the segmented watermark area to obtain an image with the watermark removed, the image watermark removal method based on deep learning further includes: Adjust the boundary of the segmented watermark area to obtain an adjusted segmented watermark area; Remove the watermark in the adjusted segmented watermark area to obtain an image with the watermark removed Optionally, the adjustment of the boundary of the segmented watermark region to obtain the adjusted segmented watermark region includes: Performing image recognition and orientation recognition on the segmented watermark region to obtain an image recognition result and an orientation recognition result of the segmented watermark region; Generating a reference watermark image according to the image recognition result and the orientation recognition result; Registering the reference watermark image with the segmented watermark region to obtain a registered image; Obtaining edge pixel information of the registered image, and each piece of the edge pixel information forms the adjusted segmented watermark region.
[0009] Optionally, the performing image recognition and orientation recognition on the segmented watermark region to obtain an image recognition result and an orientation recognition result of the segmented watermark region includes: Performing feature extraction on the segmented watermark region to obtain a feature vector; Obtaining an orientation recognition network; Performing an initial orientation recognition result through the orientation recognition network according to the feature vector and the segmented watermark region; Performing orientation correction on the initial orientation recognition result to obtain a final orientation recognition result; Performing image recognition on the segmented watermark region based on the orientation recognition result to obtain an image recognition result, where the image recognition result includes font information and glyph information.
[0010] Optionally, generating a reference watermark image according to the image recognition result and the orientation recognition result includes: Obtaining a preset character database, where the preset character database includes at least one preset character contour image and a preset font corresponding to each preset character; Obtaining a preset character contour image corresponding to both the font information and the glyph information of the image recognition result; Performing orientation adjustment on the preset character contour image according to the obtained final orientation recognition result to obtain an adjusted preset character contour image; Obtaining pixel point size information of each part of the segmented watermark region; Setting the size information of the preset character contour image according to the pixel point size information of each part of the segmented watermark region to obtain a reference watermark image.
[0011] Optionally, the registering the reference watermark image with the segmented watermark region to obtain a registered image includes: Obtaining a trained deep convolutional neural network; Input the reference watermark image and the segmented watermark region into the deep convolutional neural network respectively, and output the features extracted by the pooling layers in different convolutional blocks in the deep convolutional neural network; For the outputs of each of the last three pooling layers in the convolutional neural network, divide different search regions in each layer, and use the nearest neighbor matching algorithm to complete rough matching based on the feature points obtained from different search regions; After screening the feature points based on the rough matching and using the nearest neighbor matching algorithm, obtain the corresponding set of feature points, and use the point set registration algorithm based on the set of feature points to complete the registration of the reference watermark image and the segmented watermark region.
[0012] Optionally, the output of the bottleneck layer in the modified C3 module is expressed as follows:
[0013]
[0014] * represents the convolution operation, ω primary,k is the k-th standard convolution kernel, let ω k be the k-th dynamic convolution kernel, and there are K dynamic convolution kernels in total, α k is the dynamically generated weight. Optionally, the comprehensive loss function of the YOLOv5 model is: ζ = λ cls ζ cls + λ loc ζ loc + λ pose ζ pose ; Among them, ζ cls refers to the classification loss, which measures the gap between the predicted class and the true class, and uses the cross-entropy loss; ζ loc refers to the localization loss, which measures the gap between the predicted bounding box and the true bounding box; ζ pose refers to the pose estimation loss, which measures the gap between the predicted key point positions and the true key point positions; uses the mean squared error, λ cls 、λ loc 、λ pose These are the weight hyperparameters corresponding to the loss terms, which are used to balance the contributions of different loss terms to the total loss.
[0015] This application also provides an image watermark removal device based on deep learning. The image watermark removal device based on deep learning includes: A module for obtaining an image to be watermark-removed, which is used to obtain an image to be watermark-removed; A module for obtaining a model, which is used to obtain a trained YOLOv5 model; The watermark area box detection module is configured to input the image to be de-watermarked into the YOLOv5 model to obtain the watermark area box; The segmented watermark area acquisition module is configured to perform threshold segmentation on the watermark area box to obtain the segmented watermark area; The watermark removal module is configured to remove the watermark within the segmented watermark area to obtain the watermark-removed image.
[0016] The image de-watermarking method based on deep learning of the present application uses a YOLO model to identify the watermark area. Compared with the prior art, the recognition of the watermark area box is more accurate, and the model used is designed with lightweight and does not require a large amount of space for storage. Description of the Drawings
[0017] Figure 1 is a schematic flow chart of the image de-watermarking method based on deep learning in an embodiment of the present application; Figure 2 is a schematic structural diagram of the prior art of the YOLOv5 model in an embodiment of the present application.
[0018] Figure 3 is a schematic diagram of the watermark in the prior art.
[0019] Figure 4 is a schematic diagram of the reference watermark image obtained by the method of the present application.
[0020] Figure 5 is an image with a certain offset after registration. Detailed Embodiments
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] As Figure 1 shown, the image de-watermarking method based on deep learning includes: Step 1: Obtain the image to be de-watermarked; Step 2: Obtain the trained YOLOv5 model; Step 3: Input the image to be de-watermarked into the YOLOv5 model to obtain the watermark area box; Step 4: Perform threshold segmentation on the watermark area box to obtain the segmented watermark area; Step 5: Remove the watermark within the segmented watermark area to obtain an image with the watermark removed.
[0023] In this embodiment, after threshold segmentation of the watermark area box to obtain the segmented watermark area, and before removing the watermark within the segmented watermark area to obtain an image with the watermark removed, the deep learning-based image watermark removal method further includes: Adjust the boundary of the segmented watermark area to obtain an adjusted segmented watermark area; Remove the watermark within the adjusted segmented watermark area to obtain an image with the watermark removed In some cases, there may be a situation where some edge pixels in the segmented watermark area should be removed but are not. This is an unrecognized problem caused by threshold segmentation. Through the adjustment of this application, this situation can be avoided as much as possible.
[0024] In this embodiment, the method of threshold segmentation of the watermark area box to obtain the segmented watermark area can be as follows: Obtain the final threshold TK based on the watermark area box. When the threshold is the final threshold TK, the average gray value of the foreground is equal to the average gray value of the background. Specifically, convert the watermark part in the watermark area box into a grayscale image; determine the initial threshold T0 based on the maximum gray value and the minimum gray value; divide the foreground and background of the watermark area according to the initial threshold T0; calculate the average gray value of the foreground and the average gray value of the background respectively; determine the new threshold Ti according to the average gray value of the foreground and the average gray value of the background; divide the foreground and background of the watermark area with the new threshold Ti and iterate the above steps until the average gray value of the foreground is equal to the average gray value of the background to obtain the final threshold TK; Segment the watermark area box with the final threshold TK to obtain the segmented watermark area.
[0025] In this embodiment, adjusting the boundary of the segmented watermark area to obtain an adjusted segmented watermark area includes: Perform image recognition and orientation recognition on the segmented watermark area to obtain the image recognition result and orientation recognition result of the segmented watermark area; Generate a reference watermark image according to the image recognition result and the orientation recognition result; Register the reference watermark image with the segmented watermark area to obtain a registered image; Obtain the edge pixel information of the registered image, and each piece of the edge pixel information forms the adjusted segmented watermark area.
[0026] In this embodiment, performing image recognition and orientation recognition on the segmented watermark region to obtain an image recognition result and an orientation recognition result of the segmented watermark region includes: Performing feature extraction on the segmented watermark region to obtain a feature vector; Obtaining an orientation recognition network; Performing an initial orientation recognition result through the orientation recognition network according to the feature vector and the segmented watermark region; Performing orientation correction on the initial orientation recognition result to obtain a final orientation recognition result; specifically, performing connectivity detection on the segmented watermark region to obtain a connectivity detection result, and performing orientation correction based on the connectivity detection result and the initial orientation recognition result to obtain a final orientation recognition result; Performing image recognition on the segmented watermark region based on the orientation recognition result to obtain an image recognition result, where the image recognition result includes font information and glyph information.
[0027] In this embodiment, generating a reference watermark image according to the image recognition result and the orientation recognition result includes: Obtaining a preset character database, where the preset character database includes at least one preset character contour image and a preset font corresponding to each preset character; Obtaining a preset character contour image corresponding to both the font information and the glyph information of the image recognition result; Performing orientation adjustment on the preset character contour image according to the obtained final orientation recognition result to obtain an adjusted preset character contour image; Obtaining pixel point size information of each part of the segmented watermark region; Setting the size information of the preset character contour image according to the pixel point size information of each part of the segmented watermark region to obtain a reference watermark image.
[0028] In this embodiment, the ratio between each stroke of each character is fixed, regardless of the complexity and font of the character. For example, for a large Song typeface character, the ratio between the lengths of both ends of one horizontal stroke and the sizes of other parts of the character always remains a certain ratio. Therefore, as long as the pixel point size information of each part of the segmented watermark region is known, the size information of the preset character contour image can also be known. For example, taking a large character as an example, as long as the distance between both ends of one horizontal stroke is known, based on the equal ratio scaling principle, the size information of each part (such as the character "one" and the character "person") of the preset character contour image can be determined.
[0029] In this embodiment, the following scheme can be adopted for image recognition: Performing preprocessing on the segmented watermark region to obtain a character contour image; Enhancing the saturation and contrast of the text outline image to obtain a text outline enhanced image; Dividing the text contour enhanced image into a set number of regions; Obtaining feature values of each region in the text contour enhanced image; The characteristic values of the text outline enhanced image are matched with the characteristic values of each character in the font library for similarity to obtain multiple matching values, and the text information in the font library corresponding to the highest matching value is output, wherein the text information includes the glyph and font type; each piece of information in the font library includes the glyph, font type and characteristic values of each area; the characteristic values of the text outline enhanced image include the characteristic values of each area of the text outline enhanced image, and the characteristic values of each character in the font library include the characteristic values of each area corresponding to each character.
[0030] In this embodiment, registering the reference watermark image with the segmented watermark region to obtain a registered image includes: Get a trained deep convolutional neural network; Inputting the reference watermark image and the segmented watermark area into the deep convolutional neural network respectively, and outputting the features extracted by the pooling layers in different convolutional blocks in the deep convolutional neural network; For the outputs of the last three pooling layers in the convolutional neural network, different search areas are divided in each layer, and the nearest neighbor matching algorithm is used to complete the rough matching based on the feature points obtained in different search areas; After screening the feature points based on the rough matching and using the nearest neighbor matching algorithm, a corresponding feature point set is obtained, and a point set registration algorithm is used to complete the registration of the reference watermark image and the segmented watermark area based on the feature point set.
[0031] See to Figure 5 , Figure 3 It means that the part of the segmented watermark area is obtained by performing threshold segmentation, from Figure 3 As can be seen from the figure, the threshold segmentation method may cause the actual watermark text to be missing. For example, the word "strictly prohibit copying" is actually missing a part (see Figure 5 It can be seen from the figure), it can be understood that when the threshold segmentation is actually performed, there will not be so many missing pixels. In order to be able to more intuitively reflect the role of the present application, there are more missing pixels in this example. The actual situation is that each part of each word may lack several pixels. This situation is difficult to see intuitively with the naked eye. Therefore, in the example of the present application, the missing pixels are increased so that the missing parts can be intuitively seen.
[0032] In this embodiment, Figure 4That is the reference watermark image obtained by the method of this application. It can be understood that through image recognition, the text in Figure 3 is strictly prohibited from being copied, and the inclination angle can be known through the recognition of the orientation. Then, the specific font and glyph can be recognized through the font and glyph, so as to obtain the reference watermark image of this application.
[0033] See Figure 5 , Figure 5 which is actually the registered image. Since the actually registered image will directly cover the original 4 characters of "strictly prohibited from copying", in order to show the difference between the two, Figure 5 the image shown is the image that has been offset to a certain extent after registration, so that it can be seen that if Figure 3 and Figure 4 are registered, the actual pixel range of Figure 4 will be formed, so that a more complete removal of the watermark can be achieved when removing the watermark.
[0034] See Figure 2 , in this embodiment, the trained YOLOv5 model of this application includes multiple improved C3 modules. Among them, the bottleneck layer ( Figure 2 the bottleneck in
[0035] ) in each improved C3 module is changed to a ghost module and a dynamic convolution layer.
[0036]
[0037] * represents the convolution operation, ω primary,k is the k-th standard convolution kernel. Let ω k be the k-th dynamic convolution kernel, and there are K dynamic convolution kernels in total. α k is the dynamically generated weight.
[0038] In this embodiment, the comprehensive loss function of the YOLOv5 model is:
[0039] ζ = λ cls ζ cls + λ loc ζ loc + λ pose ζ pose ;
[0040] Among them, ζ cls refers to the classification loss, which measures the gap between the predicted class and the true class, and uses the cross-entropy loss; ζ loc refers to the localization loss, which measures the gap between the predicted bounding box and the true bounding box; ζ poseFinger pose estimation loss, which measures the gap between the predicted key point positions and the true key point positions; using mean squared error, λ cls , λ loc , λ pose These are the weight hyperparameters corresponding to the loss terms, used to balance the contributions of different loss terms to the total loss.
[0041] In this embodiment, ζ cls The classification loss calculation formula is as follows:
[0042]
[0043] where N is the number of samples, C is the number of classes, y i,c is the true label that sample i belongs to class c, is the probability of class c predicted by the model.
[0044] In this embodiment, ζ loc The localization loss calculation formula is as follows:
[0045] where b i is the true bounding box of sample i, is the bounding box predicted by the model. In this embodiment, ζ pose The pose estimation loss calculation formula is as follows: where K is the number of key points, p i,j is the true position of the j-th key point of sample i, is the position of the j-th key point predicted by the model. This application also provides an image watermark removal device based on deep learning. The image watermark removal device based on deep learning includes a watermark-containing image acquisition module, a model acquisition module, a watermark area box detection module, a segmented watermark area acquisition module, and a watermark removal module. Among them, The watermark-containing image acquisition module is used to acquire a watermark-containing image; The model acquisition module is used to acquire a trained YOLOv5 model; The watermark area box detection module is used to input the watermark-containing image into the YOLOv5 model to obtain a watermark area box; The segmented watermark area acquisition module is used to perform threshold segmentation on the watermark area box to obtain a segmented watermark area; The watermark removal module is used to remove the watermark within the segmented watermark area to obtain an image with the watermark removed. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A deep learning-based image watermarking method, characterized in that: The image watermarking method based on deep learning includes: Obtain the image to be dewatered; Get the trained YOLOv5 model; Input the image to be dewatermarked into the YOLOv5 model to obtain the watermark area frame; Perform threshold segmentation on the watermark area frame to obtain the segmented watermark area; The watermark in the segmented watermark area is removed to obtain an image with the watermark removed.
2. The method of claim 1, wherein the boundary of the segmented watermark area is adjusted, After performing threshold segmentation on the watermark region frame to obtain the segmented watermark region, and before removing the watermark in the segmented watermark region to obtain the watermark-removed image, the image watermark removal method based on deep learning further includes: Adjusting the boundary of the segmented watermark area to obtain an adjusted segmented watermark area; The watermark in the adjusted segmented watermark area is removed to obtain an image with the watermark removed.
3. The image watermark removal method based on deep learning as claimed in claim 2, characterized in that: The step of adjusting the boundary of the segmented watermark area to obtain the adjusted segmented watermark area includes: Performing image recognition and orientation recognition on the segmented watermark area, thereby obtaining an image recognition result and an orientation recognition result of the segmented watermark area; Generate a reference watermark image according to the image recognition result and the orientation recognition result; Registering the reference watermark image with the segmented watermark region to obtain a registered image; The edge pixel information of the registered image is obtained, and each of the edge pixel information constitutes an adjusted segmented watermark area.
4. The image watermark removal method based on deep learning as claimed in claim 3, characterized in that: The performing image recognition and orientation recognition on the segmented watermark area to obtain the image recognition result and orientation recognition result of the segmented watermark area includes: Performing feature extraction on the segmented watermark area to obtain a feature vector; Acquire a location recognition network; According to the feature vector and the segmented watermark area, the initial orientation recognition result is obtained through the orientation recognition network; Performing orientation correction on the initial orientation recognition result to obtain the final orientation recognition result; Image recognition is performed on the segmented watermark area based on the orientation recognition result, thereby obtaining an image recognition result, wherein the image recognition result includes font information and glyph information.
5. The image watermark removal method based on deep learning as claimed in claim 4, characterized in that: Generating a reference watermark image based on the image recognition result and the orientation recognition result includes: Acquire a preset text database, wherein the preset text database includes at least one preset text outline image and a preset font corresponding to each preset text; Acquire a preset text outline image corresponding to both the font information and the glyph information of the image recognition result; Adjusting the orientation of the preset character outline image according to the final orientation recognition result, thereby obtaining an adjusted preset character outline image; Get the pixel size information of each part of the segmented watermark area; The size information of the preset text outline image is set according to the pixel size information of each part of the segmented watermark area, so as to obtain a reference watermark image.
6. The image watermark removal method based on deep learning as claimed in claim 5, characterized in that: The registering the reference watermark image with the segmented watermark region to obtain a registered image comprises: Get a trained deep convolutional neural network; Inputting the reference watermark image and the segmented watermark area into the deep convolutional neural network respectively, and outputting the features extracted by the pooling layers in different convolutional blocks in the deep convolutional neural network; For the outputs of the last three pooling layers in the convolutional neural network, different search areas are divided in each layer, and the nearest neighbor matching algorithm is used to complete the rough matching based on the feature points obtained in different search areas; After screening the feature points based on the rough matching and using the nearest neighbor matching algorithm, a corresponding feature point set is obtained, and a point set registration algorithm is used to complete the registration of the reference watermark image and the segmented watermark area based on the feature point set.
7. The image watermark removal method based on deep learning as claimed in claim 1, characterized in that: The trained YOLOv5 model includes multiple improved C3 modules, wherein the bottleneck layer in each improved C3 module is changed to a ghost module and a dynamic convolution layer.
8. The image watermark removal method based on deep learning as claimed in claim 7, characterized in that: The bottleneck layer output in the modified C3 module is expressed as follows: * represents the convolution operation, ω primary,k is the kth standard convolution kernel, let ω k is the kth dynamic convolution kernel, there are K dynamic convolution kernels in total, α k is a dynamically generated weight.
9. The image watermark removal method based on deep learning as claimed in claim 8, characterized in that: The comprehensive loss function of the YOLOv5 model is: g=l cls g cls +λ loc g loc +λ pose g pose ; Among them, cls Refers to classification loss, which measures the gap between the predicted category and the true category, using cross entropy loss; ζ loc The position loss measures the gap between the predicted bounding box and the true bounding box; ζ pose Refers to the pose estimation loss, which measures the gap between the predicted key point position and the true key point position; using the mean square error, λ cls , loc , pose These are the weight hyperparameters for the corresponding loss terms, which are used to balance the contribution of different loss terms to the total loss.
10. An image watermark removal device based on deep learning, characterized in that: The image watermark removal device based on deep learning includes: A module for acquiring an image to be dewatered, wherein the module is used to acquire an image to be dewatered; A model acquisition module, wherein the model acquisition module is used to acquire a trained YOLOv5 model; A watermark region frame detection module, wherein the watermark region frame detection module is used to input the image to be dewatered into the YOLOv5 model to obtain a watermark region frame; A segmented watermark region acquisition module, wherein the segmented watermark region acquisition module is used to perform threshold segmentation on the watermark region frame to obtain the segmented watermark region; The watermark removal module is used to remove the watermark in the segmented watermark area, thereby obtaining an image with the watermark removed.