Watermark extraction model training method and device, equipment, storage medium and product
By performing multiple distortion simulations in the watermark embedded image, generating noise simulated images and extracting watermark information, the problem of incomplete noise simulation of the watermark extraction model in specific scenarios is solved, and correct watermark extraction is achieved in various scenarios.
Patent Information
- Application Number
- CN202510740358.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-16
AI Technical Summary
In the existing technology, the watermark extraction model is not comprehensive enough in noise simulation, which makes it difficult to correctly extract watermark information in certain scenarios.
The sample watermark information is embedded into the carrier image through the watermark information embedding model to generate a watermark-embedded image. Noise simulation is performed using various distortion simulation methods such as perspective distortion, illumination distortion, moiré distortion and JPEG compression to generate a noise-simulated image. Finally, the watermark information is extracted through the watermark extraction model, and the extraction judgment result is constructed to determine that the model training is completed.
This ensures that the trained watermark extraction model can correctly extract watermark information in various scenarios, thereby improving the applicability and accuracy of the model.
Smart Images

Figure CN120655482A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to watermark extraction model training methods, devices, equipment, storage media and products. Background Art
[0002] Screen capture-resistant image watermarking technology is an advanced digital watermarking system designed to prevent the illegal copying of image content through screen capture or screenshots. By embedding a hidden and difficult-to-remove watermark into an image, the watermark information can be successfully extracted from the copy even if the image is copied through screen capture or other secondary photography techniques. This technology helps verify the originality and copyright ownership of an image, providing a strong technical guarantee for copyright protection of image content.
[0003] This technology currently generally relies on model implementation (using the model to embed and extract watermarks). However, when training the model corresponding to this technology, the noise simulation is not comprehensive enough, which makes it difficult for the model to correctly extract watermark information for images of specific scenes. Summary of the Invention
[0004] The main purpose of this application is to provide a watermark extraction model training method, device, equipment, storage medium and product, aiming to solve the technical problem that the noise simulation of the existing technology is not comprehensive enough, which makes it difficult for the model to correctly extract watermark information for images of specific scenes.
[0005] To achieve the above objectives, this application proposes a watermark extraction model training method, which includes:
[0006] Embed the sample watermark information into the carrier image through the watermark information embedding model to generate a watermark embedded image;
[0007] Performing noise simulation on the watermark embedded image to generate a noise simulated image, wherein the noise simulation is performed using at least one distortion simulation method;
[0008] Performing watermark extraction on the noise simulation image through a watermark extraction model to obtain watermark information;
[0009] An extraction determination result is constructed according to the sample watermark information and the watermark information, and whether the watermark extraction model is trained is determined based on the extraction determination result.
[0010] Optionally, the distortion simulation method includes perspective distortion simulation;
[0011] The performing noise simulation on the watermark embedded image to generate a noise simulated image includes:
[0012] Performing text recognition on the watermark-embedded image to determine a plurality of dispersed regions, wherein the dispersed regions are image regions where text may be present;
[0013] merging the plurality of scattered regions into a complete connected region;
[0014] Perform perspective distortion simulation processing on the connected area to generate a noise simulation image.
[0015] Optionally, the distortion simulation method includes illumination distortion simulation;
[0016] The performing noise simulation on the watermark embedded image to generate a noise simulated image includes:
[0017] Performing text recognition on the watermark embedded image to determine a plurality of pixel-level dispersed regions, wherein the pixel-level dispersed regions are image regions at the pixel level where text may exist;
[0018] Performing single-time illumination distortion processing on other areas in the watermark-embedded image, and performing multiple-time illumination distortion processing on each of the pixel-level dispersed areas to generate an illumination-distorted image;
[0019] Performing line light source distortion processing on the illumination-distorted image to generate a noise-simulated image.
[0020] Optionally, performing line light source distortion processing on the illumination-distorted image to generate a noise-simulated image includes:
[0021] performing line light source distortion processing on the illumination-distorted image to generate a light source-distorted image;
[0022] Performing bilateral filtering on the light source distorted image to generate a noise simulation image.
[0023] Optionally, the watermark information embedding model includes a watermark information encoding layer and a watermark information embedding layer;
[0024] The method of embedding the sample watermark information into the carrier image through the watermark information embedding model to generate the watermark embedded image includes:
[0025] Performing cyclic error correction coding on the sample watermark information through the watermark information coding layer to generate a watermark bit sequence;
[0026] Converting the watermark bit sequence through the watermark information embedding layer to generate multi-dimensional watermark information, wherein the multi-dimensional watermark information is information of the same size as the carrier image;
[0027] The multi-dimensional watermark information is fused with the carrier image at the channel level through the watermark information embedding layer, and convolution processing is performed on the fused data to generate a watermark embedded image.
[0028] Optionally, the step of performing channel-level fusion of the multi-dimensional watermark information and the carrier image through the watermark information embedding layer and performing convolution processing on the fused data to generate a watermark-embedded image includes:
[0029] The multi-dimensional watermark information is fused with the carrier image at the channel level through the watermark information embedding layer, and convolution processing is performed on the fused data to generate a first watermark image;
[0030] performing weighted processing on the first watermark image through the watermark information embedding layer to generate a second watermark image, wherein the weighted processing is performed based on a gradient mask matrix obtained by performing gradient recognition on the first watermark image;
[0031] The carrier image and the second watermark image are superimposed through the watermark information embedding layer to generate a watermark embedded image.
[0032] In addition, to achieve the above-mentioned purpose, the present application also proposes a watermark extraction model training device, which includes:
[0033] An embedding module, used for embedding sample watermark information into a carrier image through a watermark information embedding model to generate a watermark-embedded image;
[0034] a simulation module, configured to perform noise simulation on the watermark embedded image to generate a noise simulated image, wherein the noise simulation is performed using at least one distortion simulation method;
[0035] An extraction module, configured to extract watermarks from the noise simulation image using a watermark extraction model to obtain watermark information;
[0036] A determination module is used to construct an extraction determination result according to the sample watermark information and the watermark information, and determine whether the watermark extraction model is trained based on the extraction determination result.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a watermark extraction model training device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the watermark extraction model training method described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the watermark extraction model training method described above are implemented.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the watermark extraction model training method described above.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] After the watermark is embedded in the image, one or more distortion simulation methods are used to perform noise simulation on the watermark-embedded image, which ensures that the generated noise image fits various actual scenes as much as possible, and ensures that the trained watermark extraction model can still correctly extract watermark information in various scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 A flowchart of the first embodiment of the watermark extraction model training method of the present application is provided;
[0045] Figure 2 A flow chart of the second embodiment of the watermark extraction model training method provided in this application;
[0046] Figure 3 This is a schematic diagram of the module structure of the watermark extraction model training device according to an embodiment of the present application;
[0047] Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the watermark extraction model training method in the embodiment of the present application.
[0048] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0050] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0051] Based on this, the embodiment of the present application provides a watermark extraction model training method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the watermark extraction model training method of this application.
[0052] In this embodiment, the watermark extraction model training method includes steps S10 to S40:
[0053] Step S10: embedding the sample watermark information into the carrier image through the watermark information embedding model to generate a watermark-embedded image.
[0054] It should be noted that the executor of this embodiment can be the watermark extraction model training device, and the watermark extraction model training device can be an electronic device such as a personal computer, a server, or other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the watermark extraction model training method of this application is explained using the watermark extraction model training device as an example.
[0055] It should be noted that the sample watermark information can be standard information for use in the training model, pre-provided by the administrator of the watermark extraction model training device, and its specific content can be set or adjusted according to actual needs. The carrier image can be an image for use in the training model, pre-provided by the administrator of the watermark extraction model training device, and its specific content can be set or adjusted according to actual needs.
[0056] In actual use, there can be at least one sample watermark information and carrier image. If there are multiple, the sample watermark information and carrier image can be combined. For example, assuming that the sample watermark information is A1 and A2, and the carrier images are B1 and B2, then there can be combinations such as A1-B1, A1-B2, A2-B1, and A2-B2. For each combination, a corresponding watermark embedded image can be generated.
[0057] The watermark information embedding model may be a model pre-set by a manager of a watermark extraction model training device.
[0058] Step S20: performing noise simulation on the watermark embedded image to generate a noise simulated image.
[0059] It should be noted that text images, also known as document images, are images that specifically contain text information. Their main characteristic is that they contain readable text content, which may be handwritten, printed, or other types of symbols and characters. These include scanned document images, PDF document images, and handwritten text images.
[0060] Text images typically contain a large amount of textual information, and this text is often located in specific areas of the image, with high contrast and distinct structural boundaries. Therefore, when embedding a watermark, special attention must be paid to the area where the watermark is embedded to avoid interference with the textual content. Unlike natural images, which have complex and dispersed textures and structures, watermarks can be embedded in different areas of the image without significantly affecting the visual effect. However, in text images, the area for watermark embedding requires particular caution. Usually, the background, blank areas, or low-contrast areas of the image are selected for embedding to avoid affecting the readability of the text.
[0061] The text in text images typically has a clear structure and distinct boundaries, so when embedding a watermark, it's crucial to maintain its clarity and readability. Watermark embedding must consider the edges and structure of the text to avoid disrupting its form. This differs from other image types, where watermarks can be embedded within textured areas to minimize disruption. In text images, the integrity of the text is a top priority.
[0062] In text images, the strength and visibility of the watermark need to be adjusted according to the distribution and importance of the text. Since text images usually have high-contrast text areas and low-contrast background areas, the strength of the watermark embedding can be dynamically adjusted based on the local characteristics of the image.
[0063] It is understandable that in order to ensure that the trained watermark extraction model can correctly extract watermark information for various scenarios, it is necessary to simulate various scenarios as much as possible. Therefore, a variety of distortion simulation methods can be used to simulate the impact of various scenarios on the image to enhance the performance of the model and ensure that it can cope with various scenarios.
[0064] Based on this, a screen capture noise simulation layer can be pre-set. This layer uses the characteristics of text images to identify text regions in watermarked images. Simulating the characteristics of screen capture images, the images after text region identification are subjected to perspective distortion and illumination distortion, respectively. Furthermore, moiré distortion and JPEG compression distortion can be applied sequentially to the watermarked images to generate a watermark image that simulates the screen capture process during the training phase. This image is used to perform end-to-end training on the neural networks in the watermark encoding and decoding systems, enabling the watermark decoding system to better embed and extract watermarks from text images.
[0065] The image distortion that occurs during screen capture can be represented by a combination of perspective distortion, lighting distortion, moiré distortion, JPEG compression distortion, and other noise. That is, in actual use, at least one distortion simulation method can be used to simulate noise. For example, one or more combinations of distortion simulation methods such as perspective distortion simulation, lighting distortion simulation, moiré distortion simulation, and JPEG compression simulation can be used to simulate noise.
[0066] Step S30: extracting watermarks from the noise simulation image using a watermark extraction model to obtain watermark information.
[0067] It should be noted that the watermark extraction model can be a model pre-set by the administrator of the watermark extraction model training equipment, and the watermark extraction model and the watermark information embedding model can be adversarial networks to each other. For example, the adversarial network consists of a group of models, one model in the group of models serves as an encoder, and the other model serves as a decoder. The two are trained together. At this time, the watermark extraction model can be the decoder in the adversarial network, and the watermark information embedding model can be the encoder in the adversarial network.
[0068] In actual use, the watermark extraction model can extract watermarks from noise-simulated images, thereby obtaining the watermark information contained in the noise-simulated images.
[0069] Step S40: constructing an extraction determination result according to the sample watermark information and the watermark information, and determining whether the watermark extraction model has been trained based on the extraction determination result.
[0070] In actual use, the extracted watermark information can be compared with the sample watermark information to determine whether the watermark extraction model has correctly extracted the watermark information, thereby generating an extraction judgment result.
[0071] It is understandable that during the training process, multiple training sessions are generally conducted. At this time, the extraction judgment results obtained from multiple training sessions can be collected, and the extraction accuracy of the watermark extraction model for watermark extraction can be determined based on the extraction judgment results. If the extraction accuracy reaches the preset accuracy threshold, it can be determined that the watermark extraction model training is completed.
[0072] Of course, when actually training the model, the number of training times or the number of training rounds can also be set. If the number of training times reaches a preset threshold, or the number of training rounds reaches a preset threshold, it can also be determined that the watermark extraction model training is completed.
[0073] Among them, the preset correct threshold, the preset number threshold and the preset round number threshold can all be set in advance by the administrator of the watermark extraction model training device.
[0074] It is understandable that if the watermark extraction model is not trained, the parameters of the watermark extraction model and the watermark information embedding model may be adjusted, and the process returns to step S10 to continue training.
[0075] If the watermark extraction model has been trained, it can be put into use to extract the watermark from the target image using the trained watermark extraction model. Of course, the watermark information embedding model can also be put into use. When actually needed, the watermark information embedding model can be used to embed specific watermark information into the image.
[0076] This embodiment provides a watermark extraction model training method. After the watermark is embedded in the image, one or more distortion simulation methods are used to perform noise simulation on the watermark-embedded image. This ensures that the generated noise image fits various actual scenarios as closely as possible, ensuring that the trained watermark extraction model can still correctly extract watermark information in various scenarios.
[0077] In a feasible implementation, the distortion simulation method includes perspective distortion simulation, and step S20 may include steps S21 to S23:
[0078] Step S21: performing text recognition on the watermark embedded image to determine a plurality of dispersed areas.
[0079] It should be noted that the scattered area may be an image area where text may exist.
[0080] In practice, perspective distortion simulation first requires vertex localization. For text images, this can be done by analyzing the image, identifying the text edges, and selecting vertices around the edges. This protects the content within the text area and reduces the impact of direct perturbations on the text.
[0081] In practical applications, we can first use the Sobel operator to calculate the gradient magnitude G(x, y) and gradient direction θ(x, y) of each pixel in the image, then:
[0082]
[0083] To simplify calculation, the gradient direction θ(x, y) may be quantized into four main directions, namely 0°, 45°, 90°, and 135°, to obtain the quantized direction θ'(x, y).
[0084] According to the quantized gradient direction, compare the gradient magnitude of the current pixel G(x, y) with the two adjacent pixels along the quantized direction θ'(x, y). If the gradient magnitude of the current pixel is not a local maximum, it is suppressed to 0. The specific logic can be represented as follows:
[0085]
[0086] The gradient image G is suppressed by non-maximum NMS Each pixel in (x, y) is compared with a preset threshold value, which is automatically selected by analyzing the grayscale histogram of the image to maximize the contrast between the foreground and the background. Pixels above the threshold are set to 1 (white) and pixels below the threshold are set to 0 (black), thus generating a binary image edge_intensity(x, y), which can be represented as:
[0087]
[0088] Using the connected component labeling algorithm, the pixels in the binarized edge image can be divided into several connected components, each of which corresponds to a possible text area.
[0089] L(x,y)=Label_Components(Edge_Intensity(x,y)) Formula (5)
[0090] Where L(x, y) is the labeling matrix, representing the connected component label to which each pixel in the image belongs. Label_Components() represents the connected component labeling algorithm, which iterates over each pixel and groups spatially connected pixels with the same value (usually 1, i.e., an edge) into a connected region. This algorithm is used to identify and label spatially connected groups of pixels in binary images.
[0091] Thus, a plurality of dispersed areas can be determined.
[0092] Step S22: merging the multiple scattered areas into a complete connected area.
[0093] Step S23: performing perspective distortion simulation processing on the connected area to generate a noise simulation image.
[0094] It should be noted that for each scattered area (i.e., the connected component mentioned above), all pixels belonging to the connected component can be found by scanning L(x, y), and then their minimum and maximum x, y values are calculated. These values define the boundaries of the connected component, and each value can be represented as:
[0095]
[0096] By merging the boundaries of all connected components, we can finally get the global bounding box:
[0097]
[0098] GlobalBoundingBox=(x global_min , xglobal_max ,y global_min ,y global_max ) Formula (6)
[0099] Among them, GlobalBoundingBox is the global bounding box.
[0100] Finally, points outside the text area are selected as vertices to reduce the direct impact of perspective transformation on the text. Each vertex V1-V4 can be represented as:
[0101] V1=(x global_min -δ,y global_min -δ)
[0102] V2=(x global_max +δ,y global_min -δ)
[0103] V3=(x global_min -δ,y gtobal_max +δ)
[0104] V4=(x global_max +δ,y global_max +δ)
[0105] After vertex positioning based on the text border, we can obtain four vertices V1(x1, y1), V2(x2, y2), V3(x3, y3), and V4(x4, y4). After random perturbation, each vertex becomes a new vertex coordinate V'1(x'1, y'1), V2(x'2, y'2), V3(x'3, y'3), and V'4(x'4, y'4). The perturbation of the vertex is achieved by adding or subtracting a random value δ within a fixed range (for example, ±6 pixels).
[0106] Next, the homography transformation parameters are calculated based on the original vertices and the perturbed vertices. The perspective transformation can be achieved by solving the following homography transformation equation, which allows points in one plane to be mapped to another plane:
[0107]
[0108] Among them, a0, b0, a1, b1, a2, b2, c1, c2 are transformation parameters, which can be calculated by setting the initial vertex and the transformed vertex coordinates and solving 8 linear equations.
[0109] Once all the transformation parameters are obtained, they can be applied to create a perspective-distorted version of the entire image. First, the above homography transformation equation is applied to each pixel. Finally, the transformed coordinates are sampled using the bicubic interpolation method. Bicubic interpolation improves image quality through the following interpolation function:
[0110]
[0111] Where h(t) is the cubic Hermite interpolation kernel and I(x, y) is the pixel value of the image at position (x, y).
[0112] After the above processing, a perspective-distorted image can be finally obtained, that is, a noise-simulated image generated through perspective distortion simulation processing.
[0113] In a feasible implementation, the distortion simulation method includes illumination distortion simulation, and step S20 may include steps S21' to S23':
[0114] Step S21 ′: performing text recognition on the watermark embedded image to determine a plurality of pixel-level dispersed areas.
[0115] It should be noted that pixel-level dispersion refers to areas of the image that may contain text. During screen capture, different shooting conditions will result in different lighting distributions in the captured image due to the influence of ambient lighting and screen lighting. To simulate these lighting conditions, light sources are divided into two categories: point light sources and line light sources.
[0116] In order to simulate the distortion of point light sources, a distribution weight matrix IW is needed p (x, y). The calculation formula of this matrix is as follows:
[0117]
[0118] Among them, (p x , p y ) are the coordinates of the simulated point light source, which can be randomly selected within the entire image range. maxDistance(p x , p y ) means from (p x , p y ) to the maximum value of the distances from the four vertex corners of the image. min Indicates the minimum ratio of illumination change, uniformly sampled from [0.6, 0.9]. max Indicates the maximum ratio of illumination change, uniformly sampled from [1.0,1.3].
[0119] Considering that text images are highly sensitive to lighting conditions, especially around text edges, illumination distortion can cause text to become blurred or distorted. Therefore, local enhancement can be performed based on the characteristics of the text region. This means giving greater weight to text regions (such as characters and edges) to ensure that the illumination distortion effect in these areas is more significant. A fully convolutional network (FCN) can be used to predict the geometric shape of the text region and the score of text presence.
[0120] First, the input image is input into a convolutional neural network to extract the feature map, then F = CNN(I), where F is the feature map, I is the input image, and CNN is a commonly used deep convolutional neural network, such as ResNet.
[0121] Next, a regression network consisting of a convolutional layer and a fully connected layer is used to predict the geometric shape G and score map S of each text region from the feature map F:
[0122] G,S=Regression(F) Formula (10)
[0123] Secondly, the final bounding box B of each text region (i.e., pixel-level scattered area) is generated by non-maximum suppression NMS, B = NMS (S, G).
[0124] Step S22 ′: performing single-time illumination distortion processing on other areas in the watermark-embedded image, and performing multiple-time illumination distortion processing on each pixel-level dispersed area to generate an illumination-distorted image.
[0125] It should be noted that after determining the bounding boxes B of each text area, the pixel values in the text area can be set to 1, and the pixel values in the non-text area can be set to 0, and the binary text mask matrix W of the input image can be obtained. text (x, y).
[0126] Based on the detected text area mask, the optimized point light source distortion weight matrix can be expressed as:
[0127]
[0128] Among them, α is the enhancement factor, which can be pre-set by the administrator of the watermark extraction model training equipment and applied to the text area to emphasize the lighting distortion effect of the text area. text When (x, y) = 1 (i.e. text area), the illumination distortion is enhanced by α times, i.e. multiple illumination distortion processing; when W text When (x, y)=0 (ie, non-text area), the illumination distortion is normal, ie, single illumination distortion processing is performed.
[0129] Step S22 ′: performing line light source distortion processing on the illumination-distorted image to generate a noise simulation image.
[0130] In actual use, in order to simulate the distortion of line light sources, the processing flow requires another weight distribution matrix IW1, which is the matrix from T0, T 90 , T 180 , T 270 A randomly selected matrix from :
[0131] IW l ~U{T o ,T 90 , T 180 , T 270 Formula (12)
[0132] Where: H represents the height of the image, and the calculation formula of T0(x, y) is:
[0133]
[0134] At this point, both the line light source and the point light source have been simulated, so the processed image can be used as a noise simulation image.
[0135] In a specific implementation, in order to make the edge of the text area clearer, step S22' in this embodiment may include:
[0136] performing line light source distortion processing on the illumination-distorted image to generate a light source-distorted image;
[0137] Performing bilateral filtering on the light source distorted image to generate a noise simulation image.
[0138] It should be noted that bilateral filtering has the advantage of smoothing images while preserving edges. Especially for text images, it can make the edges of text more obvious and enhance the difference between text areas (pixel values less than or equal to 128) and non-text areas (pixel values greater than 128).
[0139] To improve the illumination distortion simulation of text images, bilateral filtering is used for smoothing, better adapting to the characteristics of text images. In addition to considering the spatial distance between pixels, bilateral filtering also considers the difference between pixel values, thereby smoothing the image while preserving edge information, which is particularly important for text images.
[0140] The basic formula of the bilateral filter is as follows:
[0141]
[0142] Among them, I(x, y) is the pixel value of the original image at position (x, y), I' BW (x, y) is the pixel value of the filtered image at position (x, y), f r is a Gaussian function of intensity difference, used to adjust the weight caused by intensity difference. s is a Gaussian function of spatial variance, used to adjust the weights according to the spatial distance between pixels. p is a normalization factor that ensures that the sum of the weights is 1 and is calculated as:
[0143]
[0144] Here, k is the radius of the convolution kernel. In this example, k=3. k determines the size of the neighborhood considered in the filtering process.
[0145] For the illumination distortion weight matrix IW' p (x, y) and IW1, this method applies bilateral filtering to enhance its smoothness, especially in places where the text edges are kept clear while reducing noise or where the lighting changes are too drastic, which can be characterized as:
[0146] IW p_smooth =BilateralFilter(IW′ p (x,y)) Formula (16)
[0147] IW l_smooth =BilateralFilter(IW l (x,y)) Formula (17)
[0148] Among them, the specific applications of bilateral filtering are:
[0149]
[0150] in, is the standard deviation of the Gaussian function used for intensity differences, is the standard deviation of the Gaussian function used for spatial differences. W is a normalization factor given by the sum of all weights.
[0151] Final Lighting Distortion I D It is the point light source distortion IW after random selection of RandomChoice and smoothing operation p_smooth Or line light source distortion IW l_smooth The formula is as follows:
[0152] I D =RandomChoice{IW p_smoot , IW l_smooth Formula (19)
[0153] The illumination distortion matrix I D It will be used to simulate the impact of different lighting conditions on text images during screen capture.
[0154] In practical applications, the distortion simulation method may also include moiré distortion simulation and JPEG compression distortion.
[0155] In actual use, the execution process of moiré distortion simulation is as follows:
[0156] During screen capture, the mismatch between the screen display and the camera sampling frequency often causes moiré distortion in the captured image. This distortion appears as irregular textures on the image, seriously affecting image quality. This distortion can be simulated using the following equation:
[0157]
[0158] Among them, γ is uniformly sampled in [0,π], (z x , z y ) represents the coordinates of a point randomly sampled from the entire image, MD() represents the moiré distortion, and the final image after moiré distortion is obtained. MD ,After that, the image after moiré distortion can be used as a noise simulation image.
[0159] In actual use, the execution process of JPEG compression distortion is:
[0160] First, convert the image from RGB to YCbCr. In YCbCr, the Y channel represents luminance, while the Cb and Cr channels represent chrominance, corresponding to the color difference between blue and red, respectively. The steps for converting an image from RGB to YCbCr and performing JPEG compression are as follows:
[0161] Y=0.299R+0.587G+0.114B Formula (23)
[0162] Cb=-0.1687R-0.3313G+0.5B+128 Formula (24)
[0163] Cr=0.5R-0.4187G-0.0813B+128 Formula (25)
[0164] First, the image is divided into multiple 8x8 blocks, and the following DCT (Discrete Cosine Transform) operation is applied to the Y, Cb, Cr data of each block:
[0165]
[0166] Where I(x, y) represents the pixel value in the image block, and D(u, v) is the DCT coefficient.
[0167] Next, quantization is performed using the JPEG standard quantization matrix.
[0168]
[0169] Among them, Q(u, v) is the standard JPEG quantization matrix, which is used to adjust the compression rate and control the loss of image quality; D'(u, v) represents the DCT coefficient after the quantization operation, and τ represents the parameter that controls the smoothness of the function.
[0170] Next, we simulate the reconstruction process. During the reconstruction process, we need to perform an inverse operation on these quantized coefficients to restore the approximate original DCT coefficients. The DCT coefficients are then converted back to pixel space using the inverse discrete cosine transform (IDCT):
[0171] D(u,v)=D′(u,v)×Q(u,v) Formula (29)
[0172]
[0173] After all blocks are processed, the YCbCr to RGB color space is inversely converted for each pixel of the entire image to ensure accurate color restoration. The inverse conversion formula is as follows:
[0174] R=Y+1.402×(Cr-128) Formula (31)
[0175] G=Y-0.344136×(Cb-128)-0.714136×(Cr-128) Formula (32)
[0176] B=Y+1.772×(Cb-128) Formula (33)
[0177] After the color space is inversely converted, the image I is obtained after JPEG compression distortion. JD ;
[0178] The final complete screen noise layer is composed of the following noise simulation image I after the above noise processing no :
[0179] I no =δ1×I D ×I PD +δ2×I MD +δ3×I JD +G N Formula (34)
[0180] Among them, I D represents the illumination distortion matrix, I PD Represents the perspective distorted image, I MD Represents the image after moiré distortion, I JD Represents an image after JPEG compression distortion.
[0181] G N represents the main part of the remaining noise and is simulated using Gaussian noise δ1, δ2, and δ3 represent the corresponding distortion rates, which are set to 0.7, 0.15, and 0.15 by default and can be adjusted according to actual needs.
[0182] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 ,The watermark information embedding model includes a watermark information encoding layer and a watermark information embedding layer;
[0183] Step S10 includes steps S101 to S103:
[0184] Step S101: performing cyclic error correction coding on sample watermark information through the watermark information coding layer to generate a watermark bit sequence.
[0185] It should be noted that the embedded sample watermark information is a binary bit sequence m of length 45. Because the original embedded watermark information may be partially destroyed during image transmission or processing, the use of BCH encoding can effectively recover the watermark information at the receiving end, thereby improving the robustness of the watermark. Based on 45 information bits and 3 bits of error correction capability, BCH encoding can produce watermark data with a finite field order of 6 and a code length of 63.
[0186] First, for any binary bit string containing k information bits (k is 45), denoted as b k-1 , b k-2 ,...,b0. Among them, b k-1 is the most significant bit, b0 is the least significant bit, and can be expressed as an information polynomial i(x) containing all information bits, where x is the independent variable:
[0187] i(x)=b0+b1x 1 +b2x 2 +…+b k-1 x k-1 Formula (35)
[0188] Secondly, the generating polynomial g(x) is the core of constructing BCH code.
[0189] For the preset error correction capability t=3, the polynomial is composed of the minimum polynomial of the original element α and the least common multiple (LCM) of its power, which is expressed as follows:
[0190]
[0191] Where α is a finite field GF(2 6 ) of the original element, m αi (x) represents αi The minimum polynomials of these polynomials can be obtained by using GF(2 6 ) in calculating α,α 2 ,α 3 The continuous powers of and solving the corresponding polynomials are obtained.
[0192] Finally, encoding according to the information polynomial i(x) and the generator polynomial g(x) can generate the codeword c(x):
[0193] c(x)=(i(x)×g(x))mod(x n +1) Formula (37)
[0194] After expansion, c(x) can be expressed as follows:
[0195] c(x)=c0+c1x+c2x 2 +…+c 62 x 62 Formula (38)
[0196] Combining the coefficients can obtain the 63-bit watermark bit sequence m after BCH encoding. BCH .
[0197] Step S102: converting the watermark bit sequence through the watermark information embedding layer to generate multi-dimensional watermark information, wherein the multi-dimensional watermark information is information of the same size as the carrier image.
[0198] Step S103: performing channel-level fusion of the multi-dimensional watermark information and the carrier image through the watermark information embedding layer, and performing convolution processing on the fused data to generate a watermark-embedded image.
[0199] In actual use, the watermark bit sequence m can be BCH Expand to obtain a multidimensional feature vector w; perform an upsampling operation on the multidimensional feature vector w to obtain a watermark information w' of the same size as the carrier image, and merge the watermark information and the carrier image at the channel level to obtain a merged 6-channel fusion input input; use the U-shaped structure network to fusion network I fusion Perform convolution processing to obtain the first watermark image I with the same size as the carrier image R1 and embed the first watermark image into the image as a watermark.
[0200] In actual use, in order to improve the richness of watermark information and convert the discrete watermark bit sequence into a continuous multi-dimensional feature representation suitable for neural network processing, a multi-layer fully connected neural network is used to expand the original watermark sequence to obtain the multi-dimensional feature vector w of the watermark information.
[0201] w=W·m BCH +b formula (39)
[0202] Where W is the weight matrix of the fully connected layer of the neural network, with a shape of p×m. b is the bias vector, with a length of p. w is the expanded watermark sequence, with a length of p, where p is 256 and m is m. BCH The length is the same, 63.
[0203] Secondly, the multidimensional feature vector w of the watermark information is adjusted to the same size as the carrier image through upsampling technology, so that the multidimensional watermark information can be obtained, so that the watermark can be spatially aligned with the carrier image.
[0204] w′=Upsample(w,H,W) Formula (40)
[0205] The upsampling function Upsample() upsamples the multidimensional feature vector w according to the given carrier image height H and width W. w' is the upsampled watermark information, that is, the multidimensional watermark information, which has the same spatial dimension as the carrier image.
[0206] Finally, the channels are merged and the watermark information w' obtained by upsampling is combined with the carrier image I c Merge at the channel level to form a six-channel fused input I suitable for input to the encoder fusion . Fusion Input I fusion It not only retains the rich visual information of the carrier image, but also embeds the watermark features, providing strong support for deep feature extraction and watermark embedding in the subsequent encoder. Concatenate means splicing in the channel dimension. Two three-channel images will be spliced into an image with 6 channels. Then I fusion =Concatenate(I c , w')
[0207] Next, in order to better integrate the watermark content with the carrier image, it is necessary to extract the carrier image features and complete the deep watermark embedding based on the shallow and deep features of the carrier image. The carrier image feature extraction method adopts a U-shaped structure that combines downsampling and upsampling.
[0208] In the downsampling stage, the fused input I fusion First, the spatial resolution is gradually reduced through a series of convolutional layers, and nonlinear activation functions (such as ReLU) and pooling operations are applied after each convolution operation. These layers extract deep features of the image by reducing the image size.
[0209]
[0210] in, Represents the downsampled feature map of the kth layer. It is obtained by applying convolution, ReLU activation function and pooling operation to the feature map of the previous layer Each downsampling process reduces the image's length and width to half of its original size. Pool represents a pooling operation, which is used to reduce the spatial size of the feature map while retaining important features. ReLU is a nonlinear activation function that sets all negative values to 0, increasing the network's nonlinearity and alleviating the vanishing gradient problem. Conv represents a convolution operation, which extracts spatial features from the input feature map. represents the downsampled feature map of layer k-1. In the calculation of layer k, it is the input of the convolution operation. K represents the total number of layers in the downsampling stage, and here K is 4.
[0211] In the upsampling stage, the feature map obtained by downsampling The image is gradually restored to its original size through the deconvolution layer. In this process, by connecting the image features passed from the downsampling stage, it is possible to restore image details and fuse image features at different levels.
[0212]
[0213] in, Represents the upsampled feature map of the kth layer. This is obtained by applying the deconvolution operation to the combination of the upsampled feature map of the previous layer and the feature map of the corresponding downsampling stage. Deconv indicates that the deconvolution operation is used to increase the spatial resolution of the feature map from a smaller size to a larger size. The symbol represents the concatenation of feature maps, which can combine shallow features with deep features and make full use of features at different levels.
[0214] In addition, to compensate for the information loss caused by the downsampling operation and expand the receptive field of the feature to capture a wider range of contextual information, three dilated convolutional layers with dilation rates of 4, 8, and 16 are used in the middle stage of upsampling to increase the receptive field of the feature.
[0215]
[0216] F atrous Represents the feature map after dilated convolution, integrating contextual information from different scales. Represents the concatenation of feature maps in the channel dimension, which is used to fuse feature maps generated by different dilation rates. AtrousConv represents the operation of dilated convolution, where r is the dilation rate of dilated convolution. Represents the downsampled feature map on the Lth layer, which is used as the input of the dilated convolution. r∈{4,8,16} represents the three different expansion rates used.
[0217] After the above image feature extraction and fusion operations, the obtained feature map contains comprehensive information from different network layers, which needs to be further processed through a series of convolutional layers. The purpose of these convolutional layers is to adjust and optimize the feature map in order to generate the first watermark image I that can retain the carrier image and watermark information. w1 .
[0218] In a specific implementation, in order to improve the effect of watermark embedding, step S103 in this embodiment may include:
[0219] The multi-dimensional watermark information is fused with the carrier image at the channel level through the watermark information embedding layer, and convolution processing is performed on the fused data to generate a first watermark image;
[0220] performing weighted processing on the first watermark image through the watermark information embedding layer to generate a second watermark image, wherein the weighted processing is performed based on a gradient mask matrix obtained by performing gradient recognition on the first watermark image;
[0221] The carrier image and the second watermark image are superimposed through the watermark information embedding layer to generate a watermark embedded image.
[0222] It should be noted that before generating the final watermark image, the area where the watermark is embedded needs to be confirmed. First, the Sobel operator is used to calculate the gradient map of the input image. The Sobel operator is an edge detection operator that calculates the spatial gradient of image brightness, thereby highlighting the structural boundaries in the image.
[0223]
[0224] Among them, Sobel x and Sobel y Represents the Sobel operators in the horizontal and vertical directions respectively. Represents the gradient intensity map of the carrier image after calculation by the Sobel operator, I c Represents a carrier image.
[0225] Afterwards, the gradient intensity map of the carrier image is used To generate the adjusted gradient mask matrix MASK v ,This map determines the watermark strength in each region to reduce visual noise and improve the ,invisibility of the watermark.
[0226]
[0227] Here, the Normalize function normalizes the gradient to the range [0, 1]. Higher gradient values indicate that image edges will be given lower weights, meaning that these areas will carry less watermark information to maintain image quality.
[0228] Finally, based on MASK v , generate the second watermark image I w2 :
[0229] I w2 =I w1 MASK v Formula (47)
[0230] Here, Iw1 is the first watermark image obtained by convolution operation based on the feature map, and is multiplied by MASK v Ensure that the residual image has a lower influence in areas with complex textures and a higher influence in smooth areas.
[0231] Afterwards, the second watermark image I w2 With the original carrier image I c Add together to generate the final watermark embedded image I watermarked , then
[0232] I watermarked =I c +I w2 Formula (48)
[0233] In the specific implementation, since the text image requires high readability, it is crucial to ensure the concealment of the embedded watermark and the image quality. Based on this, the watermark information embedding model can also be equipped with a watermark concealment enhancement module, which uses the adversarial network (PatchGAN) represented as Ad to ensure the invisibility of the generated embedded watermark image and the visual quality of the original carrier image. For the complete encoding model of the watermark information embedding layer, we use E to represent it, θ E Represents the parameters of the encoder, m BCH The watermark information is encoded by BCH in the watermark information encoding layer. Then the watermark information embedding layer can be expressed as:
[0234] I watermarked =E(θ E ,I c ,m BCH ) Formula (49)
[0235] Next, use the adversarial network Ad to distinguish the watermark embedded image I watermarked and the original carrier image I c First, the encoded image attempts to mislead the adversarial network into producing an image similar to the original support image. AdBy adjusting θ E Then improve I watermarked image quality.
[0236] L Ad =log(Ad(I watermarked ))=log(Ad(E(θ E ,I c , m BCH Formula (50)
[0237] At the same time, the parameters of the adversarial network Ad need to be updated to correctly classify I watermarked with I c :
[0238] L Adv =log(1-Ad(θ Ad , I watermarked ))=log(1-Ad(θ Ad ,E(L c , m BCH Formula (51)
[0239] In the specific implementation, the watermark information may be encoded during the embedding process. In order to correctly extract it, the watermark extraction model can include a watermark information extraction layer and a watermark information decoding layer:
[0240] The watermark information extraction layer needs to extract the watermark information from the noise simulated image I no The decoder used in this process includes a preliminary six-layer convolutional structure. Each layer includes basic convolution operations, batch normalization, and ReLU activation functions, followed by a maximum pooling operation to gradually extract image features and reduce the spatial size of the feature map.
[0241] F l =MaxPool(ReLU(B(Conv(F l-1 Formula (52)
[0242] Among them F l Represents the feature map of layer l.
[0243] After the six-layer convolutional structure is processed, the final feature map F6 needs to be converted to a one-dimensional form for further neural network processing. The flattening operation can convert the multi-dimensional feature map into a one-dimensional vector so that it can be input into the subsequent fully connected layer.
[0244] F flat =Flatten(F6) Formula (53)
[0245] Then it is further compressed through two fully connected layers, and finally outputs a watermark information m BCH To map the output value to a binary value close to 0 or 1, the final output layer uses the Sigmoid activation function.
[0246] m out =Sigmoid(ReLU(Dense(ReLU(Dense(F flat Formula (54)
[0247] Since the watermark information is a binary bit sequence, it can be regarded as a two-class classification problem. Therefore, the cross entropy loss function is used to calculate the decoded watermark information M' i and the original watermark information (ie, sample watermark information) M' i The difference between .
[0248]
[0249] Where n is the number of watermark bits.
[0250] During the training phase, since the entire network architecture and each stage of the noise layer are differentiable, the entire model can be trained end-to-end. The final loss L of the entire network includes the aforementioned adversarial image loss L Ad and decoding loss L D , which can be expressed as:
[0251] L=λ1L D +λ2L Ad Formula (56)
[0252] Here, λ1 and λ2 represent the weights of decoding loss and adversarial image loss, respectively, with default values of 0.8 and 0.2.
[0253] The watermark information decoding layer uses the end-to-end deep learning framework to obtain the watermark information m out After that, BCH decoding is needed to restore all its information. First, the received watermark information m out Convert the codeword r(x) into polynomial form:
[0254]
[0255] Next, divide r(x) by g(x) to determine if there is an error (if the remainder is non-zero). If the remainder is non-zero, an error has occurred, and the Berlekamp-Massey algorithm can be used to locate and correct the error. BCH decoding and restoration ensure the accuracy of the embedded watermark information, ultimately allowing the correct and complete watermark information to be extracted from the watermarked image after the screen capture process.
[0256] This embodiment provides a watermark extraction model training method. Since watermark encoding is performed during watermark embedding, it ensures that the extracted watermark information can be correctly compared even after a noise attack, thereby improving the error correction capability of the watermark information.
[0257] This application also provides a watermark extraction model training device, please refer to Figure 3 , the watermark extraction model training device includes:
[0258] The embedding module 10 is used to embed the sample watermark information into the carrier image through the watermark information embedding model to generate a watermark embedded image;
[0259] a simulation module 20, configured to perform noise simulation on the watermark embedded image to generate a noise simulated image, wherein the noise simulation is performed using at least one distortion simulation method;
[0260] An extraction module 30 is configured to extract a watermark from the noise simulation image using a watermark extraction model to obtain watermark information;
[0261] The determination module 40 is configured to construct an extraction determination result according to the sample watermark information and the watermark information, and determine whether the watermark extraction model has been trained based on the extraction determination result.
[0262] The watermark extraction model training device provided in this application, employing the watermark extraction model training method of the aforementioned embodiment, can address the technical issue in the prior art where noise simulation is insufficiently comprehensive, resulting in the model's difficulty in correctly extracting watermark information from images of specific scenes. Compared to the prior art, the beneficial effects of the watermark extraction model training device provided in this application are the same as those of the watermark extraction model training method of the aforementioned embodiment. Other technical features of the watermark extraction model training device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0263] The present application provides a watermark extraction model training device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the watermark extraction model training method in the above-mentioned embodiment 1.
[0264] Reference below Figure 4, which shows a schematic diagram of the structure of a watermark extraction model training device suitable for implementing the embodiments of the present application. The watermark extraction model training device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The watermark extraction model training device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0265] like Figure 4 As shown, the watermark extraction model training device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the watermark extraction model training device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the watermark extraction model training device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a watermark extraction model training device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.
[0266] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0267] The watermark extraction model training device provided in this application, employing the watermark extraction model training method of the aforementioned embodiment, can address the technical issue of the prior art where noise simulation is incomplete, resulting in the model's difficulty in correctly extracting watermark information from images of specific scenes. Compared to the prior art, the beneficial effects of the watermark extraction model training device provided in this application are the same as those of the watermark extraction model training method of the aforementioned embodiment. Other technical features of the watermark extraction model training device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0268] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0269] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0270] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the watermark extraction model training method in the above embodiment.
[0271] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0272] The computer-readable storage medium may be included in the watermark extraction model training device; or it may exist independently without being assembled into the watermark extraction model training device.
[0273] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the watermark extraction model training device, the watermark extraction model training device: embeds sample watermark information into the carrier image through the watermark information embedding model to generate a watermark embedded image; performs noise simulation on the watermark embedded image to generate a noise simulated image, wherein the noise simulation is performed using at least one distortion simulation method; performs watermark extraction on the noise simulated image through the watermark extraction model to obtain watermark information; constructs an extraction judgment result according to the sample watermark information and the watermark information, and determines whether the watermark extraction model has been trained based on the extraction judgment result.
[0274] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0275] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0276] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0277] The computer-readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the watermark extraction model training method described above. This computer-readable storage medium can address the technical issue in the prior art where noise simulation is insufficiently comprehensive, making it difficult for the model to correctly extract watermark information for images in specific scenes. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the watermark extraction model training method provided in the above-mentioned embodiment, and are not further elaborated here.
[0278] The present application also provides a computer program product, comprising a computer program, which implements the steps of the watermark extraction model training method as described above when the computer program is executed by a processor.
[0279] The computer program product provided in this application can address the technical problem of incomplete noise simulation in existing technologies, which can make it difficult for the model to correctly extract watermark information for images in specific scenes. Compared with existing technologies, the beneficial effects of the computer program product provided in this application are the same as those of the watermark extraction model training method provided in the above-mentioned embodiments, and will not be elaborated here.
[0280] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A watermark extraction model training method, characterized in that: The watermark extraction model training method includes: Embed the sample watermark information into the carrier image through the watermark information embedding model to generate a watermark embedded image; Performing noise simulation on the watermark embedded image to generate a noise simulated image, wherein the noise simulation is performed using at least one distortion simulation method; Performing watermark extraction on the noise simulation image through a watermark extraction model to obtain watermark information; An extraction determination result is constructed according to the sample watermark information and the watermark information, and whether the watermark extraction model is trained is determined based on the extraction determination result.
2. The watermark extraction model training method according to claim 1, characterized in that: The distortion simulation method includes perspective distortion simulation; The performing noise simulation on the watermark embedded image to generate a noise simulated image includes: Performing text recognition on the watermark-embedded image to determine a plurality of dispersed regions, wherein the dispersed regions are image regions where text may be present; merging the plurality of scattered regions into a complete connected region; Perform perspective distortion simulation processing on the connected area to generate a noise simulation image.
3. The watermark extraction model training method according to claim 1, characterized in that: The distortion simulation method includes illumination distortion simulation; The performing noise simulation on the watermark embedded image to generate a noise simulated image includes: Performing text recognition on the watermark embedded image to determine a plurality of pixel-level dispersed regions, wherein the pixel-level dispersed regions are image regions at the pixel level where text may exist; Performing single-time illumination distortion processing on other areas in the watermark-embedded image, and performing multiple-time illumination distortion processing on each of the pixel-level dispersed areas to generate an illumination-distorted image; Performing line light source distortion processing on the illumination-distorted image to generate a noise-simulated image.
4. The watermark extraction model training method according to claim 3, characterized in that: The performing line light source distortion processing on the illumination-distorted image to generate a noise simulation image includes: performing line light source distortion processing on the illumination-distorted image to generate a light source-distorted image; Performing bilateral filtering on the light source distorted image to generate a noise simulation image.
5. The watermark extraction model training method according to any one of claims 1 to 4, characterized in that: The watermark information embedding model includes a watermark information encoding layer and a watermark information embedding layer; The method of embedding the sample watermark information into the carrier image through the watermark information embedding model to generate the watermark embedded image includes: Performing cyclic error correction coding on the sample watermark information through the watermark information coding layer to generate a watermark bit sequence; Converting the watermark bit sequence through the watermark information embedding layer to generate multi-dimensional watermark information, wherein the multi-dimensional watermark information is information of the same size as the carrier image; The multi-dimensional watermark information is fused with the carrier image at the channel level through the watermark information embedding layer, and convolution processing is performed on the fused data to generate a watermark embedded image.
6. The watermark extraction model training method according to claim 5, characterized in that: The multi-dimensional watermark information is fused with the carrier image at the channel level through the watermark information embedding layer, and convolution processing is performed on the fused data to generate a watermark embedded image, including: The multi-dimensional watermark information is fused with the carrier image at the channel level through the watermark information embedding layer, and convolution processing is performed on the fused data to generate a first watermark image; performing weighted processing on the first watermark image through the watermark information embedding layer to generate a second watermark image, wherein the weighted processing is performed based on a gradient mask matrix obtained by performing gradient recognition on the first watermark image; The carrier image and the second watermark image are superimposed through the watermark information embedding layer to generate a watermark embedded image.
7. A watermark extraction model training device, characterized in that: The watermark extraction model training device comprises: An embedding module, used for embedding sample watermark information into a carrier image through a watermark information embedding model to generate a watermark-embedded image; a simulation module, configured to perform noise simulation on the watermark embedded image to generate a noise simulated image, wherein the noise simulation is performed using at least one distortion simulation method; An extraction module, configured to extract watermarks from the noise simulation image using a watermark extraction model to obtain watermark information; A determination module is used to construct an extraction determination result according to the sample watermark information and the watermark information, and determine whether the watermark extraction model is trained based on the extraction determination result.
8. A watermark extraction model training device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the watermark extraction model training method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the watermark extraction model training method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the watermark extraction model training method according to any one of claims 1 to 6 are implemented.