Watermark embedding method, related device, equipment and storage medium

Through the dual-splitting collaborative watermark embedding method, offset prediction deformation convolution and feature interaction are used to solve the problems of weak attack resistance and difficult to guarantee visual quality in the existing watermark embedding technology, and achieve high visual quality and robust watermark embedding.

CN120278870AActive Publication Date: 2025-07-08IFLYTEK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510758988.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing watermark embedding technology has weak attack resistance and difficult to guarantee visual quality, making it difficult to meet application needs.

Method used

By combining offset prediction deformation convolution and feature interaction based on the original image feature encoding and target string feature projection, the dual-splitting collaborative watermark embedding method, including feature fusion decoding, ensure the visual quality after watermark embedding and improve the attack resistance.

Benefits of technology

While maintaining high visual quality, it significantly improves the robustness and concealment of watermark embedding, and can maintain excellent robustness under new types of attacks that have not been seen before.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278870A_ABST
    Figure CN120278870A_ABST
Patent Text Reader

Abstract

The invention discloses a watermark embedding method, a related device, equipment and a storage medium, and the method comprises the steps: carrying out the feature coding based on an original image, obtaining the image features of the original image, carrying out the feature projection based on a to-be-embedded target character string, and obtaining the sequence features of the target character string; performing deformation convolution on the image features on the basis of offset obtained by performing offset prediction on the image features to obtain deformation features of the image features, performing feature interaction with the sequence features on the basis of the deformation features to obtain first features, and performing modulation convolution on the basis of the image features and the sequence features to obtain second features; and performing fusion decoding based on the first feature and the second feature to obtain a target character string which is used as a target image after the image watermark is embedded into the original image. According to the scheme, the visual quality after watermark embedding can be guaranteed as much as possible, and the anti-attack capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a watermark embedding method and related devices, equipment, and storage media. Background Art

[0002] As an important means of digital copyright protection, digital image watermarking technology has broad application prospects in fields such as multimedia content authentication and intellectual property protection.

[0003] Currently, existing watermark embedding technologies have technical problems such as weak anti-attack ability and easy generation of visual artifacts, resulting in difficult-to-guarantee visual quality, and thus it is difficult to meet application requirements. In view of this, how to ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a watermark embedding method and related devices, equipment, and storage media, which can ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability.

[0005] To solve the above technical problem, in the first aspect of this application, a watermark embedding method is provided, including: performing feature encoding on an original image to obtain image features of the original image, and performing feature projection on a target string to be embedded to obtain sequence features of the target string; performing deformable convolution on the image features based on the offset obtained by offset prediction of the image features to obtain deformed features of the image features, and performing feature interaction based on the deformed features and the sequence features to obtain first features, and performing modulation convolution based on the image features and the sequence features to obtain second features; performing fusion decoding based on the first features and the second features to obtain the target string as the target image after embedding the image watermark into the original image.

[0006] To solve the above technical problem, in the second aspect of this application, a watermark embedding device is provided, including: a feature preparation module, a splitting processing module, and a fusion decoding module. The feature preparation module is used to perform feature encoding on an original image to obtain image features of the original image, and perform feature projection on a target string to be embedded to obtain sequence features of the target string; the splitting processing module is used to perform deformable convolution on the image features based on the offset obtained by offset prediction of the image features to obtain deformed features of the image features, and perform feature interaction based on the deformed features and the sequence features to obtain first features, and perform modulation convolution based on the image features and the sequence features to obtain second features; the fusion decoding module is used to perform fusion decoding based on the first features and the second features to obtain the target string as the target image after embedding the image watermark into the original image.

[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other. At least program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the watermark embedding method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the watermark embedding method in the first aspect above.

[0009] In the above solution, feature encoding is performed based on the original image to obtain the image features of the original image, and feature projection is performed based on the target string to be embedded to obtain the sequence features of the target string. Then, based on the offset obtained by offset prediction of the image features, deformable convolution is performed on the image features to obtain the deformed features of the image features, and feature interaction is performed based on the deformed features and the sequence features to obtain the first features. In addition, modulation convolution is performed based on the image features and the sequence features to obtain the second features. Furthermore, fusion decoding is performed based on the first features and the second features to obtain the target string as the target image after the image watermark is embedded in the original image. Therefore, after obtaining the image features and the sequence features, on the one hand, in one path of the feature processing flow, offset prediction is performed according to the image features to obtain the offset, and deformable convolution is performed on the image features based on this, which can adaptively capture the local feature structure of the original image. Subsequently, when performing feature interaction, it helps to make the watermark embedding avoid the visually perceptible areas as much as possible, that is, it can ensure the visual quality after watermark embedding as much as possible. On the other hand, in the other path of the feature processing flow, modulation convolution is performed through the image features and the sequence features, and fusion decoding is performed with the aforementioned one path of the feature processing flow, which can realize dual-path collaborative watermark embedding and help improve the anti-attack ability. Therefore, it can ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flowchart of an embodiment of the watermark embedding method of the present application; Figure 2a is a schematic diagram of the process of an embodiment of the watermark embedding method of the present application; Figure 2b is a schematic diagram of the process of an embodiment of the watermark attack process of the present application; Figure 3 is a schematic framework diagram of an embodiment of the watermark embedding device of the present application; Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application; Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] The following will combine the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of the present application.

[0012] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0013] The terms "system" and "network" are often used interchangeably herein. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "multiple" in this article means two or more than two.

[0014] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the watermark embedding method of the present application. Specifically, it may include the following steps: Step S11: Perform feature encoding on the original image to obtain the image features of the original image, and perform feature projection on the target string to be embedded to obtain the sequence features of the target string.

[0015] In an implementation scenario, the original image may be the image data for embedding an image watermark, and specifically may include but is not limited to: video frame images of movies and TV shows, scanned images of paper paintings, electronic painting images, etc. The specific content of the original image is not limited herein.

[0016] In an implementation scenario, an encoder may be used to perform feature encoding on the original image to compress the original image into the latent space and obtain the image features of the original image. It should be noted that the encoder may include but is not limited to a convolutional neural network, etc. The network structure of the encoder is not limited herein. For the convenience of description, taking the original image as an RGB image with a resolution of 256*256 as an example, the original image may be denoted as I∈[0,1] 256*256*3, where the superscript 3 represents the three channels of R, G, and B, and [0, 1] indicates that the pixel value is between 0 and 1. That is, the original image can be represented as a matrix with 3 channels and a size of 256 * 256. Of course, the above example is only a possible representation of the original image, and does not limit the resolution and number of channels of the original image. In addition, in addition to encoding the features of the original image through an encoder, the features of the original image can also be encoded through manually designed feature operators such as histogram of oriented gradients, local binary pattern, scale-invariant feature transform, speeded up robust features, etc. Or, the features of the original image can also be encoded through statistical learning-based coding methods such as principal component analysis and independent component analysis. The specific method of encoding the features of the original image is not limited, and no further examples will be given. For the sake of convenience of description, still taking the aforementioned original image I as an example, the image features obtained by its feature encoding can be denoted as z img ∈R 4*32*32 .

[0017] In an implementation scenario, the target string to be embedded can be a binary string, which can be denoted as s ∈ {0, 1} 64 , that is, any character in the target string is binary 0 or binary 1, and the length of the target string is 64. Of course, the above example is only a possible example of the target string. For example, the target string to be embedded can also be a decimal string, and even the string to be embedded can include letters. Other possible situations are not limited here, and no further examples will be given.

[0018] In an implementation scenario, in order to perform feature projection on the target string to be embedded, the hash network can be first used to process the target string to obtain a conditional vector. For the sake of convenience of description, still taking the aforementioned target string s as an example, the conditional vector can be expressed as c = H ψ (s). Exemplarily, the conditional vector c can be a 128-dimensional vector. Then, the conditional vector can be projected to the projection features that are an integer multiple of the total number of single-channel elements of the aforementioned image features through a fully connected layer. Still taking the aforementioned image feature z img as an example (the total number of single-channel elements is 32 * 32 = 1024), the above 128-dimensional conditional vector is projected through a fully connected layer. For example, 2048-dimensional projection features can be obtained. On this basis, reshaping can be performed based on the projection features to obtain sequence features with the same resolution as the image features. Still taking the aforementioned image feature z img as an example, the aforementioned 2048-dimensional projection features can be reshaped to obtain 64 * 32 * 32-dimensional sequence features wm featIn this way, the global distribution of the watermark information can be repeatedly extended in space, enabling the complete watermark information to be perceived at each position in space. Of course, the above example is only one possible implementation manner of feature projection in the actual application process, and other possible implementation manners are not limited herein (for example, a neural network model including network layers such as convolutional layers can be directly used to perform feature projection on the target sequence, etc.), and no further examples will be given one by one.

[0019] Step S12: Perform deformable convolution on the image features based on the offset obtained by offset prediction from the image features to obtain the deformed features of the image features, perform feature interaction based on the deformed features and the sequence features to obtain the first feature, and perform modulated convolution based on the image features and the sequence features to obtain the second feature.

[0020] In an implementation scenario, to obtain the deformed features of the image features, offset prediction can be first performed based on the image features to obtain the offset, and the offset can specifically include the offset values of each feature position in the image features at each convolution position within the convolution range. On this basis, during the execution of the deformable convolution, for each feature position, the weighted sum of the image sub-features at the target positions after the convolution positions in the image features are offset by the offset values can be respectively weighted based on the weight factors of each convolution position within the convolution range of the feature position to obtain the deformed sub-features of the feature position, and then the deformed features can be obtained based on the deformed sub-features of each feature position in the image features. The above method can refine the granularity of feature processing as much as possible during the process of adaptively capturing the local feature structure of the original image by performing feature processing such as offset prediction and deformable convolution with the feature position as the basic unit, which helps to improve the fine degree of capturing the local feature structure.

[0021] In a specific implementation scenario, offset prediction can be performed on the image features based on a neural network model to obtain the offset values of each feature position in the image features at each convolution position within the convolution range. It should be noted that the neural network model can include but is not limited to network layers such as standard convolutional layers, and the network structure of the neural network model used for offset prediction is not limited herein, and no further examples will be given one by one. In addition, the offset value at the convolution position is used to represent the offset position of the local feature structure relative to the convolution position, so as to assist in capturing the local feature structure of the original image subsequently. As a special example, in order to prevent possible damage to the semantic content due to excessive deformation, the offset can also be constrained during the offset prediction process, such as regularization constraint with ||Δp||2 < 1.5.

[0022] In a specific implementation scenario, taking the use of a 3*3 convolutional kernel as an example, the convolution range is 3*3. For any feature position, the 3*3 range centered on this feature position is the convolution range (in other words, there are 9 different offset values for each feature position). For any convolution position within the 3*3 convolution range, corresponding offset values are predicted (which are the offset values of the image feature in the width dimension and height dimension outside the channel dimension). For the sake of description, still taking the aforementioned image feature z img ∈R 4*32*32 as an example, in the case of using a 3*3 convolutional kernel, the offset can be expressed as Δp ∈ R 2*9*32*32 . Of course, the above example is only one possible example of the convolution range and offset in the actual application process. Other possible situations are not limited here and will not be exemplified one by one.

[0023] In a specific implementation scenario, still taking the use of a 3*3 convolutional kernel as an example, after obtaining the offset, for any feature position p in the image feature, within the 3*3 convolution range centered on the feature position p, the weighted sum of the image sub-features at the target positions after offsetting the corresponding convolution positions by the offset values using the weight factors of each convolution position can be calculated to obtain the deformed sub-feature F out (p):

[0024] In the above formula, W(k) represents the weight factor at the convolution position k within the 3*3 convolution range centered on the feature position p, F in represents the image feature, p k represents the preset offset value at the convolution position k within the 3*3 convolution range centered on the feature position p (such as, [0,0], [0,1], [1,0], [0,-1], [-1,0], [1,1], [1,-1], [-1,1], [-1,-1]), Δp(k,p) represents the aforementioned predicted offset value at the convolution position k within the 3*3 convolution range centered on the feature position p, then p + p k + Δp(k,p) represents the target position, and F in (p + p k + Δp(k,p)) represents the image sub-feature of the image feature F in at the target position.

[0025] In a specific implementation scenario, after obtaining the deformed sub-features at each feature position in the image feature, the deformed sub-features at each feature position can be combined according to the original orientation of each feature position in the image feature to obtain the deformed feature of the image feature. It should be noted that the above processing flow of the deformed convolution can be repeated multiple times. For example, after obtaining the deformed feature, it can be used as the new image feature, and the steps of predicting the offset based on the image feature to obtain the offset amount can be returned and executed for the new image feature, and so on in a loop until the deformed convolution is executed for the target number of times (such as 3 times), and the latest deformed feature can be used as the final deformed feature of the image feature. In addition, to facilitate the implementation of the deformed convolution, the above steps of the deformed convolution execution process can be implemented through a feature extraction network composed of several groups (such as 3 groups) of deformable convolutions. For the technical details of the deformable convolution, please refer to the relevant content, and will not be elaborated here.

[0026] In an implementation scenario, after obtaining the deformed feature of the image feature, it can be used to perform feature interaction with the sequence feature to obtain the first feature. It should be noted that the first feature obtained through the above steps of deformed convolution, feature interaction, etc. in the embodiments of the present disclosure integrates the feature information of both the image and the watermark. As a possible implementation manner, an attention mechanism such as a cross-attention mechanism can be directly used to perform feature interaction on the deformed feature and the sequence feature to establish the context relationship between the watermark information and the global image, and strengthen the association degree of the texture complex regions. It should be noted that for the specific process of feature interaction in this case, please refer to the technical details of attention mechanisms such as the cross-attention mechanism, and will not be elaborated here.

[0027] In an implementation scenario, as another possible implementation manner, different from directly using the attention mechanism to perform feature interaction on the image feature and the sequence feature, it is also possible to first splice the deformed feature and the sequence feature in the channel dimension to obtain a fused feature, and then divide the fused feature in the resolution dimension to obtain several fused sub-features. On this basis, for each fused sub-feature, a self-attention mechanism introducing positional encoding can be used to process the fused sub-feature to obtain an enhanced sub-feature of the fused sub-feature, and the enhanced sub-features of each fused sub-feature can be combined to obtain an enhanced feature. Furthermore, the enhanced feature can be fused with the sequence feature to obtain the first feature. In the above manner, by introducing positional encoding in each fused sub-feature through window attention for self-attention mechanism processing, the relationship between the watermark information and the global image context can be established, and the association strength of the texture complex regions can be strengthened.

[0028] In a specific implementation scenario, for the sake of description, let the deformed feature be denoted as F def ∈R 64*32*32 as an example, then for the 64*32*32-dimensional sequence feature wmfeat In terms of this, the two can be concatenated in the channel dimension to obtain a fused feature of 128 * 32 * 32 dimensions. Then, the above fused feature is divided in the resolution dimension. Taking the example of dividing it into 8 * 8 small windows, 16 fused sub-features of 128 * 8 * 8 dimensions can be obtained. Of course, the above example is only a possible example of dividing fused sub-features in the actual application process, and other possible situations will not be exemplified one by one here.

[0029] In a specific implementation scenario, the attention mechanism adopted by the fused sub-features may include but is not limited to the multi-head self-attention mechanism (e.g., the number of heads n head = 4), which is not limited here. In addition, for each fused sub-feature, based on the self-attention mechanism with positional encoding introduced, an enhanced sub-feature can be obtained:

[0030] In the above formula, Q, K, and V respectively represent the query feature after the fused sub-feature is transformed by the query matrix, the key feature after being transformed by the key matrix, and the value feature after being transformed by the value matrix. The superscript T of K represents the transpose, d represents the feature dimension, and B represents the positional encoding. That is to say, in the process of the self-attention mechanism, after the interaction sub-feature between the query feature and the key feature of the fused sub-feature is fused with the positional encoding, it interacts with the value feature of the fused sub-feature to obtain the enhanced sub-feature of the fused sub-feature. In addition, as a possible example, the positional encoding can be specifically generated by learnable parameters, which is not limited here.

[0031] In a specific implementation scenario, after obtaining the enhanced sub-features of each fused sub-feature, the enhanced sub-features of each fused sub-feature can be combined according to the original positions of each fused sub-feature to obtain the enhanced feature. For the convenience of description, the enhanced feature can be denoted as F main . Exemplarily, still taking the aforementioned sequence feature and image feature as an example, an enhanced feature such as F main ∈ R 128*32*32 can be finally obtained. Of course, the above example is only a possible example of the enhanced feature, and other possible situations of the enhanced feature are not limited here, nor will they be exemplified one by one.

[0032] In a specific implementation scenario, after obtaining the enhanced features, the enhanced features and the sequence features can be fused to obtain the first feature. Specifically, weight prediction can be performed based on the enhanced features and the sequence features to obtain a weight image for fusing the enhanced features and the sequence features. The weight image has the same resolution as the image features and the sequence features. Then, the enhanced features and the sequence features are weighted based on the weight image to obtain the first feature. Exemplarily, the enhanced features and the sequence features can be concatenated in the channel dimension first. Since the number of channels increases at this time, one-dimensional convolution can be used to reduce the number of channels, and an activation function such as sigmoid is used to constrain the output data obtained after one-dimensional convolution to the range of 0 to 1 to serve as the weight image. For ease of description, the weight image G can be expressed as σ(conv 1*1 ([F main ,wm feat )). On this basis, the enhanced features and the sequence features can be weighted and summed using the weight image to obtain the first feature z main :

[0033] In the above formula, z main represents the first feature, F main represents the enhanced feature, wm feat represents the sequence feature, represents the dot product operation, G represents the weight image, and 1 - G represents another weight image obtained by subtracting each element in the weight image from 1, that is, the sum of the corresponding elements in this weight image and the foregoing weight image G is 1. Exemplarily, the first feature z main ∈R 64 *32*32 . It should be noted that the above example is only a possible example of the first feature in the actual application process, and other possible situations are not listed one by one here. In the above method, weight prediction is performed based on the enhanced features and the sequence features to obtain a weight image for fusing the enhanced features and the sequence features. The weight image has the same resolution as the image features and the sequence features. Then, the enhanced features and the sequence features are weighted based on the weight image to obtain the first feature. Therefore, it can adaptively focus on fusing the enhanced features and the sequence features.

[0034] In an implementation scenario, similar to the aforementioned first feature, the second feature in the embodiments of the present disclosure also incorporates the feature information of both the image and the watermark. The main difference from the first feature lies in their different implementation manners. Specifically, in order to perform modulated convolution on the image feature and the sequence feature, a first prediction can be made based on the image feature to obtain the convolution parameters of the modulated convolution, and a second prediction can be made based on the image feature and the sequence feature to obtain the modulation parameters of the modulated convolution. The convolution parameters can include weight parameters and bias parameters, and the modulation parameters can include the intensity values at each feature position when performing the modulated convolution. On this basis, the image feature can be subjected to modulated convolution based on the convolution parameters and the modulation parameters to obtain the second feature. It should be noted that for the specific process of the modulated convolution, reference can be made to the technical details of the modulated convolution, which will not be elaborated here. Through the first prediction and the second prediction in the above manner, the convolution parameters and the modulation parameters are respectively obtained, and then the modulated convolution is performed accordingly, which can adaptively perform the modulated convolution in terms of intensity, contribute to the precise modulation of the feature channels, and can implicitly control the watermark intensity.

[0035] In a specific implementation scenario, a multi-layer perceptron can be used to perform the first prediction on the image feature to obtain the convolution parameters. For ease of description, the weight parameter in the convolution parameters can be denoted as W dyn , and the bias parameter in the convolution parameters can be denoted as b. As a possible example, the weight parameter W dyn ∈ R 64*64*3*3 , and the bias parameter b ∈ R 64 . Of course, the above example is only a possible example in the actual application process, and other possible situations will not be exemplified one by one here.

[0036] In a specific implementation scenario, the image feature and the sequence feature can be first concatenated in the channel dimension, and then the first convolution operation, the activation function operation, the second convolution operation, and the normalization operation are sequentially performed to obtain the modulation parameters. For ease of description, the modulation parameters can be denoted as M, and it can be expressed as M = σ(conv(GELU((conv([z img , wm feat ))))), where the square brackets indicate concatenation in the channel dimension, and from the inside out are the first convolution conv, the activation function operation GELU, the second convolution conv, and the normalization operation σ. As a possible example, the modulation parameter M ∈ R 1*32*32 . Of course, the above example is only a possible example in the actual application process, and other possible situations will not be exemplified one by one here.

[0037] In a specific implementation scenario, the convolution operation can be first performed on the image feature using the weight parameter and the bias parameter in the modulation parameters, and the convolution feature z mod = conv2d(zimg ,W dyn ) + b. Based on this, the modulation parameter can be used to control the intensity of the above convolution features to obtain the second feature z mod_out = z mod ⊗M. As a possible example, z mod_out ∈R 64*32*32 . Of course, the above example is only a possible example in the actual application process, and other possible situations will not be exemplified one by one here.

[0038] Step S13: Based on the first feature and the second feature, perform fusion decoding to obtain the target string as the target image after the image watermark is embedded in the original image.

[0039] In an implementation scenario, after obtaining the first feature and the second feature, weighted fusion can be performed based on the first feature and the second feature to obtain a weighted feature, and then point-by-point fusion is performed on the decoded image based on the weighted feature and the original image to obtain the target image. It should be noted that the decoded image and the original image can have the same resolution. In addition, the above weighted fusion can be linear fusion, that is, a weighted coefficient can be preset to fuse the first feature and the second feature. For the convenience of description, taking the preset weighted coefficient α as an example, the weighted feature can be expressed as z fuse = αz main +(1 - α)z mod_out . Based on this, the above weighted feature can be upsampled to the image size of the original image through a lightweight convolution decoder to obtain the decoded image, and the pixel point values at the corresponding positions of the decoded image and the original image are fused to obtain the target image. For the convenience of understanding, the target image can be expressed as I marked = D φ (z fuse ) + I, where I represents the original image, and D φ represents the decoder, and D φ (z fuse ) represents the decoded image. In the above manner, weighted fusion is performed based on the first feature and the second feature to obtain a weighted feature, and then point-by-point fusion is performed on the decoded image based on the weighted feature and the original image to obtain the target image, and the decoded image and the original image have the same resolution, which can achieve the residual connection of the original image and help to ensure the visual inapprehensibility of the image watermark in the target image as much as possible.

[0040] In an implementation scenario, as a special example, please refer to Figure 2a , Figure 2a is a schematic diagram of the process of an embodiment of the watermark embedding method of this application. As Figure 2aAs shown, the original image can first be encoded by an encoder to obtain image features, and the target string serving as the image watermark can successively pass through a hash network and a fully connected layer projection to obtain sequence features. On this basis, on the one hand, in one processing flow, the image features can be subjected to deformable convolution to obtain deformed features, and then the deformed features can be concatenated with the sequence features and fed into window attention (such as the relevant descriptions of dividing out fusion sub-features and then introducing positional encoding for self-attention processing, etc.) to obtain the first feature that fuses image information and watermark information. On the other hand, in another processing flow, the image features can be subjected to a first prediction to obtain convolution parameters, and a second prediction can be performed after the fusion of the image features and the sequence features to obtain modulation parameters, so as to perform modulated convolution on the image features according to the convolution parameters and the modulation parameters to obtain the second feature that fuses image information and watermark information. Furthermore, based on the first feature and the second feature, fusion decoding can be performed to obtain a target image that has the same resolution as the original image and uses the target string as the image watermark. In this way, the target string can be embedded in the original image as the image watermark, and the visual quality after watermark embedding can be ensured as much as possible, and the anti-attack ability can be improved. After testing, the above watermark embedding process steps in the embodiments of the present disclosure can significantly improve the robustness and invisibility of watermark embedding while maintaining a visual quality of PSNR > 38dB.

[0041] In one implementation scenario, to improve the implementation efficiency of watermark embedding, the target image can be obtained by a watermark embedding model processing the original image and the target string, and the watermark embedding model can be jointly trained with the watermark extraction model based on sample images in the same training process. In each round of the training process: the sample original image and the sample string can be processed by the watermark embedding model to obtain a sample target image after embedding the sample string as the image watermark into the sample original image, and the sample target image can be attacked according to the selection probability of each of several watermark attack methods to obtain a sample attacked image. The sample attacked image can be used for the watermark extraction model to predict the predicted string as the image watermark, and the selection probability of the watermark attack method can be determined by the bit error between the sample string in the previous round of the training process and the predicted string extracted after adopting the watermark attack method. In the above manner, in the simulated attack process of model training, since the selection probability of the watermark attack method is determined by the bit error between the sample string in the previous round of the training process and the predicted string extracted after adopting the watermark attack method, it helps to force the model to automatically focus on the currently more effective attack method while maintaining the exploration ability, and ultimately can improve the transferability and robustness of adversarial samples.

[0042] In a specific implementation scenario, as a possible example, the watermark embedding model can include Figure 2aMany network modules such as the encoder, multi-layer perceptron, deformable convolution, window attention, modulated convolution, hash network, fully-connected projection layer, and fusion decoding shown in can be specifically referred to the foregoing relevant content about watermark embedding. Here, the network structure of the watermark embedding model is not limited. In addition, the original sample image and the sample string can be processed by the watermark embedding model to obtain a sample string, and the sample target image after the sample string is embedded in the sample original image as an image watermark can be specifically referred to the foregoing relevant description of the target string as an image watermark embedded in the original image, which will not be elaborated here. Of course, although the watermark embedding model and the watermark extraction model are jointly trained, it does not mean that they must also be used together after the training is completed. That is to say, they can be independently applied separately after the training is completed. For example, when only watermark embedding is required, only the watermark embedding model can be used, or when only watermark extraction is required, only the watermark extraction model can be used.

[0043] In a specific implementation scenario, several watermark attack methods may include, but are not limited to: geometric transformation implemented by STN, compression attack simulated by DiffJPEG, masking, etc. Here, the watermark attack methods are not limited. In addition, in the first-round training process, the selection probabilities of each watermark attack method can be the same.

[0044] In a specific implementation scenario, for the watermark extraction model, when obtaining the test image with the string as the image watermark (in the training process, the test image is the sample attack image), in response to detecting the watermark extraction instruction for the test image, it can encode the test image to obtain a token sequence, and the token sequence can contain the feature information of the image watermark. Exemplarily, a coding and decoding structure combining Vision Transformer (i.e., ViT) and an improved U-Net upsampling structure can be used to extract the watermark information. That is to say, in terms of the network structure, the encoder can use the basic ViT architecture to divide the test image (e.g., I aug ∈R 256*256*3 ) into several non-overlapping blocks (e.g., 16*16 non-overlapping blocks). Each block can be converted into an embedding vector (384-dimensional embedding vector) through linear projection, and then a token sequence can be obtained after being encoded by multiple layers of Transformer, such as token∈R 256*384 . On this basis, reshaping can be performed based on the token sequence to obtain the feature to be sampled. Reshaping is mainly used to integrate the token sequence into a three-dimensional vector. For example, for the foregoing token sequence token∈R 256*384In terms of this, 256 dimensions can be reshaped into 16 * 16, that is, the feature to be sampled can be transformed into 16 * 16 * 384. Of course, the above example is just a possible example of feature reshaping. In other cases, it can be deduced by analogy, and no more examples will be given here. After obtaining the feature to be sampled, several transposed convolutions can be sequentially performed based on the feature to be sampled to gradually increase the feature resolution, and the output features of each transposed convolution are obtained. Still taking the above-mentioned feature to be sampled of 16 * 16 * 384 as an example, when 4 transposed convolutions are sequentially performed, the output feature S1 ∈ R of the first transposed convolution can be obtained 32*32*256 , the output feature S2 ∈ R of the second transposed convolution 64*64*128 , the output feature S3 ∈ R of the third transposed convolution 128*128*64 , the output feature S4 ∈ R of the fourth transposed convolution 256*256*32 . On this basis, in the process of the i-th upsampling: the channels of the output feature of the i-th transposed convolution of the feature to be sampled can be adjusted to obtain the current feature, and the transposed convolution and attention processing are sequentially performed based on the output feature of the (i - 1)-th upsampling to obtain the reference feature, and the current feature and the reference feature are fused to obtain the output feature of the i-th upsampling. It should be noted that the channel adjustment can be achieved through 1 * 1 convolution, and the attention processing can be a channel-spatial dual attention mechanism. For the convenience of description, the output feature of the i-th upsampling can be expressed as: F i_layer = Attention(conv_transpose(F i-1_layer )) + skipconnection(S i ), where F i-1_layer represents the output feature of the (i - 1)-th upsampling, conv_transpose represents the transposed convolution, Attention represents the attention processing (such as the channel-spatial dual attention mechanism), and S iDenote the output feature of the i-th execution of the transposed convolution as output_feature_i, and skipconnection is used to adjust the channel dimension (e.g., using 1*1 convolution). In addition, the number of times of upsampling can be the same as the number of times of successively executing several transposed convolutions based on the feature to be sampled. For example, in the above example, if several transposed convolutions are successively executed 4 times based on the feature to be sampled, then upsampling can be correspondingly executed 4 times. Finally, the output feature of the last execution of upsampling can be quantized to obtain the predicted string as the image watermark in the image to be tested. In the above manner, when obtaining the image to be tested with the string as the image watermark, in response to detecting a watermark extraction instruction for the image to be tested, encoding is performed based on the image to be tested to obtain a token sequence, and the token sequence contains the feature information of the image watermark, and then reshaping is performed based on the token sequence to obtain the feature to be sampled, and then several transposed convolutions are successively executed based on the feature to be sampled to gradually increase the feature resolution to obtain the output features of each transposed convolution. Thus, in the process of the i-th execution of upsampling: adjusting the channels of the output feature of the i-th execution of the transposed convolution on the feature to be sampled to obtain the current feature, and successively executing transposed convolution and attention processing based on the output feature of the (i - 1)-th execution of upsampling to obtain the reference feature, and fusing the current feature and the reference feature to obtain the output feature of the i-th execution of upsampling. Furthermore, quantizing the output feature of the last execution of upsampling to obtain the predicted string as the image watermark in the image to be tested, which can achieve watermark extraction through attention processing of residual connection at each upsampling stage, helping to improve the accuracy of watermark extraction.

[0045] In a specific implementation scenario, after predicting the string in the sample attack image obtained by attacking the sample target image in the watermark attack manner, it can be compared with the sample string of the sample original image to obtain the bit error between the two, and then the selection probability of this watermark attack manner can be determined accordingly. For example, the selection probability can be positively correlated with the bit error, that is, the larger the bit error, the larger the selection probability. On the contrary, the smaller the bit error, the smaller the selection probability can also be. As a possible example, in order to improve the accuracy of the selection probability, in the i-th round of training process, each batch can select a watermark attack manner according to the selection probability. Then, for each sample original image in this batch, the corresponding bit error can be determined according to the above process, and the statistical calculation (such as taking the average) of all bit errors in this batch can be used as the statistical bit error corresponding to this watermark formula manner. After obtaining the statistical bit errors corresponding to each watermark attack manner respectively, the selection probabilities of each watermark attack manner in the (i + 1)-th round can be determined accordingly. Specifically, for each watermark attack manner, the bit error between the sample string in the previous round of training process and the predicted string extracted after adopting the watermark attack manner can be obtained, and then normalized based on the bit errors corresponding to each watermark attack manner in the previous round of training process to obtain the selection probabilities of each watermark attack manner in this round of training process. For the sake of description, for the i-th watermark attack manner, its selection probability in the (t + 1)-th round of training process can be expressed as:

[0046] In the above formula, represents the bit error of the i-th watermark attack manner in the t-th round of training process, represents the temperature coefficient, which is used to control the sharpening degree of the probability distribution, represents the selection probability of the i-th watermark attack manner in the (t + 1)-th round of training process. In addition, in order to prevent early convergence or some watermark formula manners from being completely eliminated, when the selection probability of the watermark attack manner is lower than the probability threshold, the selection instruction of the watermark attack manner is reset to the probability threshold. In the above manner, for each watermark attack manner, the bit error between the sample string in the previous round of training process and the predicted string extracted after adopting the watermark attack manner is obtained, and normalized based on the bit errors corresponding to each watermark attack manner in the previous round of training process to obtain the selection probabilities of each watermark attack manner in this round of training process, which can force the model to automatically focus on the currently more effective watermark attack manner while maintaining the exploration ability, and finally improve the transferability and robustness of the adversarial samples.

[0047] In a specific implementation scenario, after obtaining the bit error, the training loss for joint training can be obtained based on the reconstruction loss and the bit error between the sample target image and the sample original image. That is, based on the training loss, the network parameters of the watermark embedding model and the watermark extraction model can be adjusted. It should be noted that the training loss is positively correlated with the reconstruction loss and positively correlated with the bit error. That is to say, during the joint training process, the network parameters can be adjusted with the goal of minimizing the training loss, forcing the watermark embedding model to ensure as much visual consistency as possible after embedding the string as an image watermark into the image, and forcing the watermark extraction model to extract the watermark information as accurately as possible.

[0048] In one implementation scenario, as a specific example, please refer to Figure 2b , Figure 2b which is a schematic diagram of the process of an embodiment of the watermark attack process of this application. As Figure 2b shown, when obtaining the image to be tested with the string as the image watermark, after the image to be tested is attacked by the watermark attack method, the attacked image can be encoded (such as the relevant descriptions of ViT mentioned above), to obtain a token sequence, and then reshaped based on the token sequence to obtain the feature to be sampled. The feature to be sampled is passed through the residual attention with transposed convolution (such as the relevant descriptions of transposed convolution and upsampling mentioned above) to obtain the output feature, and the output feature is quantized to obtain the predicted string as the image watermark. By combining the difference between the actually embedded string in the attacked image to be tested and the predicted string, the bit error can be obtained, and based on this, the effectiveness of the watermark attack method can be measured. For the watermark formula method that induces a higher bit error, its selection probability can be increased proportionally in the next round of training process, while for the watermark attack method with poor performance, its selection probability can be correspondingly reduced. After testing, through the above method, it is possible to force the model to continuously be exposed to the most challenging attack environment while ensuring the training stability as much as possible, and finally achieve excellent robustness with a bit error still lower than 2% under unseen new attacks. Of course, Figure 2b what is shown is only a possible implementation example of the watermark attack process, and it does not limit the implementation of the watermark attack using other processes, and no further examples are given here.

[0049] Based on the above solution, feature encoding is performed on the original image to obtain the image features of the original image, and feature projection is performed on the target string to be embedded to obtain the sequence features of the target string. Then, the image features are subjected to deformable convolution based on the offset obtained by offset prediction on the image features, resulting in the deformed features of the image features. Feature interaction is performed based on the deformed features and the sequence features to obtain the first feature, and modulated convolution is performed based on the image features and the sequence features to obtain the second feature. Furthermore, fusion decoding is performed based on the first feature and the second feature to obtain the target image with the target string embedded as the image watermark in the original image. Therefore, after obtaining the image features and the sequence features, on the one hand, in one path of the feature processing flow, offset prediction is performed according to the image features to obtain the offset, and the image features are subjected to deformable convolution based on this, which can adaptively capture the local feature structure of the original image. Subsequently, when performing feature interaction, it helps to make the watermark embedding avoid the visually perceptible areas as much as possible, that is, it can ensure the visual quality after watermark embedding as much as possible. On the other hand, in the other path of the feature processing flow, modulated convolution is performed through the image features and the sequence features, and fusion decoding is performed with the aforementioned one path of the feature processing flow, which can achieve dual-path collaborative watermark embedding and help improve the anti-attack ability. Therefore, it can ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability.

[0050] Please refer to Figure 3 , Figure 3 which is a schematic framework diagram of an embodiment of the watermark embedding device of the present application. The watermark embedding device 30 includes: a feature preparation module 31, a split-path processing module 32, and a fusion decoding module 33. The feature preparation module 31 is configured to perform feature encoding on the original image to obtain the image features of the original image, and perform feature projection on the target string to be embedded to obtain the sequence features of the target string. The split-path processing module 32 is configured to perform deformable convolution on the image features based on the offset obtained by offset prediction on the image features to obtain the deformed features of the image features, perform feature interaction based on the deformed features and the sequence features to obtain the first feature, and perform modulated convolution based on the image features and the sequence features to obtain the second feature. The fusion decoding module 33 is configured to perform fusion decoding based on the first feature and the second feature to obtain the target image with the target string embedded as the image watermark in the original image.

[0051] In the above solution, the watermark embedding device 30 performs feature encoding based on the original image to obtain the image features of the original image, and performs feature projection based on the target string to be embedded to obtain the sequence features of the target string. Then, based on the offset obtained by offset prediction from the image features, the image features are subjected to deformable convolution to obtain the deformed features of the image features. Based on the deformed features and the sequence features, feature interaction is performed to obtain the first feature, and based on the image features and the sequence features, modulation convolution is performed to obtain the second feature. Furthermore, based on the first feature and the second feature, fusion decoding is performed to obtain the target string as the target image after the image watermark is embedded in the original image. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed according to the image features to obtain the offset, and the image features are subjected to deformable convolution based on this, which can adaptively capture the local feature structure of the original image. Subsequently, when performing feature interaction, it helps to make the watermark embedding avoid the visually sensitive areas as much as possible, that is, it can ensure the visual quality after watermark embedding as much as possible. On the other hand, in another feature processing flow, modulation convolution is performed through the image features and the sequence features, and fusion decoding is performed with the aforementioned one feature processing flow, which can realize dual-path collaborative watermark embedding and help improve the anti-attack ability. Therefore, it can ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability.

[0052] In some publicly disclosed embodiments, the split processing module 32 includes an offset prediction sub-module for performing offset prediction based on the image features to obtain an offset; wherein, the offset includes: the offset values of each feature position in the image features at each convolution position within the convolution range; the split processing module 32 includes a deformable convolution sub-module for, for each feature position during the execution of the deformable convolution, based on the weight factors of each convolution position within the convolution range of the feature position, respectively performing weighted summation on the image sub-features at the target positions after the convolution positions in the image features are offset by the offset values to obtain the deformed sub-features of the feature position; the split processing module 32 includes a first combination sub-module for obtaining the deformed features based on the deformed sub-features of each feature position in the image features.

[0053] In some disclosed embodiments, the image features and the sequence features have the same resolution. The split processing module 32 includes a feature splicing sub-module for splicing the deformation features and the sequence features in the channel dimension to obtain fused features; the split processing module 32 includes a feature division sub-module for dividing the fused features in the resolution dimension to obtain a plurality of fused sub-features; the split processing module 32 includes a feature enhancement sub-module for, for each fused sub-feature, processing the fused sub-feature based on the self-attention mechanism introducing positional encoding to obtain an enhanced sub-feature of the fused sub-feature; the split processing module 32 includes a second combination sub-module for combining the enhanced sub-features of the respective fused sub-features to obtain enhanced features; the split processing module 32 includes a feature fusion sub-module for fusing the enhanced features and the sequence features to obtain first features.

[0054] In some disclosed embodiments, the positional encoding is generated from learnable parameters; and / or, during the processing of the self-attention mechanism, after the interaction sub-feature between the query feature and the key feature of the fused sub-feature is fused with the positional encoding, it is then interacted with the value feature of the fused sub-feature to obtain the enhanced sub-feature of the fused sub-feature.

[0055] In some disclosed embodiments, the feature fusion sub-module includes a weight prediction unit for predicting weights based on the enhanced features and the sequence features to obtain a weight image for fusing the enhanced features and the sequence features; wherein, the weight image has the same resolution as the image features and the sequence features; the feature fusion sub-module includes a weighted processing unit for performing weighted processing on the enhanced features and the sequence features based on the weight image to obtain first features.

[0056] In some disclosed embodiments, the split processing module 32 includes a parameter prediction sub-module for performing a first prediction based on the image features to obtain the convolution parameters of the modulated convolution, and performing a second prediction based on the image features and the sequence features to obtain the modulation parameters of the modulated convolution; wherein, the convolution parameters include weight parameters and bias parameters, and the modulation parameters include the intensity values of each feature position when performing the modulated convolution; the split processing module 32 includes a modulated convolution sub-module for performing modulated convolution on the image features based on the convolution parameters and the modulation parameters to obtain second features.

[0057] In some disclosed embodiments, the fusion decoding module 33 includes a weighted fusion sub-module for performing weighted fusion on the first features and the second features to obtain weighted features; the fusion decoding module 33 includes a pointwise fusion sub-module for performing pointwise fusion on the decoded image of the weighted features and the original image to obtain the target image; wherein, the decoded image has the same resolution as the original image.

[0058] In some disclosed embodiments, the target image is obtained by processing an original image and a target string with a watermark embedding model. The watermark embedding model and the watermark extraction model are jointly trained based on sample images in the same training process. In each round of the training process: the sample original image and the sample string are processed by the watermark embedding model to obtain a sample string, which is used as an image watermark to be embedded into the sample original image to obtain a sample target image. Then, the sample target image is attacked according to the selection probability of each of several watermark attack methods to obtain a sample attacked image. The sample attacked image is used for the watermark extraction model to predict a predicted string as the image watermark. And the selection probability of the watermark attack method is determined by the bit error between the sample string and the predicted string extracted after adopting the watermark attack method in the previous round of the training process.

[0059] In some disclosed embodiments, the watermark embedding device 30 includes an error metric module, which is configured to obtain, for each watermark attack method, the bit error between the sample string and the predicted string extracted after adopting the watermark attack method in the previous round of the training process. The watermark embedding device 30 includes a normalization module, which is configured to normalize based on the respective bit errors corresponding to various watermark attack methods in the previous round of the training process to obtain the selection probabilities of various watermark attack methods in the current round of the training process.

[0060] In some disclosed embodiments, the watermark embedding model and the watermark extraction model are independently applied separately after the training is completed; and / or, in the first round of the training process, the selection probabilities of each of the watermark attack methods are the same; and / or, the selection probability is positively correlated with the bit error; and / or, in the case where the selection probability of the watermark attack method is lower than the probability threshold, the selection instruction of the watermark attack method is reset to the probability threshold.

[0061] In some disclosed embodiments, the watermark embedding device 30 includes an image encoding module, which is configured to, when obtaining a to-be-tested image with a string as an image watermark and in response to detecting a watermark extraction instruction for the to-be-tested image, perform encoding based on the to-be-tested image to obtain a token sequence; wherein the token sequence contains feature information of the image watermark; the watermark embedding device 30 includes a feature reshaping module, which is configured to reshape based on the token sequence to obtain a to-be-sampled feature; the watermark embedding device 30 includes a transposed convolution module, which is configured to sequentially perform a plurality of transposed convolutions based on the to-be-sampled feature to gradually increase the feature resolution and obtain output features of each transposed convolution; the watermark embedding device 30 includes an upsampling module, which is configured to, during the i-th execution of upsampling: adjust the channels of the output feature of the i-th execution of transposed convolution on the to-be-sampled feature to obtain a current feature, perform transposed convolution and attention processing based on the output feature of the (i - 1)-th execution of upsampling in sequence to obtain a reference feature, and fuse the current feature and the reference feature to obtain the output feature of the i-th execution of upsampling; the watermark embedding device 30 includes a feature quantization module, which is configured to quantize the output feature of the last execution of upsampling to obtain a predicted string serving as the image watermark in the to-be-tested image.

[0062] Please refer to Figure 4 , Figure 4 is a schematic framework diagram of an embodiment of an electronic device according to the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. At least program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above-described embodiments of the watermark embedding method. Specifically, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 40 may include, but is not limited to, devices such as mobile phones, tablet computers, learning machines, smart large screens, servers, etc. The specific type of the electronic device 40 is not limited herein.

[0063] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned watermark embedding method embodiments. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Additionally, the processor 42 can be implemented jointly by integrated circuit chips.

[0064] In the above solution, the electronic device 40 performs feature encoding on the original image to obtain the image features of the original image, and performs feature projection on the target string to be embedded to obtain the sequence features of the target string. Then, based on the offset obtained by offset prediction from the image features, the image features are subjected to deformable convolution to obtain the deformed features of the image features. Based on the deformed features and the sequence features, feature interaction is performed to obtain the first feature, and based on the image features and the sequence features, modulation convolution is performed to obtain the second feature. Furthermore, based on the first feature and the second feature, fusion decoding is performed to obtain the target string as the target image after the image watermark is embedded in the original image. Therefore, after obtaining the image features and the sequence features, on the one hand, in one path of the feature processing flow, offset prediction is performed based on the image features to obtain the offset, and the image features are subjected to deformable convolution based on this, which can adaptively capture the local feature structure of the original image. Subsequently, when performing feature interaction, it helps to make the watermark embedding avoid the visually perceptible areas as much as possible, that is, it can ensure the visual quality after watermark embedding as much as possible. On the other hand, in another path of the feature processing flow, modulation convolution is performed through the image features and the sequence features, and fusion decoding is performed with the aforementioned one path of the feature processing flow, which can realize dual-path collaborative watermark embedding and help improve the anti-attack ability. Therefore, it can ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability.

[0065] Please refer to Figure 5 , Figure 5 is a framework schematic diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor. The program instructions 51 are used to implement the steps in any of the above-mentioned watermark embedding method embodiments.

[0066] In the above solution, the computer-readable storage medium 50 performs feature encoding based on the original image to obtain the image features of the original image, and performs feature projection based on the target string to be embedded to obtain the sequence features of the target string. Then, based on the offset obtained by offset prediction from the image features, deformable convolution is performed on the image features to obtain the deformed features of the image features. Based on the deformed features and the sequence features, feature interaction is performed to obtain the first feature, and based on the image features and the sequence features, modulated convolution is performed to obtain the second feature. Furthermore, based on the first feature and the second feature, fusion decoding is performed to obtain the target string as the target image after the image watermark is embedded in the original image. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed according to the image features to obtain the offset, and deformable convolution is performed on the image features based on this, which can adaptively capture the local feature structure of the original image. Subsequently, when performing feature interaction, it helps to make the watermark embedding avoid the visually perceptible areas as much as possible, that is, it can ensure the visual quality after watermark embedding as much as possible. On the other hand, in another feature processing flow, modulated convolution is performed through the image features and the sequence features, and fusion decoding is performed with the aforementioned one feature processing flow, which can achieve dual-branch collaborative watermark embedding and help improve the anti-attack ability. Therefore, it can ensure the visual quality after watermark embedding as much as possible and improve the anti-attack ability.

[0067] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0068] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0069] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0070] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0071] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0072] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0073] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A watermark embedding method, characterized in that, Including: Performing feature encoding on the original image to obtain the image features of the original image, and performing feature projection on the target string to be embedded to obtain the sequence features of the target string; Performing deformable convolution on the image features based on the offset obtained by offset prediction of the image features to obtain the deformed features of the image features, performing feature interaction based on the deformed features and the sequence features to obtain the first feature, and performing modulated convolution based on the image features and the sequence features to obtain the second feature; Performing fusion decoding based on the first feature and the second feature to obtain the target image after embedding the target string as an image watermark into the original image.

2. The method according to claim 1, wherein The performing deformable convolution on the image features based on the offset obtained by offset prediction of the image features to obtain the deformed features of the image features includes: Performing offset prediction on the image features to obtain the offset; wherein, the offset includes: the offset values of each feature position in the image features at each convolution position within the convolution range; During the process of performing the deformable convolution, for each of the feature positions, based on the weight factors of each convolution position within the convolution range of the feature position, respectively performing weighted summation on the image sub-features at the target positions after offsetting the convolution positions in the image features by the offset values to obtain the deformed sub-features of the feature position; Based on the deformed sub-features of each of the feature positions in the image features, obtaining the deformed features.

3. The method according to claim 1, wherein The image features and the sequence features have the same resolution. The performing feature interaction based on the deformed features and the sequence features to obtain the first feature includes: Performing splicing on the deformed features and the sequence features in the channel dimension to obtain a fused feature; Performing division on the fused feature in the resolution dimension to obtain a plurality of fused sub-features; For each of the fused sub-features, processing the fused sub-feature based on the self-attention mechanism introducing positional encoding to obtain the enhanced sub-feature of the fused sub-feature; Combining the enhanced sub-features of each of the fused sub-features to obtain an enhanced feature; Performing fusion based on the enhanced feature and the sequence features to obtain the first feature.

4. The method according to claim 3, characterized in that, The positional encoding is generated from learnable parameters; And / or, during the processing of the self-attention mechanism, after the interaction sub-feature between the query feature and the key feature of the fused sub-feature is fused with the positional encoding, then interacting with the value feature of the fused sub-feature to obtain the enhanced sub-feature of the fused sub-feature.

5. The method according to claim 3, characterized in that, The performing fusion based on the enhanced feature and the sequence features to obtain the first feature includes: Performing weight prediction based on the enhanced feature and the sequence features to obtain a weight image for fusing the enhanced feature and the sequence features; wherein, the weight image has the same resolution as the image features and the sequence features; Performing weighted processing on the enhanced feature and the sequence features based on the weight image to obtain the first feature.

6. The method according to claim 1, characterized in that Performing modulation convolution based on the image features and the sequence features to obtain a second feature, including: Performing a first prediction based on the image features to obtain the convolution parameters of the modulation convolution, and performing a second prediction based on the image features and the sequence features to obtain the modulation parameters of the modulation convolution; wherein, the convolution parameters include weight parameters and bias parameters, and the modulation parameters include intensity values at each feature position when performing the modulation convolution; Performing modulation convolution on the image features based on the convolution parameters and the modulation parameters to obtain the second feature.

7. The method according to claim 1, wherein Performing fusion decoding based on the first feature and the second feature to obtain the target string as the target image after embedding the image watermark into the original image, including: Performing weighted fusion based on the first feature and the second feature to obtain a weighted feature; Performing point-by-point fusion on the decoded image of the weighted feature and the original image to obtain the target image; wherein, the decoded image has the same resolution as the original image.

8. The method according to claim 1, wherein The target image is obtained by processing the original image and the target string by a watermark embedding model. The watermark embedding model and the watermark extraction model are jointly trained based on sample images in the same training process. In each round of the training process: the sample original image and the sample string are processed by the watermark embedding model to obtain the sample string as the image watermark embedded into the sample original image to obtain the sample target image, and the sample target image is attacked according to the selection probability of each of several watermark attack methods to obtain a sample attacked image. The sample attacked image is used for the watermark extraction model to predict the predicted string as the image watermark, and the selection probability of the watermark attack method is determined by the bit error between the sample string and the predicted string extracted after adopting the watermark attack method in the previous round of the training process.

9. The method according to claim 8, wherein The selection probability of the watermark attack method is updated through the following steps: For each of the watermark attack methods, obtaining the bit error between the sample string and the predicted string extracted after adopting the watermark attack method in the previous round of the training process; Normalizing based on the respective bit errors corresponding to the watermark attack methods in the previous round of the training process to obtain the selection probabilities of the watermark attack methods in this round of the training process.

10. The method according to claim 8, characterized in that The watermark embedding model and the watermark extraction model are independently applied separately after the training is completed; and / or, in the first round of the training process, the selection probabilities of the watermark attack methods are the same; and / or, the selection probability is positively correlated with the bit error; and / or, in the case where the selection probability of the watermark attack method is lower than the probability threshold, the selection instruction of the watermark attack method is reset to the probability threshold.

11. The method according to any one of claims 1 to 10, characterized in that, When obtaining a test image with a string as the image watermark, the method further includes: In response to detecting a watermark extraction instruction for the test image, encoding the test image to obtain a token sequence; wherein, the token sequence contains the feature information of the image watermark; Performing reshaping based on the token sequence to obtain a feature to be sampled; Perform a number of transposed convolutions in sequence based on the feature to be sampled to gradually increase the feature resolution, and obtain the output features of each transposed convolution; During the i-th upsampling process: adjust the channels of the output feature of the i-th transposed convolution of the feature to be sampled to obtain the current feature, and perform transposed convolution and attention processing on the output feature of the (i-1)-th upsampling in sequence to obtain the reference feature, and fuse the current feature and the reference feature to obtain the output feature of the i-th upsampling; Quantize the output feature of the last upsampling to obtain the predicted string as the image watermark in the image to be tested.

12. A watermark embedding device, characterized in that, Comprising: A feature preparation module for performing feature encoding on the original image to obtain the image feature of the original image, and performing feature projection on the target string to be embedded to obtain the sequence feature of the target string; A shunt processing module for performing deformable convolution on the image feature based on the offset obtained by offset prediction of the image feature to obtain the deformed feature of the image feature, performing feature interaction between the deformed feature and the sequence feature to obtain the first feature, and performing modulated convolution on the image feature and the sequence feature to obtain the second feature; A fusion decoding module for performing fusion decoding on the first feature and the second feature to obtain the target image after embedding the target string as the image watermark into the original image.

13. An electronic device, characterized in that, At least comprising a memory and a processor, at least program instructions are stored in the memory, and the processor is used to execute the program instructions to implement the watermark embedding method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, Store program instructions that can be run by the processor, and the program instructions are used to implement the watermark embedding method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Robust image watermarking method and system based on hierarchical attention feature fusion

    CN115908095A

  • Deep watermarking method for immune geometric distortion

    CN119273523A

  • Image watermark processing method and device, storage medium and electronic equipment

    CN119963390A

  • Insertion and detection of hidden watermark in digital image or audio data uses decoded component coefficients that are modulated by signal representing watermarking information to form watermark coefficients

    FR2785426A1

  • Cross-modal image-watermark joint generation and detection device and method thereof

    US12125119B1