Watermark embedding method and related device, equipment and storage medium

Through dual-split collaborative processing of feature coding, offset prediction deformation convolution and modulated convolution, the existing watermark embedding technology has solved the problems of weak attack resistance and insufficient visual quality, and achieved high visual quality and robust watermark embedding.

CN120278870BActive Publication Date: 2025-08-22IFLYTEK CO LTD

Patent Information

Application Number
CN202510758988.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-22
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing watermark embedding technology has weak attack resistance and difficult to guarantee visual quality, making it difficult to meet application needs.

Method used

By combining offset prediction deformation convolution and feature interaction based on original image feature encoding and target string feature projection, image watermark embedding is used to realize dual-split collaborative processing of image features and sequence features, improving attack resistance and ensuring visual quality.

Benefits of technology

While maintaining high visual quality, it significantly improves the robustness and concealment of watermark embedding, and can effectively resist various attacks under PSNR>38dB.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278870B_ABST
    Figure CN120278870B_ABST
Patent Text Reader

Abstract

The present application discloses a watermark embedding method and related devices, equipment, and storage medium, wherein the watermark embedding method includes: performing feature encoding based on the original image to obtain image features of the original image, and performing feature projection based on the target string to be embedded to obtain sequence features of the target string; performing deformation convolution on the image features based on the offset obtained by offset prediction of the image features to obtain deformation features of the image features, and performing feature interaction based on the deformation features and the sequence features to obtain a first feature, and performing modulated convolution based on the image features and the sequence features to obtain a second feature; performing fusion decoding based on the first feature and the second feature to obtain a target image after the target string is embedded as an image watermark in the original image. The above scheme can ensure the visual quality after the watermark is embedded as much as possible and improve the anti-attack capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a watermark embedding method and related devices, equipment and storage media. Background Art

[0002] As an important means of digital copyright protection, digital image watermarking technology has broad application prospects in multimedia content authentication, intellectual property protection and other fields.

[0003] Currently, existing watermark embedding technologies have technical issues such as weak anti-attack capabilities and the easy generation of visual artifacts, which makes it difficult to ensure visual quality. As a result, it is difficult to meet application requirements. In view of this, how to ensure the visual quality of the embedded watermark as much as possible and improve the anti-attack capabilities has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a watermark embedding method and related devices, equipment and storage media, which can ensure the visual quality of the watermark after embedding as much as possible and improve the anti-attack capability.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a watermark embedding method, including: performing feature encoding based on the original image to obtain image features of the original image, and performing feature projection based on the target character string to be embedded to obtain sequence features of the target character string; performing deformation convolution on the image features based on the offset obtained by offset prediction of the image features to obtain deformation features of the image features, and performing feature interaction based on the deformation features and the sequence features to obtain a first feature, and performing modulated convolution based on the image features and the sequence features to obtain a second feature; performing fusion decoding based on the first feature and the second feature to obtain a target image after the target character string is embedded in the original image as an image watermark.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a watermark embedding device, including: a feature preparation module, a branch processing module and a fusion decoding module, the feature preparation module is used to perform feature encoding based on the original image to obtain the image features of the original image, and perform feature projection based on the target character string to be embedded to obtain the sequence features of the target character string; the branch processing module is used to perform deformation convolution on the image features based on the offset obtained by offset prediction of the image features to obtain the deformation features of the image features, and perform feature interaction based on the deformation features and the sequence features to obtain the first feature, and perform modulated convolution based on the image features and the sequence features to obtain the second feature; the fusion decoding module is used to perform fusion decoding based on the first feature and the second feature to obtain the target image after the target character string is embedded in the original image as an image watermark.

[0007] In order to solve the above technical problems, the third aspect of this application provides an electronic device, which at least includes a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the watermark embedding method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the watermark embedding method of the first aspect.

[0009] The above scheme performs feature encoding based on the original image to obtain the image features of the original image, and performs feature projection based on the target string to be embedded to obtain the sequence features of the target string, thereby deforming and convolving the image features based on the offset obtained by offset prediction of the image features to obtain the deformed features of the image features, and performing feature interaction based on the deformed features and the sequence features to obtain the first feature, and performing modulated convolution based on the image features and the sequence features to obtain the second feature, and then performing fusion decoding based on the first feature and the second feature to obtain the target image after the target string is embedded in the original image as an image watermark. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed according to the image features to obtain the offset and deformed convolution is performed on the image features based on this, which can adaptively capture the local feature structure of the original image, and then when performing feature interaction, it helps to make the watermark embedding avoid the visually perceptible area as much as possible, so as to ensure the visual quality after the watermark is embedded as much as possible. On the other hand, in another feature processing flow, modulated convolution is performed by the image features and the sequence features, and fused decoding is performed with the above-mentioned feature processing flow, which can realize dual-path collaborative watermark embedding, which helps to improve the anti-attack capability. Therefore, the visual quality of the watermark after embedding can be guaranteed as much as possible and the anti-attack capability can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flow chart of an embodiment of the watermark embedding method of the present application;

[0011] Figure 2a This is a process diagram of an embodiment of the watermark embedding method of the present application;

[0012] Figure 2b This is a process diagram of an embodiment of the watermark attack process of the present application;

[0013] Figure 3 This is a schematic diagram of the framework of an embodiment of the watermark embedding device of the present application;

[0014] Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0015] Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0016] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0017] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0018] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the fragment " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document refers to two or more than two.

[0019] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the watermark embedding method of the present application. Specifically, it may include the following steps:

[0020] Step S11: performing feature encoding based on the original image to obtain image features of the original image, and performing feature projection based on the target character string to be embedded to obtain sequence features of the target character string.

[0021] In one implementation scenario, the original image may be image data to be embedded with an image watermark, which may specifically include but is not limited to: video frame images of film and television dramas, scanned images of paper paintings, electronic painting images, etc. The specific content of the original image is not limited here.

[0022] In one implementation scenario, an encoder can be used to perform feature encoding on the original image to compress the original image into a latent space to obtain the image features of the original image. It should be noted that the encoder can include but is not limited to a convolutional neural network, and the network structure of the encoder is not limited here. For ease of description, taking the original image as an RGB image with a resolution of 256*256 as an example, the original image can be recorded as I∈[0,1] 256*256*3, where the superscript 3 represents the three channels of R, G, and B, and [0,1] represents the pixel value between 0 and 1, that is, the original image can be represented as a matrix with 3 channels and a size of 256*256. Of course, the above example is only one possible representation of the original image, and does not limit the resolution and number of channels of the original image. In addition, in addition to encoding the features of the original image through the encoder, the original image can also be feature encoded through manually designed feature operators such as directional gradient histogram, local binary pattern, scale-invariant feature transform, accelerated robust feature, etc., or the original image can be feature encoded through encoding methods based on statistical learning such as principal component analysis and independent component analysis. The specific method of feature encoding the original image is not limited and will not be given one by one. For the sake of convenience, the above-mentioned original image I is still taken as an example, and the image features obtained by feature encoding can be recorded as z img ∈R 4*32*32 .

[0023] In one implementation scenario, the target string to be embedded can be a binary string, such as s∈{0,1} 64 , that is, any character in the target string is either binary 0 or binary 1, and the length of the target string is 64. Of course, the above example is only one possible example of the target string. For example, the target string to be embedded can also be a decimal string, and even the string to be embedded can include letters. Other possible situations are not limited here, and no further examples are given.

[0024] In one implementation scenario, in order to perform feature projection on the target string to be embedded, the target string can be processed using a hash network to obtain a conditional vector. For ease of description, still taking the aforementioned target string s as an example, the conditional vector can be expressed as c=H ψ (s). For example, the conditional vector c can be a 128-dimensional vector. Then, the conditional vector can be projected to a projection feature whose total number of single-channel elements is an integer multiple of the aforementioned image feature through a fully connected layer. img For example (the total number of single channel elements is 32*32=1024), the above 128-dimensional conditional vector is projected by the fully connected layer, and a 2048-dimensional projection feature can be obtained. On this basis, the projection feature can be reshaped to obtain a sequence feature with the same resolution as the image feature. img For example, the aforementioned 2048-dimensional projection feature can be reshaped to obtain a 64*32*32-dimensional sequence feature wm featIn this way, the global distribution of the watermark information can be achieved by repeatedly expanding in space, so that the complete watermark information can be perceived at every position in the space. Of course, the above example is only one possible implementation method of feature projection in actual application. Other possible implementation methods are not limited here (for example, a neural network model containing network layers such as convolutional layers can be directly used to project features on the target sequence, etc.), and no further examples are given.

[0025] Step S12: performing deformation convolution on the image feature based on the offset obtained by offset prediction of the image feature to obtain a deformation feature of the image feature, performing feature interaction based on the deformation feature and the sequence feature to obtain a first feature, and performing modulated convolution based on the image feature and the sequence feature to obtain a second feature.

[0026] In one implementation scenario, in order to obtain the deformation feature of the image feature, an offset prediction can be first performed based on the image feature to obtain an offset, and the offset can specifically include the offset value of each convolution position of each feature position in the image feature within the convolution range. On this basis, in the process of performing the deformation convolution, for each feature position, the weight factors of each convolution position within the convolution range of the feature position can be used to perform weighted summation on the image sub-features at the target position after the convolution position in the image feature is offset by the offset value, respectively, to obtain the deformation sub-feature of the feature position, and then the deformation feature can be obtained based on the deformation sub-features of each feature position in the image feature. The above method, by successively performing feature processing such as offset prediction and deformation convolution based on the feature position as the basic unit, can refine the granularity of feature processing as much as possible in the process of adaptively capturing the local feature structure of the original image, which helps to improve the degree of fineness of capturing the local feature structure.

[0027] In a specific implementation scenario, the offset prediction of the image features can be performed based on the neural network model to obtain the offset value of each convolution position of each feature position in the image feature within the convolution range. It should be noted that the neural network model may include but is not limited to network layers such as standard convolution layers. The network structure of the neural network model used for offset prediction is not limited here, and examples are not given one by one. In addition, the offset value at the convolution position is used to characterize the offset position where the local feature structure exists relative to the convolution position, so as to assist in the subsequent capture of the local feature structure of the original image. As a special example, in order to prevent the semantic content from being destroyed due to excessive deformation, the offset amount can also be constrained in the offset prediction process, such as a regularization constraint of ||Δp||2<1.5.

[0028] In a specific implementation scenario, taking a 3*3 convolution kernel as an example, the convolution range is 3*3. For example, for any feature position, the 3*3 range centered on the feature position is the convolution range (in other words, each feature position has 9 different convolution position offset values). For any convolution position within the 3*3 convolution range, a corresponding offset value is predicted (the offset value of the image feature in the width dimension and height dimension outside the channel dimension). For the sake of convenience, we still use the aforementioned image feature z img ∈R 4*32*32 For example, when using a 3*3 convolution kernel, the offset can be expressed as Δp∈R 2*9*32*32 Of course, the above example is only one possible example of the convolution range and offset in actual application. Other possible situations are not limited here and will not be given examples one by one.

[0029] In a specific implementation scenario, still taking the 3*3 convolution kernel as an example, after obtaining the offset, for any feature position p in the image feature, the weight factor of each convolution position can be used to weight the sum of the image sub-features at the target position after the corresponding convolution position is offset by the offset value within the 3*3 convolution range centered on the feature position p, and the deformed sub-feature F of the feature position p can be obtained. out (p):

[0030]

[0031] In the above formula, W(k) represents the weight factor at the convolution position k within the 3*3 convolution range centered on the feature position p, and F in represents the image features, p k represents the preset offset value at the convolution position k within the 3*3 convolution range centered on the feature position p (e.g., [0,0], [0,1], [1,0], [0,-1], [-1,0], [1,1], [1,-1], [-1,1], [-1,-1]), Δp(k,p) represents the predicted offset value at the convolution position k within the 3*3 convolution range centered on the feature position p, then p+p k +Δp(k,p) represents the target position, F in (p+p k +Δp(k,p)) represents the image feature F in Image sub-features at target locations.

[0032] In a specific implementation scenario, after obtaining the deformed sub-features of each feature position in the image feature, the deformed sub-features of each feature position can be combined according to the original orientation of each feature position in the image feature to obtain the deformed feature of the image feature. It should be noted that the above-mentioned deformation convolution processing flow can be repeated multiple times. For example, after obtaining the deformed feature, it can be used as a new image feature, and the above-mentioned offset prediction based on the image feature can be returned to the new image feature to obtain the offset step, and this cycle is iterated until the deformation convolution is executed the target number of times (such as 3 times), and the latest deformation feature can be used as the final deformation feature of the image feature. In addition, in order to facilitate the implementation of deformation convolution, the above-mentioned process steps of performing deformation convolution can be implemented by a feature extraction network composed of several groups (such as 3 groups) of deformable convolutions. For details, please refer to the technical details of deformable convolution, which will not be repeated here.

[0033] In an implementation scenario, after obtaining the deformation feature of the image feature, it can be subjected to feature interaction with the sequence feature to obtain the first feature. It should be noted that the first feature obtained by the above-mentioned deformation convolution, feature interaction and other process steps in the embodiment of the present disclosure fuses the feature information of both the image and the watermark. As a possible implementation method, attention mechanisms such as the cross-attention mechanism can be directly used to perform feature interaction on the deformation feature and the sequence feature to establish a contextual relationship between the watermark information and the global image, and to enhance the degree of correlation in areas with complex textures. It should be noted that the specific process of feature interaction in this case can refer to the technical details of attention mechanisms such as the cross-attention mechanism, which will not be repeated here.

[0034] In one implementation scenario, as another possible implementation method, different from the aforementioned direct use of the attention mechanism to perform feature interaction on image features and sequence features, it is also possible to first splice the deformation features and sequence features in the channel dimension to obtain a fusion feature, and then divide the fusion features in the resolution dimension to obtain several fusion sub-features. On this basis, for each fusion sub-feature, the fusion sub-feature can be processed based on the self-attention mechanism that introduces position coding to obtain an enhanced sub-feature of the fusion sub-feature, and the enhanced sub-features of each fusion sub-feature are combined to obtain an enhanced feature, and then the enhanced feature can be fused with the sequence feature to obtain the first feature. The above method, by introducing position coding on each fusion sub-feature for self-attention mechanism processing through window attention, can establish the relationship between the watermark information and the global context of the image, and strengthen the correlation strength of the complex texture area.

[0035] In a specific implementation scenario, for the convenience of description, the deformation feature is denoted as F def ∈R 64*32*32 For example, for the 64*32*32 dimensional sequence feature wmfeat For example, the two can be concatenated in the channel dimension to obtain a 128*32*32 dimensional fused feature. The fused feature can then be divided in the resolution dimension. For example, if it is divided into 8*8 small windows, 16 128*8*8 dimensional fused sub-features can be obtained. Of course, the above example is only one possible example of how to divide the fused sub-features in actual applications, and other possible scenarios will not be given here one by one.

[0036] In a specific implementation scenario, the attention mechanism used to fuse sub-features may include but is not limited to a multi-head self-attention mechanism (e.g., the number of heads n head =4), which is not limited here. In addition, for each fused sub-feature, based on the self-attention mechanism introduced by position encoding, the enhanced sub-feature can be obtained:

[0037]

[0038] In the above formula, Q, K, and V represent the query feature after the fused sub-feature is transformed by the query matrix, the key feature after the key matrix is ​​transformed, and the value feature after the value matrix is ​​transformed. The superscript T of K represents the transpose, d represents the feature dimension, and B represents the positional encoding. In other words, during the self-attention mechanism, the interaction sub-features between the query feature and the key feature of the fused sub-feature are fused with the positional encoding, and then interact with the value feature of the fused sub-feature to obtain the enhanced sub-feature of the fused sub-feature. In addition, as a possible example, the positional encoding can be generated by learnable parameters, which is not limited here.

[0039] In a specific implementation scenario, after obtaining the enhancer features of each fusion sub-feature, the enhancer features of each fusion sub-feature can be combined according to the original position of each fusion sub-feature to obtain the enhancement feature. For ease of description, the enhancement feature can be denoted as F main For example, still taking the aforementioned sequence features and image features as an example, we can finally get F main ∈R 128*32*32 Of course, the above example is only one possible example of the enhancement feature, and other possible scenarios of the enhancement feature are not limited here, and no further examples are given.

[0040] In a specific implementation scenario, after obtaining the enhanced features, the enhanced features and the sequence features can be fused to obtain the first feature. Specifically, weight prediction can be performed based on the enhanced features and the sequence features to obtain a weight image for fusing the enhanced features and the sequence features, and the weight image has the same resolution as the image features and the sequence features, and then the enhanced features and the sequence features are weighted based on the weight image to obtain the first feature. For example, the enhanced features and the sequence features can be spliced ​​in the channel dimension first. Since the number of channels increases at this time, one-dimensional convolution can be used to reduce the number of channels, and activation functions such as sigmoid can be used to constrain the output data obtained after the one-dimensional convolution to within the range of 0~1 as a weight image. For the convenience of description, the weight image G can be expressed as σ(conv 1*1 ([F main ,wm feat ])). On this basis, the weight image can be used to perform weighted summation of the enhanced features and sequence features to obtain the first feature z main :

[0041]

[0042] In the above formula, z main Represents the first feature, F main Represents enhanced features, wm feat Represents sequence features, represents the dot product operation, G represents the weight image, and 1-G represents another weight image obtained by subtracting each element in the weight image from 1, that is, the sum of this weight image and the corresponding element in the previous weight image G is 1. For example, the first feature z main ∈R 64 *32*32 . It should be noted that the above example is only a possible example of the first feature in the actual application process, and other possible situations will not be given one by one here. In the above method, weight prediction is performed based on the enhancement feature and the sequence feature to obtain a weight image for fusing the enhancement feature and the sequence feature, and the weight image has the same resolution as the image feature and the sequence feature. The enhancement feature and the sequence feature are then weighted based on the weight image to obtain the first feature, so it can adaptively focus on fusing the enhancement feature and the sequence feature.

[0043] In an implementation scenario, similar to the aforementioned first feature, the second feature in the embodiment of the present disclosure also integrates feature information of both the image and the watermark. The main difference between the second feature and the first feature is that the two are implemented in different ways. Specifically, in order to implement the modulation convolution of the image features and the sequence features, a first prediction can be made based on the image features to obtain the convolution parameters of the modulation convolution, and a second prediction can be made based on the image features and the sequence features to obtain the modulation parameters of the modulation convolution, and the convolution parameters can include weight parameters and bias parameters, and the modulation parameters can include the intensity values ​​of each feature position when performing the modulation convolution. On this basis, the image features can be modulated and convolved based on the convolution parameters and the modulation parameters to obtain the second feature. It should be noted that the specific process of the modulation convolution can refer to the technical details of the modulation convolution, which will not be repeated here. The above method, through the first prediction and the second prediction, to obtain the convolution parameters and the modulation parameters respectively, and then perform the modulation convolution accordingly, can realize the modulation convolution with adaptive intensity, help to achieve accurate modulation of the feature channel, and can implicitly control the watermark intensity.

[0044] In a specific implementation scenario, a multi-layer perceptron can be used to make a first prediction of the image features to obtain the convolution parameters. For ease of description, the weight parameter in the convolution parameters can be recorded as W dyn , the bias parameter in the convolution parameter can be recorded as b. As a possible example, the weight parameter W dyn ∈R 64*64*3*3 , bias parameter b∈R 64 Of course, the above example is only one possible example in actual application, and other possible situations will not be given one by one here.

[0045] In a specific implementation scenario, the image features and sequence features can be first concatenated in the channel dimension, and then the first convolution operation, activation function operation, second convolution operation, and normalization operation are performed in sequence to obtain the modulation parameter. For the convenience of description, the modulation parameter can be recorded as M, which can be expressed as M=σ(conv(GELU((conv([z img ,wm feat ])))))), where the square brackets indicate splicing in the channel dimension, from the inside out they are the first convolution conv, the activation function operation GELU, the second convolution conv, and the normalization operation σ. As a possible example, the modulation parameter M∈R 1*32*32 Of course, the above example is only one possible example in actual application, and other possible situations will not be given one by one here.

[0046] In a specific implementation scenario, the weight parameters and bias parameters in the modulation parameters can be used to perform a convolution operation on the image features as the convolution feature z mod =conv2d(zimg ,W dyn )+b. On this basis, the modulation parameters can be used to control the intensity of the above convolution features, and then the second feature z mod_out =z mod ⊗M. As a possible example, z mod_out ∈R 64*32*32 Of course, the above example is only one possible example in actual application, and other possible situations will not be given one by one here.

[0047] Step S13: performing fusion decoding based on the first feature and the second feature to obtain a target image after the target character string is embedded as an image watermark in the original image.

[0048] In one implementation scenario, after obtaining the first feature and the second feature, a weighted fusion can be performed based on the first feature and the second feature to obtain a weighted feature, and the decoded image based on the weighted feature and the original image are fused point by point to obtain the target image. It should be noted that the decoded image and the original image can have the same resolution. In addition, the weighted fusion can be a linear fusion, that is, a weighting coefficient can be pre-set for fusing the first feature and the second feature. For the sake of ease of description, taking the example of pre-setting a weighting coefficient α, the weighted feature can be expressed as z fuse =αz main +(1-α)z mod_out On this basis, the above weighted features can be upsampled to the image size of the original image through a lightweight convolution decoder, to obtain a decoded image, and then fused with the pixel values ​​of the corresponding positions of the original image to obtain the target image. For ease of understanding, the target image can be represented as I marked =D φ (z fuse )+I, where I represents the original image, D φ Denotes the decoder, D φ (z fuse ) represents the decoded image. In the above method, weighted fusion is performed based on the first feature and the second feature to obtain a weighted feature. The decoded image based on the weighted feature is then fused point by point with the original image to obtain the target image. The decoded image and the original image have the same resolution, which can achieve residual connection with the original image, helping to ensure that the image watermark in the target image is visually imperceptible as much as possible.

[0049] In an implementation scenario, as a special example, please refer to Figure 2a , Figure 2a This is a process diagram of an embodiment of the watermark embedding method of the present application. Figure 2aAs shown, the original image can first be encoded by an encoder to obtain image features, while the target string serving as the image watermark can be sequentially projected through a hash network and a fully connected layer to obtain sequence features. Based on this, in one processing flow, deformable convolution can be performed on the image features to obtain deformed features, which are then concatenated with the sequence features and fed into windowed attention (as described above, such as dividing the fused sub-features and then performing self-attention processing with positional encoding) to obtain a first feature that combines image information and watermark information. In another processing flow, a first prediction can be performed on the image features to obtain convolution parameters. After fusing the image features with the sequence features, a second prediction can be performed to obtain modulation parameters. The image features are then modulated and convolved based on the convolution parameters and modulation parameters to obtain a second feature that combines image information and watermark information. Fusion decoding can then be performed based on the first and second features to obtain a target image with the same resolution as the original image and containing the target string as the image watermark. This allows the target string to be embedded as an image watermark in the original image, while minimizing the visual quality of the embedded watermark and improving anti-attack capabilities. After testing, the above-mentioned watermark embedding process steps of the embodiment of the present disclosure can significantly improve the robustness and concealment of watermark embedding while maintaining PSNR>38dB visual quality.

[0050] In one implementation scenario, to improve the efficiency of watermark embedding, a target image can be obtained by processing an original image and a target string using a watermark embedding model. The watermark embedding model can be jointly trained with a watermark extraction model based on sample images in the same training process. In each round of training, a sample original image and a sample string can be processed by the watermark embedding model to obtain a sample string as a sample target image after the image watermark is embedded into the sample original image. The sample target image is then attacked according to the selection probabilities of several watermark attack methods to obtain a sample attack image. The sample attack image can be used by the watermark extraction model to predict a predicted string for the image watermark. The selection probability of the watermark attack method can be determined by the bit error between the sample string in the previous training round and the predicted string extracted after the watermark attack method is applied. In this method, during the simulated attack process of model training, since the selection probability of the watermark attack method is determined by the bit error between the sample string in the previous training round and the predicted string extracted after the watermark attack method is applied, it helps force the model to automatically focus on the most effective attack method while maintaining its exploration capability, ultimately improving the transferability and robustness of adversarial examples.

[0051] In a specific implementation scenario, the watermark embedding model may include the following as a possible example: Figure 2aThe encoder, multi-layer perceptron, deformable convolution, window attention, modulated convolution, hash network, fully connected projection layer and fusion decoding and other network modules shown in the figure can be specifically referred to the relevant content about watermark embedding mentioned above. The network structure of the watermark embedding model is not limited here. In addition, the sample original image and the sample string can be processed by the watermark embedding model to obtain the sample string as the image watermark embedded in the sample original image and the sample target image after the sample target string is embedded in the sample original image. For details, please refer to the relevant description of the target string embedded in the original image as the image watermark, which will not be repeated here. Of course, although the watermark embedding model and the watermark extraction model are jointly trained, it does not mean that the two must be used together after the training is completed. In other words, after the training is completed, the two can be used separately and independently. For example, when only watermark embedding is required, only the watermark embedding model can be used, or when only watermark extraction is required, only the watermark extraction model can be used.

[0052] In a specific implementation scenario, several watermark attack methods may include, but are not limited to: geometric transformations implemented by STN, compression attacks simulated by DiffJPEG, and masking, etc. Watermark attack methods are not limited here. In addition, in the first round of training, the selection probability of each watermark attack method can be the same.

[0053] In a specific implementation scenario, for a watermark extraction model, when a test image with a character string as an image watermark is obtained (the test image is a sample attack image during the training process), in response to detecting a watermark extraction instruction for the test image, encoding can be performed based on the test image to obtain a word sequence, and the word sequence can contain characteristic information of the image watermark. For example, a codec structure combining a visual transformer (i.e., ViT) and an improved U-Net upsampling structure can be used to extract watermark information. That is, in terms of network structure, the encoder can use the basic ViT architecture to transform the test image (e.g., I aug ∈R 256*256*3 ) is divided into several non-overlapping blocks (e.g., 16*16 non-overlapping blocks), each of which can be converted into an embedding vector (384-dimensional embedding vector) through linear projection, and then encoded by multiple layers of Transformer to obtain a word sequence, such as token∈R 256*384 On this basis, we can reshape the word sequence to obtain the features to be sampled. Reshaping is mainly used to integrate the word sequence into a three-dimensional vector. For example, for the above word sequence token∈R 256*384For example, 256 dimensions can be reshaped into 16*16, that is, the features to be sampled can be converted to 16*16*384. Of course, the above example is only a possible example of feature reshaping. In other cases, it can be deduced by analogy, and no more examples are given here. After obtaining the features to be sampled, several transposed convolutions can be performed in sequence based on the features to be sampled to gradually improve the feature resolution and obtain the output features of each transposed convolution. Still taking the aforementioned 16*16*384 features to be sampled as an example, when four transposed convolutions are performed in sequence, the output feature S1∈R of the first transposed convolution can be obtained. 32*32*256 , the output feature S2∈R of the second transposed convolution 64*64*128 , the output feature S3∈R of the third transposed convolution 128*128*64 , the output feature S4∈R of the fourth transposed convolution 256*256*32 . On this basis, during the i-th upsampling process: the output feature channel of the i-th transposed convolution of the sampled feature can be adjusted to obtain the current feature, and the transposed convolution and attention processing can be performed in sequence based on the output feature of the i-1-th upsampling to obtain the reference feature, and the current feature and the reference feature can be fused to obtain the output feature of the i-th upsampling. It should be noted that the adjustment channel can be achieved through 1*1 convolution, and the attention processing can be a channel-space dual attention mechanism. In order to facilitate the description of the output feature of the i-th upsampling, it can be expressed as: F i_layer =Attention(conv_transpose(F i-1_layer ))+skipconnection(S i ), where F i-1_layer represents the output feature of the i-1th upsampling, conv_transpose represents transposed convolution, Attention represents attention processing (such as channel-space dual attention mechanism), S irepresents the output features of the i-th transposed convolution, and skipconnection indicates adjusting the channel dimension (e.g., using a 1x1 convolution). Furthermore, upsampling can be performed the same number of times as the transposed convolutions are performed sequentially based on the features to be sampled. For example, if four transposed convolutions are performed sequentially based on the features to be sampled in the above example, four upsampling operations can be performed accordingly. Finally, the output features of the last upsampling operation can be quantized to obtain the predicted string in the image to be tested, which serves as the image watermark. In the above method, when a test image with a character string as an image watermark is obtained, in response to detecting a watermark extraction instruction for the test image, encoding is performed based on the test image to obtain a word unit sequence, and the word unit sequence contains feature information of the image watermark, and is reshaped based on the word unit sequence to obtain a feature to be sampled, and then several transposed convolutions are performed in sequence based on the feature to be sampled to gradually improve the feature resolution, and the output features of each transposed convolution are obtained. Therefore, in the process of performing upsampling for the i-th time: the output feature channel of the i-th transposed convolution of the sampled feature is adjusted to obtain the current feature, and transposed convolution and attention processing are performed in sequence based on the output feature of the i-1-th upsampling to obtain the reference feature, and the current feature and the reference feature are fused to obtain the output feature of the i-th upsampling, and then the output feature of the last upsampling is quantized to obtain the predicted character string as the image watermark in the test image. Watermark extraction can be achieved through the attention processing of the residual connection in each upsampling stage, which helps to improve the accuracy of sound pickup extraction.

[0054] In a specific implementation scenario, after extracting the predicted string from the sample attack image after attacking the sample target image with the watermark attack method, it can be compared with the sample string of the sample original image to obtain the bit error between the two, and the selection probability of the watermark attack method can be determined based on this. For example, the selection probability can be positively correlated with the bit error, that is, the larger the bit error, the greater the selection probability, and conversely, the smaller the bit error, the smaller the selection probability. As a possible example, to improve the accuracy of the selection probability, during the i-th round of training, each batch can select a watermark attack method based on the selection probability. Then, the corresponding bit error of each sample original image in the batch can be determined according to the above process. All bit errors in the batch are statistically calculated (e.g., averaged) as the statistical bit error corresponding to the watermark formula method. After obtaining the statistical bit error corresponding to each watermark attack method, the selection probability of each watermark attack method in the i+1th round can be determined based on this. Specifically, for each watermark attack method, we can obtain the bit error between the sample string in the previous training process and the predicted string extracted after the watermark attack method is applied. Then, we normalize the bit error corresponding to each watermark attack method in the previous training process to obtain the selection probability of each watermark attack method in the current training process. For ease of description, for the i-th watermark attack method, its selection probability in the t+1 round of training process can be expressed as:

[0055]

[0056] In the above formula, represents the bit error of the i-th watermark attack method in the t-th round of training process, represents the temperature coefficient, which is used to control the sharpness of the probability distribution. represents the probability of selection of the i-th watermark attack method in the t+1 round of training. In addition, to prevent early convergence or the complete elimination of certain watermark formula methods, if the probability of selecting a watermark attack method is lower than the probability threshold, the watermark attack method selection instruction is reset to the probability threshold. In the above method, for each watermark attack method, the bit error between the sample string in the previous round of training and the predicted string extracted after the watermark attack method is applied is obtained, and the bit error corresponding to each watermark attack method in the previous round of training is normalized to obtain the selection probability of each watermark attack method in the current round of training. This can force the model to automatically focus on the current more effective watermark attack method while maintaining its exploration capability, ultimately improving the transferability and robustness of adversarial samples.

[0057] In a specific implementation scenario, after obtaining the bit error, the training loss for joint training can be obtained based on the reconstruction loss and bit error between the sample target image and the sample original image. The network parameters of the watermark embedding model and the watermark extraction model can then be adjusted based on the training loss. It should be noted that the training loss is positively correlated with the reconstruction loss and the bit error. In other words, during the joint training process, the network parameters can be adjusted with the goal of minimizing the training loss. This forces the watermark embedding model to ensure visual consistency between the image before and after embedding, while embedding the string as an image watermark, and forces the watermark extraction model to extract the watermark information as accurately as possible.

[0058] In an implementation scenario, as a special example, please refer to Figure 2b , Figure 2b This is a process diagram of an embodiment of the watermark attack process of this application. Figure 2b As shown, when a test image with a string as an image watermark is obtained, after the test image is attacked by the watermark attack method, the image after the attack can be further encoded (such as the aforementioned ViT and other related descriptions) to obtain a word sequence, and then reshaped based on the word sequence to obtain the features to be sampled. The residual attention of the transposed convolution is introduced to the sampled features (such as the aforementioned transposed convolution and upsampling related descriptions) to obtain the output features, and quantized to obtain the predicted string as the image watermark. The difference between the actual embedded string of the attack test image and the predicted string is combined to obtain the bit error, which is used to measure the effectiveness of the watermark attack method. For watermark formula methods that induce higher bit errors, their selection probability can be proportionally increased in the next round of training process, while for watermark attack methods with poor performance, their selection probability can be reduced accordingly. After testing, the above method can force the model to be continuously exposed to the most challenging attack environment while ensuring the stability of training as much as possible, and ultimately achieve excellent robustness with a bit error of less than 2% under new attacks that have never been seen. Of course, Figure 2b The example shown is only one possible implementation example of the watermark attack process, and is not limited to using other processes to implement watermark attacks.

[0059] The above scheme performs feature encoding based on the original image to obtain the image features of the original image, and performs feature projection based on the target string to be embedded to obtain the sequence features of the target string, thereby deforming and convolving the image features based on the offset obtained by offset prediction of the image features to obtain the deformed features of the image features, and performing feature interaction based on the deformed features and the sequence features to obtain the first feature, and performing modulated convolution based on the image features and the sequence features to obtain the second feature, and then performing fusion decoding based on the first feature and the second feature to obtain the target image after the target string is embedded in the original image as an image watermark. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed according to the image features to obtain the offset and deformed convolution is performed on the image features based on this, which can adaptively capture the local feature structure of the original image, and then when performing feature interaction, it helps to make the watermark embedding avoid the visually perceptible area as much as possible, so as to ensure the visual quality after the watermark is embedded as much as possible. On the other hand, in another feature processing flow, modulated convolution is performed by the image features and the sequence features, and fused decoding is performed with the above-mentioned feature processing flow, which can realize dual-path collaborative watermark embedding, which helps to improve the anti-attack capability. Therefore, the visual quality of the watermark after embedding can be guaranteed as much as possible and the anti-attack capability can be improved.

[0060] See also Figure 3 , Figure 3 The schematic diagram of the framework of an embodiment of the watermark embedding device of the present application is shown in FIG. The watermark embedding device 30 includes: a feature preparation module 31, a branch processing module 32, and a fusion decoding module 33. The feature preparation module 31 is used to perform feature encoding based on the original image to obtain image features of the original image, and perform feature projection based on the target string to be embedded to obtain sequence features of the target string; the branch processing module 32 is used to perform deformation convolution on the image features based on the offset obtained by offset prediction of the image features to obtain deformation features of the image features, and perform feature interaction based on the deformation features and the sequence features to obtain a first feature, and perform modulated convolution based on the image features and the sequence features to obtain a second feature; the fusion decoding module 33 is used to perform fusion decoding based on the first feature and the second feature to obtain the target image after the target string is embedded as an image watermark in the original image.

[0061] In the above scheme, the watermark embedding device 30 performs feature encoding based on the original image to obtain the image features of the original image, and performs feature projection based on the target string to be embedded to obtain the sequence features of the target string, thereby deforming and convolving the image features based on the offset obtained by offset prediction of the image features to obtain the deformed features of the image features, and performing feature interaction based on the deformed features and the sequence features to obtain the first feature, and performing modulation convolution based on the image features and the sequence features to obtain the second feature, and then performing fusion decoding based on the first feature and the second feature to obtain the target image after the target string is embedded in the original image as the image watermark. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed based on the image features to obtain the offset and deformed convolution is performed on the image features based on the offset, which can adaptively capture the local feature structure of the original image, and then when the feature interaction is performed, it helps to avoid the visually perceptible area as much as possible when the watermark is embedded, so as to ensure the visual quality after the watermark is embedded as much as possible. On the other hand, in another feature processing flow, modulation convolution is performed by the image features and the sequence features, and fusion decoding is performed with the above-mentioned feature processing flow, which can realize dual-path collaborative watermark embedding, which helps to improve the anti-attack capability. Therefore, the visual quality of the watermark after embedding can be guaranteed as much as possible and the anti-attack capability can be improved.

[0062] In some disclosed embodiments, the branch processing module 32 includes an offset prediction submodule for performing offset prediction based on image features to obtain an offset; wherein the offset includes: the offset value of each convolution position of each feature position in the image feature within the convolution range; the branch processing module 32 includes a deformation convolution submodule for, for each feature position in the process of performing deformation convolution, based on the weight factors of each convolution position of the feature position within the convolution range, performing weighted summation on the image subfeatures at the target position after the convolution position in the image feature is offset by the offset value, to obtain the deformation subfeature of the feature position; the branch processing module 32 includes a first combination submodule for obtaining a deformation feature based on the deformation subfeatures of each feature position in the image feature.

[0063] In some disclosed embodiments, the image features and the sequence features have the same resolution, and the branch processing module 32 includes a feature splicing submodule for splicing based on the deformation features and the sequence features in the channel dimension to obtain a fused feature; the branch processing module 32 includes a feature division submodule for dividing based on the fusion features in the resolution dimension to obtain a number of fused sub-features; the branch processing module 32 includes a feature enhancement submodule for processing the fused sub-features based on the self-attention mechanism that introduces position encoding for each fused sub-feature to obtain an enhanced sub-feature of the fused sub-feature; the branch processing module 32 includes a second combination submodule for combining based on the enhanced sub-features of each fused sub-feature to obtain an enhanced feature; the branch processing module 32 includes a feature fusion submodule for fusing based on the enhanced feature and the sequence feature to obtain a first feature.

[0064] In some disclosed embodiments, the position encoding is generated by a learnable parameter; and / or, during the processing of the self-attention mechanism, the interaction sub-feature between the query feature and the key feature of the fused sub-feature is fused with the position encoding, and then interacts with the value feature of the fused sub-feature to obtain an enhanced sub-feature of the fused sub-feature.

[0065] In some disclosed embodiments, the feature fusion submodule includes a weight prediction unit for performing weight prediction based on the enhancement feature and the sequence feature to obtain a weight image for fusing the enhancement feature and the sequence feature; wherein the weight image has the same resolution as the image feature and the sequence feature; the feature fusion submodule includes a weighted processing unit for performing weighted processing on the enhancement feature and the sequence feature based on the weight image to obtain the first feature.

[0066] In some disclosed embodiments, the branch processing module 32 includes a parameter prediction submodule, which is used to perform a first prediction based on image features to obtain convolution parameters of modulated convolution, and to perform a second prediction based on image features and sequence features to obtain modulation parameters of modulated convolution; wherein the convolution parameters include weight parameters and bias parameters, and the modulation parameters include intensity values ​​of each feature position when performing modulated convolution; the branch processing module 32 includes a modulation convolution submodule, which is used to perform modulated convolution on the image features based on the convolution parameters and modulation parameters to obtain second features.

[0067] In some disclosed embodiments, the fusion decoding module 33 includes a weighted fusion submodule for performing weighted fusion based on the first feature and the second feature to obtain a weighted feature; the fusion decoding module 33 includes a point-by-point fusion submodule for performing point-by-point fusion of the decoded image and the original image based on the weighted feature to obtain a target image; wherein the decoded image has the same resolution as the original image.

[0068] In some disclosed embodiments, a target image is obtained by processing an original image and a target character string by a watermark embedding model. The watermark embedding model and the watermark extraction model are jointly trained based on sample images in the same training process. In each round of the training process: the sample original image and the sample character string are processed by the watermark embedding model to obtain a sample character string as a sample target image after the image watermark is embedded in the sample original image, and the sample target image is attacked according to the respective selection probabilities of several watermark attack methods to obtain a sample attack image. The sample attack image is used for the watermark extraction model to predict a predicted character string as the image watermark, and the selection probability of the watermark attack method is determined by the bit error between the sample character string in the previous round of training process and the predicted character string extracted after the watermark attack method is adopted.

[0069] In some disclosed embodiments, the watermark embedding device 30 includes an error measurement module for obtaining, for each watermark attack method, a bit error between a sample string in the previous training process and a predicted string extracted after the watermark attack method is adopted; the watermark embedding device 30 includes a normalization module for normalizing the bit errors corresponding to each watermark attack method in the previous training process to obtain the selection probability of each watermark attack method in the current training process.

[0070] In some disclosed embodiments, the watermark embedding model and the watermark extraction model are applied separately and independently after training; and / or, in the first round of training process, the selection probability of each watermark attack method is the same; and / or, the selection probability is positively correlated with the bit error; and / or, when the selection probability of the watermark attack method is lower than the probability threshold, the watermark attack method selection instruction is reset to the probability threshold.

[0071] In some disclosed embodiments, the watermark embedding device 30 includes an image encoding module for, when obtaining a test image with a character string as an image watermark, encoding the test image in response to detecting a watermark extraction instruction for the test image to obtain a word sequence; wherein the word sequence contains feature information of the image watermark; the watermark embedding device 30 includes a feature reshaping module for reshaping based on the word sequence to obtain a feature to be sampled; the watermark embedding device 30 includes a transposed convolution module for sequentially performing a plurality of transposed convolutions based on the feature to be sampled to gradually improve the feature resolution, and obtain the output of each transposed convolution. The watermark embedding device 30 includes an upsampling module, which is used to: adjust the channel of the output feature of the i-th transposed convolution performed on the sampled feature during the i-th upsampling, obtain the current feature, and perform transposed convolution and attention processing in sequence based on the output feature of the i-1-th upsampling to obtain the reference feature, and fuse the current feature and the reference feature to obtain the output feature of the i-th upsampling; the watermark embedding device 30 includes a feature quantization module, which is used to quantize the output feature of the last upsampling to obtain a predicted character string in the image to be tested as an image watermark.

[0072] See also Figure 4 , Figure 4 : This is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 40 includes at least a memory 41 and a processor 42 coupled to each other. The memory 41 stores at least program instructions, and the processor 42 is used to execute the program instructions to implement the steps of any of the above-mentioned watermark embedding method embodiments. For details, please refer to the aforementioned disclosed embodiments and will not be repeated here. As a possible example, the electronic device 40 may include but is not limited to mobile phones, tablet computers, learning machines, smart large screens, servers and other devices. The specific type of the electronic device 40 is not limited here.

[0073] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps of any of the above-mentioned watermark embedding method embodiments. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor. In addition, the processor 42 may be implemented by an integrated circuit chip.

[0074] In the above scheme, the electronic device 40 performs feature encoding based on the original image to obtain the image features of the original image, and performs feature projection based on the target string to be embedded to obtain the sequence features of the target string, thereby deforming and convolving the image features based on the offset obtained by offset prediction of the image features to obtain the deformed features of the image features, and performing feature interaction based on the deformed features and the sequence features to obtain the first feature, and performing modulated convolution based on the image features and the sequence features to obtain the second feature, and then performing fusion decoding based on the first feature and the second feature to obtain the target image after the target string is embedded in the original image as an image watermark. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed based on the image features to obtain the offset and deformed convolution is performed on the image features based on the offset, which can adaptively capture the local feature structure of the original image, and then when the feature interaction is performed, it helps to avoid the visually perceptible area as much as possible when the watermark is embedded, so as to ensure the visual quality of the watermark as much as possible. On the other hand, in another feature processing flow, modulated convolution is performed by the image features and the sequence features, and fusion decoding is performed with the above-mentioned feature processing flow, which can realize dual-path collaborative watermark embedding, which helps to improve the anti-attack capability. Therefore, the visual quality of the watermark after embedding can be guaranteed as much as possible and the anti-attack capability can be improved.

[0075] See also Figure 5 , Figure 5 The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned watermark embedding method embodiments.

[0076] In the above scheme, the computer-readable storage medium 50 performs feature encoding based on the original image to obtain image features of the original image, and performs feature projection based on the target string to be embedded to obtain sequence features of the target string, thereby deforming and convolving the image features based on the offset obtained by offset prediction of the image features to obtain deformed features of the image features, and performing feature interaction based on the deformed features and the sequence features to obtain a first feature, and performing modulated convolution based on the image features and the sequence features to obtain a second feature, and then performing fusion decoding based on the first feature and the second feature to obtain the target image after the target string is embedded in the original image as an image watermark. Therefore, after obtaining the image features and the sequence features, on the one hand, in one feature processing flow, offset prediction is performed based on the image features to obtain the offset and deformed convolution is performed on the image features based on the offset, which can adaptively capture the local feature structure of the original image. Then, when performing feature interaction, it helps to avoid visually perceptible areas as much as possible when embedding the watermark, thereby ensuring the visual quality of the embedded watermark as much as possible. On the other hand, in another feature processing flow, modulated convolution is performed through the image features and the sequence features, and fusion decoding is performed with the aforementioned feature processing flow, which can achieve dual-path collaborative watermark embedding, which helps to improve anti-attack capabilities. Therefore, the visual quality of the watermark after embedding can be guaranteed as much as possible and the anti-attack capability can be improved.

[0077] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0078] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0079] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0080] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0081] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0082] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0083] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A watermark embedding method, characterized in that: include: Performing feature encoding based on the original image to obtain image features of the original image, and performing feature projection based on the target character string to be embedded to obtain sequence features of the target character string; Performing deformation convolution on the image feature based on an offset obtained by offset prediction of the image feature to obtain a deformation feature of the image feature, performing feature interaction with the sequence feature based on the deformation feature to obtain a first feature, and performing modulated convolution based on the image feature and the sequence feature to obtain a second feature; Performing fusion decoding based on the first feature and the second feature to obtain a target image after the target character string is embedded in the original image as an image watermark; The step of performing deformation convolution on the image feature based on the offset obtained by offset prediction of the image feature to obtain the deformation feature of the image feature includes: Performing offset prediction based on the image features to obtain the offset; wherein the offset includes: offset values ​​of each convolution position within the convolution range of each feature position in the image features; During the deformation convolution, for each feature position, based on the weight factors of each convolution position within the convolution range of the feature position, weighted summation is performed on the image sub-features at the target position after the convolution position in the image feature is shifted by the offset value, to obtain a deformation sub-feature of the feature position; The deformation feature is obtained based on the deformation sub-features of each feature position in the image feature.

2. The method according to claim 1, characterized in that The image feature and the sequence feature have the same resolution, and the first feature is obtained by performing feature interaction based on the deformation feature and the sequence feature, including: Based on the deformation feature and the sequence feature, splicing is performed in the channel dimension to obtain a fusion feature; Based on the fusion feature, a plurality of fusion sub-features are obtained by dividing the fusion feature in the resolution dimension; For each of the fused sub-features, processing the fused sub-feature based on a self-attention mechanism that introduces position encoding to obtain an enhanced sub-feature of the fused sub-feature; Combining the enhancer features based on the respective fusion subfeatures to obtain an enhancement feature; The first feature is obtained by fusing the enhanced feature with the sequence feature.

3. The method according to claim 2, characterized in that The position encoding is generated by learnable parameters; And / or, during the processing of the self-attention mechanism, the interaction sub-feature between the query feature and the key feature of the fused sub-feature is fused with the position encoding, and then interacts with the value feature of the fused sub-feature to obtain an enhanced sub-feature of the fused sub-feature.

4. The method according to claim 2, characterized in that The fusing the enhancement feature with the sequence feature to obtain the first feature includes: Performing weight prediction based on the enhanced features and the sequence features to obtain a weight image for fusing the enhanced features and the sequence features; wherein the weight image has the same resolution as the image features and the sequence features; The enhanced feature and the sequence feature are weighted based on the weight image to obtain the first feature.

5. The method according to claim 1, wherein The performing modulated convolution based on the image feature and the sequence feature to obtain a second feature includes: Performing a first prediction based on the image features to obtain convolution parameters of the modulated convolution, and performing a second prediction based on the image features and the sequence features to obtain modulation parameters of the modulated convolution; wherein the convolution parameters include weight parameters and bias parameters, and the modulation parameters include intensity values ​​of each feature position when performing the modulated convolution; Performing modulation convolution on the image feature based on the convolution parameter and the modulation parameter to obtain the second feature.

6. The method according to claim 1, characterized in that The performing fusion decoding based on the first feature and the second feature to obtain a target image after the target character string is embedded in the original image as an image watermark includes: Performing weighted fusion based on the first feature and the second feature to obtain a weighted feature; The decoded image based on the weighted features and the original image are fused point by point to obtain the target image; wherein the decoded image and the original image have the same resolution.

7. The method according to claim 1, characterized in that The target image is obtained by processing the original image and the target character string by the watermark embedding model. The watermark embedding model and the watermark extraction model are jointly trained based on the sample image in the same training process. In each round of training process: the sample original image and the sample character string are processed by the watermark embedding model to obtain a sample target image after the sample character string is embedded in the sample original image as an image watermark, and the sample target image is attacked according to the respective selection probabilities of several watermark attack methods to obtain a sample attack image. The sample attack image is used for the watermark extraction model to predict the predicted character string as the image watermark, and the selection probability of the watermark attack method is determined by the bit error between the sample character string in the previous round of training process and the predicted character string extracted after the watermark attack method is adopted.

8. The method according to claim 7, characterized in that The selection probability of the watermark attack mode is updated by the following steps: For each watermark attack method, obtaining the bit error between the sample character string in the previous training process and the predicted character string extracted after the watermark attack method is adopted; Based on the bit errors corresponding to the various watermark attack modes in the previous round of training process, normalization is performed to obtain the selection probabilities of the various watermark attack modes in the current round of training process.

9. The method according to claim 7, characterized in that The watermark embedding model and the watermark extraction model are applied separately and independently after training; and / or, in the first round of training process, the selection probability of each of the watermark attack modes is the same; And / or, the selection probability is positively correlated with the bit error; And / or, when the selection probability of the watermark attack mode is lower than a probability threshold, the watermark attack mode selection instruction is reset to the probability threshold.

10. The method according to any one of claims 1 to 9, characterized in that When obtaining the image to be tested with the character string as the image watermark, the method further includes: In response to detecting a watermark extraction instruction for the image to be tested, encoding the image to be tested to obtain a word-unit sequence; wherein the word-unit sequence contains characteristic information of the image watermark; Reshape based on the word sequence to obtain features to be sampled; Performing several transposed convolutions in sequence based on the features to be sampled to gradually improve feature resolution, and obtaining output features of each transposed convolution; During the i-th upsampling process: adjusting the channel of the output feature of the i-th execution of the transposed convolution on the feature to be sampled to obtain a current feature, and sequentially performing transposed convolution and attention processing based on the output feature of the i-1-th execution of the upsampling to obtain a reference feature, and fusing the current feature and the reference feature to obtain the output feature of the i-th execution of the upsampling; The output features of the last upsampling are quantized to obtain a predicted character string as an image watermark in the image to be tested.

11. A watermark embedding device, characterized in that: include: A feature preparation module is used to perform feature encoding based on the original image to obtain image features of the original image, and perform feature projection based on the target string to be embedded to obtain sequence features of the target string; a branch processing module, configured to perform deformation convolution on the image feature based on an offset obtained by performing offset prediction on the image feature to obtain a deformation feature of the image feature, perform feature interaction with the sequence feature based on the deformation feature to obtain a first feature, and perform modulated convolution based on the image feature and the sequence feature to obtain a second feature; A fusion decoding module is configured to perform fusion decoding based on the first feature and the second feature to obtain a target image after the target character string is embedded in the original image as an image watermark; wherein the deforming convolution of the image feature based on the offset obtained by offset prediction of the image feature to obtain the deformed feature of the image feature includes: Performing offset prediction based on the image features to obtain the offset; wherein the offset includes: offset values ​​of each convolution position within the convolution range of each feature position in the image features; During the deformation convolution, for each feature position, based on the weight factors of each convolution position within the convolution range of the feature position, weighted summation is performed on the image sub-features at the target position after the convolution position in the image feature is shifted by the offset value, to obtain a deformation sub-feature of the feature position; The deformation feature is obtained based on the deformation sub-features of each feature position in the image feature.

12. An electronic device, characterized in that: The watermark embedding method comprises at least a memory and a processor, wherein the memory stores at least program instructions, and the processor is used to execute the program instructions to implement the watermark embedding method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the watermark embedding method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Robust image watermarking method and system based on hierarchical attention feature fusion

    CN115908095A

  • Deep watermarking method for immune geometric distortion

    CN119273523A

Cited By

  • Method and system for generating image watermark based on potential variable optimization of diffusion model

    CN121458514A