Vector character watermark embedding and extracting method and device based on multi-modal features
By using multimodal feature fusion and loss function design, the diversity and robustness issues of Chinese text watermarking on simple characters were solved, achieving efficient embedding and extraction of vector character watermarks in multimodal scenarios, and improving the watermark's anti-attack capability and fidelity.
Patent Information
- Application Number
- CN202511040470.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies lack diversity and robustness in watermark embedding in Chinese text, especially for characters with simple glyph structures, where there is a lack of adversarial sample training. Furthermore, they fail to effectively handle vector and raster multimodal data and have weak anti-distortion capabilities.
A multimodal feature fusion approach is adopted, and a vector character watermark embedding and extraction method based on multimodal features is designed by constructing triplet, redundancy and adversarial loss functions. This includes joint training of encoder and decoder, and multi-task learning is achieved by utilizing distribution loss, invariance loss, spatial loss, redundancy loss and adversarial loss.
It improves the anti-attack capability of vector character watermarks, ensures the robustness of watermark features during character geometric transformations, enhances the concealment and fidelity of watermarks, increases the capacity and reliability of watermark information, and adapts to watermark extraction under different conditions.
Smart Images

Figure CN120852137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital watermarking, and in particular to a watermark embedding and extraction method and device based on deep learning. Background Art
[0002] As a key means of copyright protection and leakage traceability, digital watermarking technology has been widely applied in various fields of society. With the in-depth research, remarkable achievements have been made in the fields of image, audio and video watermarking. However, as the most commonly used carrier of daily information, the research on watermarking technology for text documents is still in its infancy.
[0003] Traditional text watermarking schemes mainly include three categories: based on document format parsing, based on document images, and based on font glyphs. The recent research focus is on text watermarking schemes based on character feature changes, which can be further divided into two types: based on manually designed character features and based on deep learning to generate watermark character features. Generally, the manually designed feature scheme is more suitable for languages with rich glyph structures such as Chinese; while the generation scheme based on deep learning shows better robustness in languages with relatively simple glyphs such as English and Tibetan.
[0004] For characters that are frequently used in the Chinese scenario but have simple glyph structures (such as the numbers "one" to "nine", "chuan", etc.), the existing technology faces challenges in terms of the diversity of watermark embedding and the robustness of extraction. In addition, there are also problems such as the lack of adversarial sample training, weak anti-distortion ability, and the lack of unified processing of vector and raster multimodality in watermark extraction.
[0005] In view of this, the present invention is specifically proposed. Summary of the Invention
[0006] The technical problem solved by the present invention is: overcoming the deficiencies of the prior art, and providing a vector character watermark embedding and extraction method based on multimodal features. The vector character watermark embedding and extraction method based on multimodal features provides information redundancy and complementarity through multimodal feature fusion, constructs a loss function system with extremely strong robustness (especially triplet, redundancy, and adversarial loss), so that the vector character watermark embedding has a strong anti-attack ability.
[0007] An embodiment of the present invention provides a vector character watermark embedding and extraction method based on multimodal features, including: S1. Traverse each character c in the character set C to generate the original glyph G0 corresponding to the original vector character, the watermark variant G w of the original glyph G0, and the adversarial sample glyph G adv , annotate the metadata of the character c, the original glyph G0, the watermark variant G w and the adversarial sample glyph G adv to construct a data set; S2. Calculate the feature distribution parameters {uc, ∑c} of the character c, where uc is the mean and ∑c is the covariance matrix. Based on the encoder E, calculate the feature distribution parameters {uc, ∑c} of the original glyph G0 and the watermark variant G0. w Extract features and calculate the distribution loss L. dist Invariance loss L inv and space loss L space Let L = L be the total loss function minimized. dist + 0.5L inv + 0.3L space When the total loss L is less than the preset threshold a, the parameters of the encoder E are frozen to obtain an encoder E1 with fixed parameters. S3. Extract the watermark variant G based on the encoder E1. w Adversarial sample glyphs G adv The characteristics, combined with the triplet loss L triplet Message loss L msg Redundancy loss L red Alignment loss L align and combat losses L adv Multi-task learning is performed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The decoder D is jointly trained, and only the parameters of the decoder D are updated until the total loss L is reached. total Less than a preset threshold or the message loss L msg , countering losses L adv To achieve equilibrium, the loss function L... total =w1L triplet +w2L msg + w3L red + w4L align + w5L adv ; S4. Perform character detection on the document image from which the watermark is to be extracted, and extract the character image set G={G1,G2,...G...} of the document image. i ,...,G n For each character image G, ,} i If a vector path exists, extract vector features; if it is a raster character, extract raster features; otherwise, fuse dual-modal features to generate multimodal features. Based on the watermark strength parameter λ, adaptively configure the marginal parameters and redundancy mode of the decoder D. Input the multimodal features into the majority decoder D to decode and obtain bit value b. Reconstruct all bit value sequences to reconstruct the watermark message.
[0008] This application also provides a vector character watermark embedding and extraction device based on multimodal features, including: The dataset generation module D1 iterates through each character c in the character set C to generate the original glyph G0 corresponding to the original vector character, and the watermark variant G of the original glyph G0. w Adversarial sample glyphs G adv The character c, the original glyph G0, and the watermark variant G are labeled. w and adversarial sample glyphs G adv Metadata is used to build datasets; Encoder training module D2 calculates the feature distribution parameters {uc, ∑c} of the character c, where uc is the mean and ∑c is the covariance matrix, based on encoder E for the original glyph G0 and watermark variant G. w Extract features and calculate the distributed loss L. dist Invariance loss L inv and space loss L space Let L = L be the total loss function minimized. dist + 0.5L inv +0.3L space When the total loss L is less than the preset threshold a, the parameters of the encoder E are frozen to obtain an encoder E1 with fixed parameters. Decoder training module D3 extracts the watermark variant G based on encoder E1. w Adversarial sample glyphs G adv The characteristics, combined with the triplet loss L triplet Message loss L msg Redundancy loss L red Alignment loss L align and combat losses L adv Multi-task learning is performed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The decoder D is jointly trained, and only the parameters of the decoder D are updated until the total loss L is reached. total Less than a preset threshold or the message loss L msg , countering losses L adv To achieve equilibrium, the loss function L... total =w1L triplet +w2L msg + w3L red + w4L align + w5L adv ; The watermark extraction module D4 performs character detection on the document image from which the watermark is to be extracted, and extracts the character image set G={G1,G2,...G... i ,...,G n For each character image G, ,} iIf a vector path exists, extract vector features; if it is a raster character, extract raster features; otherwise, fuse dual-modal features to generate multimodal features. Based on the watermark strength parameter λ, adaptively configure the marginal parameters and redundancy mode of the decoder D. Input the multimodal features into the majority decoder D to decode and obtain bit value b. Reconstruct all bit value sequences to reconstruct the watermark message.
[0009] This application also provides a computer-readable storage medium storing computer-executable instructions for executing the vector character watermark embedding and extraction method based on multimodal features implemented in any of the above embodiments.
[0010] This application embodiment provides an electronic device, including a memory and a processor, wherein the memory stores the following instructions executable by the processor: for performing the steps of the vector character watermark embedding and extraction method based on multimodal features described in any of the above claims.
[0011] Compared with the prior art, the beneficial effects of the present invention are: 1) Extremely strong robustness: The "invariance loss" in encoder training and the characteristics based on vector features make the watermark features more resistant to geometric transformations such as translation, rotation, and scaling of characters. 2) High concealment and fidelity: The "distribution loss" and "spatial loss" during the encoder training phase directly constrain the impact of the embedded watermark on the visual appearance of the original character. The former ensures that the feature changes are statistically small or conform to a specific distribution, while the latter directly optimizes the degree of change to the character's geometry, ensuring that the watermarked character is almost visually indistinguishable from the original character; 3) High capacity and high reliability: Redundancy loss forces watermark information to be repeatedly or consistently encoded at multiple locations or modes. This improves resistance to local damage and decoding accuracy, especially after lossy transmission or attacks. The "watermark strength parameter-based adaptive decoder configuration" in the watermark extraction stage can dynamically adjust the decoding strategy according to the detected signal quality, optimizing watermark performance under different conditions. 4) End-to-end learning optimization: The encoder and decoder are jointly optimized and trained through a series of interrelated and complementary loss functions, which overcomes the limitations of manually designing embedding rules and extraction algorithms in traditional methods and achieves a better performance balance. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0014] Figure 1 This is a flowchart of the vector character watermark embedding and extraction method based on multimodal features proposed in this invention.
[0015] Figure 2 This is a flowchart illustrating a method for embedding and extracting vector character watermarks based on multimodal features in one embodiment of the present invention.
[0016] Figure 3 This is a structural diagram of the vector character watermark embedding and extraction device based on multimodal features proposed in this invention.
[0017] Figure labeling: D1 Dataset generation module; D2 Encoder training module; D3 Decoder training module; D4 Watermark extraction module. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0019] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "an," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two.
[0020] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0021] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0022] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.
[0023] The optional embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0024] The first embodiment of the present invention, as follows: Figure 1 As shown, a vector character watermark embedding and extraction method based on multimodal features includes: S1. Traverse each character c in the character set C to generate the original glyph G0 corresponding to the original vector character, and the watermark variant G of the original glyph G0. w Adversarial sample glyphs G adv The character c, the original glyph G0, and the watermark variant G are labeled. w and adversarial sample glyphs G adv Metadata is used to build datasets; S2. Calculate the feature distribution parameters {uc, ∑c} of the character c, where uc is the mean and ∑c is the covariance matrix. Based on the encoder E, calculate the feature distribution parameters {uc, ∑c} of the original glyph G0 and the watermark variant G0. w Extract features and calculate the distribution loss L. dist Invariance loss L inv and space loss L space Let L = L be the total loss function minimized. dist + 0.5L inv + 0.3L space When the total loss L is less than the preset threshold a, the parameters of the encoder E are frozen to obtain an encoder E1 with fixed parameters. S3. Extract the watermark variant G based on the encoder E1. w Adversarial sample glyphs G adv The characteristics, combined with the triplet loss L triplet Message loss Lmsg Redundancy loss L red Alignment loss L align and combat losses L adv Multi-task learning is performed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The decoder D is jointly trained, and only the parameters of the decoder D are updated until the total loss L is reached. total Less than a preset threshold or the message loss L msg , countering losses L adv To achieve equilibrium, the loss function L... total =w1L triplet +w2L msg + w3L red + w4L align + w5L adv ; S4. Perform character detection on the document image from which the watermark is to be extracted, and extract the character image set G={G1,G2,...G...} of the document image. i ,...,G n For each character image G, ,} i If a vector path exists, extract vector features; if it is a raster character, extract raster features; otherwise, fuse dual-modal features to generate multimodal features. Based on the watermark strength parameter λ, adaptively configure the marginal parameters and redundancy mode of the decoder D. Input the multimodal features into the majority decoder D to decode and obtain bit value b. Reconstruct all bit value sequences to reconstruct the watermark message.
[0025] Further, the construction of the dataset in step S1 includes: S11. Traverse each character c in the character set C, and generate the original glyph G0 = R(c) corresponding to the original vector character through the rendering function R; S12. For each bit value b∈{0,1}, generate a watermark variant G of the original glyph G0 using a perturbation function P. w = P(G0, b); S13. For the watermark variant G w Apply at least one distortion type d∈{rotate,occlude, noise} to generate adversarial sample glyphs G. adv = D(G w , d); S14. Label the character c, the original glyph G0, and the watermark variant G. w and adversarial sample glyphs G adv The metadata is used to construct the dataset.
[0026] Furthermore, in step S2, the preset threshold a = 0.05.
[0027] Further, in step S2, the distribution loss L dist The invariance loss L is the negative logarithmic probability of the feature in its character distribution. inv The spatial loss L is the KL divergence between the original feature distribution and the watermark feature distribution. space The Euclidean distance between the original feature and the watermark feature is given.
[0028] Furthermore, in step S3, the triplet loss L triplet The message loss L is calculated based on the feature distances between anchor points, positive samples, and negative samples. msg The cross-entropy loss, the redundancy loss L red The alignment loss L is calculated based on sub-region features. align The alignment of constrained text and formula features, the adversarial loss L adv Generated based on FGSM attack.
[0029] Further, in step S3, w2 = (0.7 - 0.4λ), w4 = 0.1, and w5 = 0.1λ; Further, in step S2, the encoder E includes a multi-scale convolutional module, a spatial attention module, and a global feature generation layer, wherein: the multi-scale convolutional module contains four convolutional layers with a stride of 2, and the number of channels is 64, 128, 256, and 512 respectively; the spatial attention module receives the output of the last convolutional layer and generates channel- and spatially weighted features; the global feature generation layer outputs a 128-dimensional feature vector through global average pooling and a fully connected layer.
[0030] Further, in step S3, the decoder D includes: a message decoding branch containing two fully connected layers and a sigmoid activation function, outputting bit probability values; a redundancy module that dynamically generates redundant codes based on the watermark strength parameter λ and fuses them into the message decoding branch; an adversarial branch containing three fully connected layers, outputting adversarial scores; a triplet optimization module that calculates the feature distance of watermark variants of the same character; and a cross-modal alignment module that fuses vector path features and raster image features and performs spatial constraints.
[0031] To further understand this technical solution, the present invention combines... Figure 2 Please provide a detailed explanation.
[0032] The glyph images (including the original glyph G0 and the watermark variant Gw, all grayscale images with a uniform size of 64×64×164×64×1) from the dataset are input into the encoder E to obtain a 128-dimensional feature vector. Specifically, this includes: inputting the glyph images into the multi-scale convolution module to obtain a 4×4×512 feature vector; inputting the 4×4×512 feature vector into the spatial attention module to obtain a 4×4×512 weighted feature vector; and inputting the 4×4×512 weighted feature vector into a fully connected layer to obtain a 128-dimensional feature vector (corresponding to the feature distribution parameters in step S2).
[0033] The 128-dimensional feature vector is input into decoder D to obtain the reconstructed bit values b ∈ {0,1}, which are used to reassemble the watermark message. The input and output of each branch or module are: The message decoding branch takes the 128-dimensional feature vector as input and outputs bit probability values. The redundant decoding module takes the 128-dimensional feature vector and the watermark strength parameter λ as input and outputs the redundant features fused into the message decoding branch. The adversarial branch takes the 128-dimensional feature vector as input and outputs the adversarial score. The triplet optimization module takes the features of different watermark variants of the same character as input and outputs feature distance triplets. The cross-modal alignment module takes a vector path or raster image feature as input and outputs aligned and fused features.
[0034] In step S2, the distributed loss L dist The invariance loss L is to constrain the consistency between the features and the feature distribution parameters {uc, ∑c}. inv This is to ensure that G0 and G w Feature space alignment; the spatial loss L space This is to preserve the glyph structure information.
[0035] In step S3, the triplet loss L triplet To increase the feature distance between different watermark variants of the same character, triplet loss constrains the sample distance relationship in the feature space, forcing the feature distance between different watermark variants of the same character (such as adversarial examples and slightly deformed versions) to increase, while forcing the feature distance between similar watermark variants of different characters to decrease. This enhances the intra-class compactness and inter-class distinguishability of watermark features, significantly improves the robustness of the watermark system to character deformation and visual confusion attacks, and ensures that the decoder can accurately identify and extract the watermark signal in the interfered character.
[0036] The message loss L msgTo ensure the correctness of the decoded bit value b, the message loss directly quantizes the error between the watermark information recovered by the decoder and the original embedded message. As the core optimization goal of the training process, it drives the entire watermarking system: forcing the encoder to generate a decodeable feature representation, while guiding the decoder to accurately reconstruct the watermark bits. Essentially, it ensures the end-to-end fidelity of the watermark embedding and extraction process, which is the most basic guarantee for the accuracy and reliability of watermark information.
[0037] The redundancy loss L red To control the redundancy of watermark information, redundancy loss forces multiple modal features of a character (such as geometric contours, raster textures, and stroke topology) or different spatial regions (such as key strokes and local structures) to decode the same watermark message during the training phase, thereby achieving distributed embedding of watermark information: when a character encounters a local attack (such as occlusion, distortion, or noise) that causes some features to be destroyed, the decoder can still accurately recover the watermark based on the complete features of the surviving modal or region, thus transforming the principle of information redundancy into a robust advantage and significantly improving the survivability of the watermark under incomplete or interfered conditions.
[0038] The alignment loss L align To coordinate the vector and raster feature spaces, alignment loss constrains the consistency of representation of different modal features of characters (such as geometric information of vector contours and texture information of rasterized images) in the embedding space, forcing them to map to a unified semantic expression. This solves the problem of fusion between multi-source heterogeneous features, enabling watermark information to synergistically utilize the complementary advantages of all modalities, thereby significantly improving the system's cross-modal fault tolerance and decoding stability when a single modality is damaged (such as contour distortion or texture noise).
[0039] The resistance loss L adv To enhance resistance to attacks, adversarial loss actively introduces adversarial perturbation samples (such as subtle character deformations invisible to the human eye) into the watermark during training and forces the model to accurately decode the watermark even under these malicious perturbations. This drives the encoder to learn feature representations that are invariant to adversarial perturbations, while improving the decoder's resistance to interference. This enables the watermarking system to effectively resist targeted attacks aimed at destroying the watermark and significantly enhances its survivability and robustness in adversarial environments.
[0040] The technical solution of this invention adopts a two-stage joint training strategy to train the encoder and decoder separately. In the first stage, the encoder is pre-trained and frozen, and the feature distance loss and distribution matching loss are used to jointly optimize the encoder to accurately capture character outline deformation (to deal with rotation / occlusion) and ensure feature consistency within character categories. In the second stage, the decoder is trained end-to-end, and the triplet loss, redundancy loss, cross-modal loss, message classification loss and adversarial loss are jointly optimized. The decoder is optimized with the encoder fixed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The bit decoding and anti-attack capabilities are optimized simultaneously, and the watermark extraction is compatible with vector / raster documents.
[0041] The second embodiment of the present invention, as follows: Figure 3 As shown, a vector character watermark embedding and extraction device based on multimodal features, based on the first embodiment, includes: The dataset generation module D1 iterates through each character c in the character set C to generate the original glyph G0 corresponding to the original vector character, and the watermark variant G of the original glyph G0. w Adversarial sample glyphs G adv The character c, the original glyph G0, and the watermark variant G are labeled. w and adversarial sample glyphs G adv Metadata is used to build datasets; Encoder training module D2 calculates the feature distribution parameters {uc, ∑c} of the character c, where uc is the mean and ∑c is the covariance matrix, based on encoder E for the original glyph G0 and watermark variant G. w Extract features and calculate the distributed loss L. dist Invariance loss L inv and space loss L space Let L = L be the total loss function minimized. dist + 0.5L inv +0.3L space When the total loss L is less than the preset threshold a, the parameters of the encoder E are frozen to obtain an encoder E1 with fixed parameters. Decoder training module D3 extracts the watermark variant G based on encoder E1. w Adversarial sample glyphs G adv The characteristics, combined with the triplet loss L triplet Message loss L msg Redundancy loss L red Alignment loss L align and combat losses L adv Multi-task learning is performed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The decoder D is jointly trained, and only the parameters of the decoder D are updated until the total loss L is reached. totalLess than a preset threshold or the message loss L msg , countering losses L adv To achieve equilibrium, the loss function L... total =w1L triplet +w2L msg + w3L red + w4L align + w5L adv ; The watermark extraction module D4 performs character detection on the document image from which the watermark is to be extracted, and extracts the character image set G={G1,G2,...G... i ,...,G n For each character image G, ,} i If a vector path exists, extract vector features; if it is a raster character, extract raster features; otherwise, fuse dual-modal features to generate multimodal features. Based on the watermark strength parameter λ, adaptively configure the marginal parameters and redundancy mode of the decoder D. Input the multimodal features into the majority decoder D to decode and obtain bit value b. Reconstruct all bit value sequences to reconstruct the watermark message.
[0042] This application also provides a computer-readable storage medium storing computer-executable instructions for executing the vector character watermark embedding and extraction method based on multimodal features implemented in any of the above embodiments.
[0043] This application embodiment provides an electronic device, including a memory and a processor, wherein the memory stores the following instructions executable by the processor: for performing the steps of the vector character watermark embedding and extraction method based on multimodal features described in any of the above claims.
[0044] Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the embodiments described above is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.
[0045] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0046] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0047] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0048] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for embedding and extracting vector character watermarks based on multimodal features, characterized in that, include: S1. Traverse each character c in the character set C to generate the original glyph G0 corresponding to the original vector character, and the watermark variant G of the original glyph G0. w Adversarial sample glyphs G adv The character c, the original glyph G0, and the watermark variant G are labeled. w and adversarial sample glyphs G adv Metadata is used to build datasets; S2. Calculate the feature distribution parameters {uc, ∑c} of the character c, where uc is the mean and ∑c is the covariance matrix. Based on the encoder E, calculate the feature distribution parameters {uc, ∑c} of the original glyph G0 and the watermark variant G0. w Extract features and calculate the distribution loss L. dist Invariance loss L inv and space loss L space Let the total loss function be minimized as L = L dist +0.5L inv +0.3L space When the total loss L is less than the preset threshold a, the parameters of the encoder E are frozen to obtain an encoder E1 with fixed parameters. S3. Extract the watermark variant G based on the encoder E1. w Adversarial sample glyphs G adv The characteristics, combined with the triplet loss L triplet Message loss L msg Redundancy loss L red Alignment loss L align and combat losses L adv Multi-task learning is performed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The decoder D is jointly trained, and only the parameters of the decoder D are updated until the total loss L is reached. total Less than a preset threshold or the message loss L msg , countering losses L adv To achieve equilibrium, the loss function L... total =w1L triplet +w2L msg +w3L red +w4L align +w5L adv ; S4. Perform character detection on the document image from which the watermark is to be extracted, and extract the character image set G = {G1, G2, ... G...} of the document image. i ,...,G n For each character image G, ,} i If a vector path exists, extract vector features; if it is a raster character, extract raster features; otherwise, fuse dual-modal features to generate multimodal features. Based on the watermark strength parameter λ, adaptively configure the marginal parameters and redundancy mode of the decoder D. Input the multimodal features into the majority decoder D to decode and obtain bit value b. Reconstruct all bit value sequences to reconstruct the watermark message.
2. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, The dataset construction described in S1 includes: S11. Traverse each character c in the character set C, and generate the original glyph G0 = R(c) corresponding to the original vector character through the rendering function R; S12. For each bit value b∈{0,1}, generate a watermark variant G of the original glyph G0 using a perturbation function P. w =P(G0,b); S13. For the watermark variant G w Apply at least one distortion type d∈{rotate,occlude,noise} to generate adversarial sample glyphs. G adv =D(G w ,d); S14. Label the character c, the original glyph G0, and the watermark variant G. w and adversarial sample glyphs G adv The metadata is used to construct the dataset.
3. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, The preset threshold a = 0.05 in S2.
4. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, The distribution loss L mentioned in S2 dist The invariance loss L is the negative logarithmic probability of the feature in its character distribution. inv The spatial loss L is the KL divergence between the original feature distribution and the watermark feature distribution. space The Euclidean distance between the original feature and the watermark feature is given.
5. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, The triplet loss L described in S3 triplet The message loss L is calculated based on the feature distances between anchor points, positive samples, and negative samples. msg The cross-entropy loss, the redundancy loss L red The alignment loss L is calculated based on sub-region features. align The alignment of constrained text and formula features, the adversarial loss L adv Generated based on FGSM attack.
6. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, As stated in S3, w2 = (0.7 - 0.4λ), w4 = 0.1, and w5 = 0.1λ.
7. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, The encoder E described in S2 includes a multi-scale convolutional module, a spatial attention module, and a global feature generation layer. The multi-scale convolutional module contains four convolutional layers with a stride of 2, and the number of channels is 64, 128, 256, and 512 respectively. The spatial attention module receives the output of the last convolutional layer and generates channel- and spatially weighted features. The global feature generation layer outputs a 128-dimensional feature vector through global average pooling and a fully connected layer.
8. The method for embedding and extracting vector character watermarks based on multimodal features according to claim 1, characterized in that, The decoder D described in S3 includes: a message decoding branch containing two fully connected layers and a Sigmoid activation function, outputting bit probability values; a redundancy module that dynamically generates redundant codes based on the watermark strength parameter λ and fuses them into the message decoding branch; an adversarial branch containing three fully connected layers, outputting adversarial scores; a triplet optimization module that calculates the feature distance of watermark variants of the same character; and a cross-modal alignment module that fuses vector path features and raster image features and performs spatial constraints.
9. A vector character watermark embedding and extraction device based on multimodal features, employing the vector character watermark embedding and extraction method based on multimodal features as described in any one of claims 1-8, characterized in that, include: The dataset generation module D1 iterates through each character c in the character set C to generate the original glyph G0 corresponding to the original vector character, and the watermark variant G of the original glyph G0. w Adversarial sample glyphs G adv The character c, the original glyph G0, and the watermark variant G are labeled. w and adversarial sample glyphs G adv Metadata is used to build datasets; Encoder training module D2 calculates the feature distribution parameters {uc, ∑c} of the character c, where uc is the mean and ∑c is the covariance matrix, based on encoder E for the original glyph G0 and watermark variant G. w Extract features and calculate the distributed loss L. dist Invariance loss L inv and space loss L space Let the total loss function be minimized as L = L dist +0.5L inv +0.3L space When the total loss L is less than the preset threshold a, the parameters of the encoder E are frozen to obtain an encoder E1 with fixed parameters. Decoder training module D3 extracts the watermark variant G based on encoder E1. w Adversarial sample glyphs G adv The characteristics, combined with the triplet loss L triplet Message loss L msg Redundancy loss L red Alignment loss L align and combat losses L adv Multi-task learning is performed, and the loss weights are dynamically adjusted based on the watermark strength parameter λ. The decoder D is jointly trained, and only the parameters of the decoder D are updated until the total loss L is reached. total Less than a preset threshold or the message loss L msg , countering losses L adv To achieve equilibrium, the loss function L... total =w1L triplet +w2L msg +w3L red +w4L align +w5L adv ; The watermark extraction module D4 performs character detection on the document image from which the watermark is to be extracted, and extracts the character image set G = {G1, G2, ... G...} from the document image. i ,...,G n For each character image G, ,} i If a vector path exists, extract vector features; if it is a raster character, extract raster features; otherwise, fuse dual-modal features to generate multimodal features. Based on the watermark strength parameter λ, adaptively configure the marginal parameters and redundancy mode of the decoder D. Input the multimodal features into the majority decoder D to decode and obtain bit value b. Reconstruct all bit value sequences to reconstruct the watermark message.
10. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the vector character watermark embedding and extraction method based on multimodal features as described in any one of claims 1-8.
Citation Information
Cited By
Document file digital watermark generation method, extraction method and electronic device
CN122636390A