Digital watermark steganography method and system based on optical text compression

By using optical text compression and deep steganography, long texts are encoded into visual semantic vectors, which solves the problems of insufficient capacity and robustness in traditional steganography methods and achieves high-capacity, robust cross-modal information embedding and semantic-level recovery.

CN121961820APending Publication Date: 2026-05-01HANGZHOU LAOHE YUNQI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU LAOHE YUNQI INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-01-07
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies cannot effectively embed long text information into images, videos, and audio in a high-capacity and robust manner, and it is difficult to achieve reliable semantic-level recovery in multimodal environments. Traditional steganography methods lack capacity and robustness and are difficult to adapt to the compression and cropping of social platforms.

Method used

Optical text compression technology is used to encode long texts into visual semantic vectors, which are then embedded into images, videos, or audio through vector quantization and deep steganography. Combined with robustness enhancement and error correction mechanisms, cross-modal embedding and recovery of semantic vectors are achieved.

Benefits of technology

It achieves high-capacity long text information embedding, maintains a high recovery rate under the damage of JPEG compression, scaling and cropping, has strong cross-modal adaptability, and can directly recover semantic content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961820A_ABST
    Figure CN121961820A_ABST
Patent Text Reader

Abstract

The invention discloses a digital watermark steganography method and system based on optical text compression, and the method comprises the steps: 1, carrying out the load generation and optical coding, carrying out the optical rendering of a steganography text or image, inputting a visual semantic encoder, and obtaining a semantic vector set; step 2, performing semantic vector discretization and robustness enhancement; step 3, carrying out deep steganography embedding; 4, carrying out steganography extraction and error correction recovery; and 5, performing semantic vector reconstruction and optical decoding content generation. According to the method, cross-modal long text steganography based on visual semantic compression is realized for the first time, the limitation of a traditional steganography method in the aspects of capacity, recoverability and aggressiveness resistance is broken through, and the method can be widely applied to the fields of copyright identification, secret communication, digital traceability, media asset security and the like and has important engineering value and application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

A Digital Watermarking Steganography Method and System Based on Optical Text Compression Technical Field

[0001] This invention relates to the fields of digital content security and artificial intelligence multimodal coding, specifically to a digital watermarking steganography method and system based on optical text compression, applicable to high-capacity digital blind watermarking of images, videos and audio, covert communication, multimodal copyright identification, and secure transmission of long text. Background Technology

[0002] Currently, digital copyright protection has become a major challenge in the digital age. Digital steganography has gained widespread attention for its unique advantage of embedding copyright "blind watermarks" without affecting the sensory experience of the original work. However, traditional digital steganography methods still mainly rely on bit-level embedding methods (such as DCT coefficient modulation, LSB embedding, or spread spectrum technology), which generally limit their capacity and robustness. With the development of multimodal large-scale models (Visual-Language Models, VLMs) and optical text compression technology, it has become possible for models to represent large-scale text semantics with a very small number of visual vectors.

[0003] On the one hand, traditional steganography methods have limited capacity. When the steganographic information is long text (such as documents, credentials, or instruction sequences), the text needs to be encoded into a large bitstream, while a single frame of an image or a single video keyframe can only hold a very small amount of bits, making it difficult for traditional methods to handle content with high semantic density. On the other hand, robust steganography and cross-modal recovery are difficult. Existing steganography methods often can only recover the original bitstream and cannot recover the semantic text content, making it even more difficult to adapt to common disruptions such as compression, scaling, and cropping on social media platforms. With models such as DeepSeek-OCR proposing the optical encoding concept based on "Vision Tokens," long text can be compressed into tens to hundreds of semantic vectors. This provides new possibilities for semantic-level steganography, namely: long text can be recovered by embedding only a small set of visual semantic vectors, without needing to embed the original text or a QR code.

[0004] However, currently, there is no technical framework that fully integrates optical text compression, vector quantization, deep steganography, and optical decoding to form a systematic solution applicable to text steganography, image steganography, and video steganography. Existing technologies cannot optically encode text into high-density vectors, nor can they robustly embed floating-point semantic vectors into images / videos / audio. Their robustness and anti-interference capabilities are also relatively weak, making them easily identifiable and cracked. Therefore, there is an urgent need for a novel steganography method that can leverage the advantages of visual semantic compression to achieve high-capacity, robust, cross-modal embedding, and information recovery. Summary of the Invention

[0005] This invention proposes a digital watermarking steganography method and system based on optical text compression. It innovatively combines optical text encoding, large-scale vector quantization, deep steganography embedding, robust vector extraction, and optical decoding recovery to form a closed-loop steganography framework of "text-optical encoding-semantic vector-steganography embedding-extraction-optical decoding".

[0006] This invention can compress thousands of characters of text into hundreds of visual semantic vectors and embed them into a single frame of image, a single audio segment, or several video frames. After decoding, the original text content can be recovered. This invention breaks through the capacity bottleneck of traditional bit-based steganography, achieving semantic-level steganography. Furthermore, it maintains a high recovery rate even under the damage caused by JPEG compression, scaling, cropping, and video recoding. It successfully achieves high-capacity, highly robust, cross-modal steganographic embedding and semantic-level readable recovery of long text information.

[0007] The technical solution of the present invention is as follows: The present invention provides a digital watermarking steganography method based on optical text compression, comprising the following steps: Step 1, generating payload and optical encoding, performing optical rendering on steganographic text or image, and inputting it into a visual semantic encoder to obtain a set of semantic vectors; Step 2, discretizing and robustly enhancing semantic vectors; Step 3, performing deep steganography embedding; Step 4, performing steganography extraction and error correction recovery; Step 5, performing semantic vector reconstruction and optical decoding content generation.

[0008] Further, step 1 includes: step 1.1, receiving steganographic content P input by the user, the form of which includes one or more of the following: natural language long text, structured data, program code, mathematical formula, QR code image, and grayscale / color image containing semantic information; step 1.2, for text content, based on a preset layout template T... render Perform optical rendering, using template parameters including font type, font size, line spacing, paragraph indentation, page margins, and alignment, to generate a high-resolution optical payload image I. secret For image-based content, normalization and resolution adjustment are performed first; Step 1.3, the optical payload image I... secret The input is a pre-trained visual semantic encoder, which is built based on a visual language model and convolutional layers. It extracts global and local semantic features of the image through a multi-layer Transformer structure using a combination of global attention and window attention. The output is a semantic vector sequence V = {v1, v2, ..., v...} T}, where T is the number of visual semantic vectors, v i This represents the i-th semantic vector, where each v i ∈R d R is a d-dimensional dense vector representing high-level semantic information of text or images. d Let R represent a d-dimensional real vector space, where R represents the set of real numbers.

[0009] Further, step 2 includes: Step 2.1, quantizing and mapping the semantic vector set V based on the pre-trained vector quantization codebook, where Codebook = {e1, e2, ..., e...} K}, where K is the codebook size, which is obtained through K-means training, e j Let e ​​represent the j-th codebook vector. j ∈R d Step 2.2, for each semantic vector v i Calculate its Euclidean distance to all codebook vectors, and select the nearest neighbor codebook index: In the formula, z i It is an integer index that represents the position of the semantic vector v in the codebook. i The number of the closest codebook vector. The overall meaning is to retrieve the i-th semantic vector v. i With the j-th codebook vector e j The square of the Euclidean distance between them The index j of the minimum value is taken for each input semantic vector v. i Iterate through all vectors e1, e2, ..., e in the codebook K Calculate v i With each e j Find the distance with v i The nearest e j The index number j is assigned to z. i Thus, the index sequence Z = {z1, z2, ..., z} is obtained. T}, where T is the number of visual semantic vectors; Step 2.3, group the index sequence Z into groups of m indices and apply systematic error correction coding, using Reed-Solomon coding or BCH coding, to generate redundant check bits, resulting in an extended index sequence Z' with enhanced error resistance. The length of the extended index sequence Z' after coding is increased compared to the original index sequence Z.

[0010] Furthermore, step 2 also includes: step 2.4, encrypting or scrambling the extended index sequence Z'.

[0011] Further, step 3 includes: step 3.1, host media preparation: receiving carrier image I host 1. Standardize and align the video keyframe sequence or audio spectrogram; 2. Feature projection and spatial alignment: Map the one-dimensional index sequence Z' to a two-dimensional spatial feature map P using the learnable feature expansion network Expand(Z'). map ∈RH×W×C Where H and W are spatial dimensions, and C is the number of channels, this two-dimensional spatial feature map P map The spatial dimension and number of channels are matched with the feature dimensions of the host media; Step 3.3, multi-scale fusion embedding: construct a steganalysis embedding network H based on U-Net or residual dense network. embed By combining host media characteristics with P map Deep fusion is performed at multiple scales, preserving host structural information through skip connections and enhancing the adaptability of semantic embedding regions through attention mechanisms; Step 3.4, Imperceptible Constraints: During the training phase, the following loss function L is jointly optimized: , where L perceptual To mitigate perceived loss, visual / auditory quality is ensured based on VGG or LPIPS metrics; L extract To extract accuracy loss, supervise the recoverability of embedded information; L adversarial To combat the loss, a discriminator network is used to enhance the indistinguishability between the steganographic content and the original media. α, β, and γ are weighting coefficients, and I... stego Indicates classified media, I host This represents the received carrier image, and Z' is the encoded extended index sequence to be embedded. Extracting network H from encrypted media through steganography extract The initial index sequence is recovered; Step 3.5, robust enhancement training: in the steganalytic embedding network H embed During the training phase, it is combined with the steganalysis extraction network H extract End-to-end joint training is performed, and during the training process, a noise perturbation module N is introduced. noise For steganography embedding networks H embed The output encrypted media is subjected to various simulated attacks, including JPEG compression, Gaussian noise, random cropping, scaling, rotation, color dithering, video recoding, and audio AAC compression. The attacked encrypted media is then input into the steganography extraction network H. extract Calculate the extraction accuracy loss L extract .

[0012] Further, step 4 includes: Step 4.1, constructing a steganalysis extraction network H based on a convolutional neural network or a Vision Transformer. extract Input is the encrypted media that may be attacked. stego The output is the index prediction probability distribution p(z) for each position. i =k); Step 4.2, obtain the preliminary index sequence through maximum likelihood estimation or sequence decoding. Step 4.3, execute error correction decoding (ECC). -1 By using redundant check bits to detect and correct erroneous indexes, a corrected index sequence is obtained. Step 4.4: In multi-frame scenarios such as video or audio, multi-frame voting or temporal fusion strategies can be adopted to further improve the robustness of extraction.

[0013] Further, step 5 includes: step 5.1, reverse lookup using the codebook to obtain the index sequence. Map back to a set of semantic vectors: ,in, This represents the set of index sequences recovered after error correction and decoding. The index of the i-th recovered element is an integer whose range is determined by the codebook size K; Codebook[·] represents the vector retrieved from the codebook using the index as the key; This is represented as the semantic vector reconstructed at the i-th position. This represents the set of semantic vectors to be reconstructed; step 5.2, ... The input is a structure-based optical decoder (Decoder), which, together with the visual semantic encoder (Encoder) in step 1, forms a symmetrical structure for the pre-trained visual language model. This decoder can reconstruct the original optical image or directly generate text descriptions based on semantic vectors. Step 5.3: If the output is an optical image, post-processing is used to improve readability. Step 5.4: The final output is the reconstructed content. This completes the full recovery link from semantic vectors to readable information.

[0014] This invention also provides a digital watermarking steganography system based on optical text compression, comprising: a payload generation module for receiving steganographic content and generating a set of semantic vectors through text rendering and a visual semantic encoder; a vector quantization module for mapping the semantic vectors to an index sequence and performing error correction coding to generate an extended index sequence; a steganography embedding module for receiving a host image and an extended index sequence and generating a steganographic image; a steganography extraction module for recovering the predicted preliminary index sequence from the steganographic medium and performing error correction decoding; and an optical decoding module for dequantizing the index sequence recovered after error correction decoding and inputting it into an optical decoder for reconstruction, thereby generating the final reconstructed content.

[0015] The beneficial effects and features of this invention are as follows: 1. Extremely high steganography capacity. Long texts can be compressed into a small number of semantic vectors through optical encoding, which is tens of times higher than the capacity of traditional bit embedding.

[0016] 2. Semantic-level readable recovery. Instead of recovering the bitstream, the semantic vector is recovered, and the Decoder module restores the readable text or image.

[0017] 3. Strong robustness. Through "quantization + ECC + noise" modeling, the semantic content of the dense graph can still be recovered under high compression and high pruning conditions.

[0018] 4. Strong cross-modal adaptability. It can be embedded in images, video frames, and audio spectra to achieve multimedia steganography.

[0019] 5. Compatible with large model ecosystems. Existing OCR / VLM models can be used directly as encoders and decoders, without the need for additional text model training. Attached Figure Description

[0020] Figure 1: Overall flowchart of the method of the present invention.

[0021] Figure 2: Schematic diagram of semantic vector quantization and index generation structure.

[0022] Figure 3: Steganographic Embedding Network H embed (U-Net structure) schematic diagram.

[0023] Figure 4: Steganography extraction network H extract Structural diagram.

[0024] Figure 5: Schematic diagram of optical decoding recovery process.

[0025] Figure 6: Schematic diagram of video frame multi-point embedding scheme. Detailed Implementation

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that those skilled in the art can replace or adjust the modules without departing from the spirit of the present invention, and all such replacements or adjustments fall within the scope of protection of the present invention.

[0027] Please refer to Figure 1. A digital watermarking steganography method based on optical text compression includes the following steps: Step 1, generating payload and optical encoding, performing optical rendering on the steganographic text or image, and inputting it into a visual semantic encoder to obtain a set of semantic vectors, including: Step 1.1, receiving steganographic content P input by the user, the form of which includes but is not limited to long natural language text, structured data, program code, mathematical formula, QR code image or grayscale / color image containing semantic information.

[0028] Step 1.2: For text-based content, based on the preset layout template T render Perform optical rendering, with template parameters including font type, font size, line spacing, paragraph indentation, page margins, alignment, etc., to generate a high-resolution optical payload image I. secret This ensures that the text is visually neat and semantically recognizable.

[0029] For image content, normalization and resolution adjustment are performed first to ensure that the input size is compatible with the encoder.

[0030] Step 1.3, transfer the optical payload image I secretThe input is a pre-trained visual semantic encoder, which is built based on a visual language model (SAM VITDET, CLIP VIT, etc.) and convolutional layers. It extracts global and local semantic features of the image through a multi-layer Transformer structure using a combination of global attention and window attention. The output is a semantic vector sequence V = {v1, v2, ..., v...} T}, where v i Let v represent the i-th semantic vector, where T is the number of visual semantic vectors, and each v i ∈R d For a d-dimensional dense vector (R) d A vector space of d-dimensional real numbers (where R represents the set of real numbers) represents high-level semantic information of text or images.

[0031] Step 2: Perform semantic vector discretization and robustness enhancement. A schematic diagram of the semantic vector quantization and index generation structure is shown in Figure 2. The specific process includes: Step 2.1: Based on the pre-trained vector quantization codebook, perform quantization mapping on the semantic vector set V, where Codebook = {e1, e2, ..., e...} K}, where K is the codebook size, which can be obtained through K-means training. The codebook size K is typically set between 256 and 65536. j Let e ​​represent the j-th codebook vector. j ∈R d .

[0032] Step 2.2, for each semantic vector v i Calculate its Euclidean distance to all codebook vectors, and select the nearest neighbor codebook index: In the formula, z i It is an integer index that represents the position of the semantic vector v in the codebook. i The number of the closest codebook vector. The overall meaning is to retrieve the i-th semantic vector v. i With the j-th codebook vector e j The square of the Euclidean distance between them The index j of the minimum value is taken. For each input semantic vector v i Iterate through all vectors e1, e2, ..., e in the codebook K Calculate v i With each e j Find the distance with v i The nearest e j The index number j is assigned to z. i Thus, the index sequence Z = {z1, z2, ..., z} is obtained. T}, where T is the number of visual semantic vectors.

[0033] Step 2.3: To improve anti-interference capability, the index sequence Z is grouped into groups of m indices, and systematic error correction coding (ECC) is applied using Reed-Solomon coding or BCH coding to generate redundant check bits, resulting in an extended index sequence Z' with enhanced error resistance. The length of the encoded sequence Z' is increased compared to the original index sequence Z.

[0034] Step 2.4 involves encrypting or scrambling Z' to enhance the security of the steganographic content. Step 2.4 is optional.

[0035] Step 3, perform deep steganography embedding, including: Step 3.1, host media preparation: receive carrier image I host Standardize and align the video keyframe sequences or audio spectrograms to their dimensions.

[0036] Step 3.2, Feature Projection and Spatial Alignment: The one-dimensional index sequence Z' is mapped to a two-dimensional spatial feature map P using the learnable feature expansion network Expand(Z'). map ∈R H×W×C Where H and W are spatial dimensions, and C is the number of channels, this two-dimensional spatial feature map P map The spatial dimensions and number of channels are matched with the feature dimensions of the host media in order to facilitate subsequent fusion.

[0037] Step 3.3, Multi-scale fusion embedding: Construct a steganalysis embedding network H based on U-Net or residual dense network. embed Steganography Embedded Network H embed The structural diagram is shown in Figure 3, which combines the host media characteristics with P. map Deep fusion is performed at multiple scales, host structure information is preserved through skip connections, and the adaptability of semantic embedding regions is enhanced through attention mechanisms.

[0038] Step 3.4, Imperceptibility Constraints: During the training phase, the following loss function L is jointly optimized:

[0039] Among them, L perceptual To mitigate perceived loss, visual / auditory quality is ensured based on VGG or LPIPS metrics; L extract To extract accuracy loss, supervise the recoverability of embedded information; L adversarial To combat the loss, a discriminator network is used to enhance the indistinguishability between the steganographic content and the original media. α, β, and γ are weighting coefficients, and I... stego Indicates classified media, I hostThis represents the received carrier image, and Z' is the encoded extended index sequence to be embedded. Extracting network H from encrypted media through steganography extract The initial index sequence has been recovered.

[0040] Step 3.5, Robustness Enhancement Training: In the Steg embedding network H embed During the training phase, it is combined with the steganalysis extraction network H extract (See step 4) Perform end-to-end joint training. During training, introduce a noise perturbation module N. noise For steganography embedding networks H embed The output encrypted media is subjected to various simulated attacks, including JPEG compression (quality factor q∈[30,95]), Gaussian noise, random cropping, scaling, rotation, color dithering, video recoding (H.264 / HEVC), and audio AAC compression. The attacked media is then input into the steganography extraction network H. extract Calculate the extraction accuracy loss L extract This training mechanism enables the steganography embedding network H... embed It can learn and generate steganalytic features with strong anti-interference capabilities, thereby improving the robustness of the entire method.

[0041] Step 4, perform steganalysis and error correction, including: Step 4.1, construct a steganalysis network H based on a convolutional neural network (CNN) or Vision Transformer. extract Steganography extraction network H extract The structural diagram is shown in Figure 4. The input is the encrypted media I that may be attacked. stego The output is the index prediction probability distribution p(z) for each position. i =k).

[0042] Step 4.2: Obtain the preliminary index sequence through maximum likelihood estimation or sequence decoding (such as Beam Search). .

[0043] Step 4.3, perform error correction decoding (ECC). -1 By using redundant check bits to detect and correct erroneous indexes, a corrected index sequence is obtained. .

[0044] Step 4.4: In multi-frame scenarios such as video or audio, multi-frame voting or temporal fusion strategies can be adopted to further improve the robustness of extraction.

[0045] Step 5, perform semantic vector reconstruction and optical decoding content generation, including: Step 5.1, reverse lookup using the codebook to generate the index sequence. Map back to a set of semantic vectors: ,in, This represents the set of index sequences recovered after error correction and decoding. The index of the i-th recovered element is an integer whose range is determined by the codebook size K; Codebook[·] represents the vector retrieved from the codebook using the index as the key; This is represented as the semantic vector reconstructed at the i-th position. Represents the set of semantic vectors for reconstruction.

[0046] Step 5.2, will The input is an optical decoder based on the MoE (Mixture of Experts) structure. This optical decoder, together with the visual semantic encoder in step 1, forms a symmetrical structure for the pre-trained visual language model, which can reconstruct the original optical image or directly generate text descriptions based on semantic vectors.

[0047] Step 5.3: If the output is an optical image, readability can be improved through post-processing (binarization, super-resolution reconstruction, etc.).

[0048] Step 5.4, final output of reconstructed content It can be text, images, or a combination of multiple modalities to complete the entire recovery chain from semantic vectors to readable information. A schematic diagram of the optical decoding recovery process is shown in Figure 5.

[0049] Step 6, implement the configurable and scalable application mechanism of the method, including: Step 6.1, support end-to-end training and modular deployment, each sub-module can be optimized independently or fine-tuned jointly.

[0050] Step 6.2 provides a user-configurable parameter interface, including codebook size K, vector dimension D, embedding strength λ, error correction redundancy rate R, etc.

[0051] Step 6.3 supports batch processing and streaming embedding, suitable for real-time steganography needs of image libraries, video streams, and audio clips.

[0052] Step 6.4 provides a steganalysis adversarial module, which can further improve the steganalysis embedding network H through adversarial training. embed The generated encrypted media has enhanced resistance to detection, thus improving the security of this method under steganography detection attacks.

[0053] The present invention also provides a digital watermarking steganography system based on optical text compression, comprising: a payload generation module for receiving steganographic content P and generating a semantic vector set V through text rendering and a visual semantic encoder, corresponding to step 1 above.

[0054] Vector quantization module: used to map semantic vector V to index sequence Z and perform error correction coding to generate extended index sequence Z', corresponding to step 2 above.

[0055] Steganography embedding module: used to receive host image I host And expand the index sequence Z', and generate the encrypted image I. stego This corresponds to step 3 above.

[0056] Steganography extraction module: This module is used to extract steganography data from stony images. stego Preliminary index sequence for recovery prediction Then perform error correction decoding, corresponding to step 4 above.

[0057] Optical decoding module: used to recover the index sequence after error correction decoding. The data is dequantized and fed into the optical decoder for reconstruction, thus generating the final reconstructed content. This corresponds to step 5 above.

[0058] Example 1: Semantic Vector Steganography Based on a Single Image This example uses a text of approximately 800 characters as the steganographic content P, and performs the following steps: Step 1: Payload Generation and Optical Text Encoding 1) Text Rendering. Using a uniform font, line spacing, and white background with black text, the text is rendered as a 1024×1024 optical image I. secret . That is: I secret =Render(P).

[0059] 2) Optical Encoder. This involves encoding the I... secret Input the Encoder to generate T=64 visual semantic vectors: Each vector has a dimension of 1024, and the overall information content is about 1 / 80 of the original text.

[0060] Step 2: Semantic Vector Quantization and Error Correction Coding. A codebook of size K=4096 is used, trained using K-means to obtain the codebook. Quantization Mapping: The resulting index sequence Z has a length of 64. To enhance robustness, Reed-Solomon error correction coding is used: The ECC extension ratio is 1.5 times, and the final embedded payload length is approximately 96 index units.

[0061] Step 3: Training the Steg embedding network and generating the stegmap. The feature projection network Expand(Z') maps the index sequence into a 64×64×32 spatial feature map P. map H, a steganography embedding network with a U-Net structure embed , host image I host With P mapThe input U-Net is concatenated, and after multi-layer convolutional feature fusion, the output is a dense image I. stego The loss function is optimized as follows (where α is set to 0.8, β to 1.2, and γ to an initial value of 0.1): During training, a noise perturbation layer is added: JPEG (q∈[30,95]), random cropping, rotation, blurring, etc. The final output is a high-density image. I cannot be distinguished by the naked eye stego Compared with the original image I host (LPIPS<0.015).

[0062] Step 4: After uploading the steganographically extracted image to a social media platform and downloading it, under heavy compression: input the extraction network H... extract : After error correction and decoding, the following is obtained: The recovery rate can still reach 97% under common JPEG 40 compression, with an average recovery rate of over 95%.

[0063] Step 5: Optical decoding is obtained from indexed inverse quantization. : The final text is obtained by the decoder. : The edit distance is less than 3% of the original text, making the content highly readable.

[0064] Example 2: Distributed Steganography Based on Video Keyframes. Building upon the single-image steganography method described in Example 1, this example applies it to video carriers. By distributively embedding the steganographic payload into multiple video keyframes, robustness against complex attacks such as video compression and frame loss is further enhanced. The main difference between this example and Example 1 lies in the specific implementation of steps 3 (steganographic embedding) and 4 (steganographic extraction). The remaining steps (payload generation, optical encoding, vector quantization, error correction encoding, and optical decoding) are the same as in Example 1.

[0065] Steps 1 and 2 are exactly the same as in Example 1. The steganographic content P is received, and after rendering, optical coding, vector quantization, and error correction coding, an enhanced index sequence Z' of length 96 is obtained.

[0066] Step 3: Distributed Steganography and Embedding. The index sequence Z' to be embedded is uniformly divided into N=8 parts, denoted as: Each sub-load It contains 12 index units (corresponding to the original 8 semantic vectors and their redundant check bits).

[0067] Eight visually complex and evenly distributed I-frames or keyframes are selected from the host video as carriers, denoted as: For each pair (sub-loads) Host keyframe Step 3 of Example 1 is executed independently. That is, the same feature projection network Expand(·) and steganalysis embedding network H are used. embed Each sub-payload is embedded into the corresponding keyframe, generating 8 payload-detailed keyframes. .

[0068] The keyframes carrying the encryption are replaced back with the original video sequence, and the non-keyframes are encoded in a regular manner to generate the final encrypted video.

[0069] Step 4: Distributed Steg extraction. From the encrypted video that may have undergone re-encoding, resolution adjustment, or other processing, re-detect and extract the corresponding 8 keyframes. .

[0070] For each keyframe carrying encryption Step 4 of Example 1 is executed independently. That is, the same steganalysis network H is used. extract Error correction decoding (ECC) -1 Eight recovered sub-index sequences were obtained. .

[0071] To correct individual frame extraction errors, a majority voting strategy is used to make fusion decisions on the final index sequence. Specifically, for each index position j: ,in, It is the j-th index extracted from the i-th frame. For the indicator function, argmax k The parentheses (·) indicate that the value of the independent variable k that maximizes the subsequent expression is returned, thus achieving majority voting. This yields the robustly enhanced final index sequence. .

[0072] When a frame is severely damaged and cannot be extracted, it can be recovered by relying on the redundant information of the remaining frames through error correction coding itself.

[0073] Step 5: This is exactly the same as step 5 in Example 1. The merged index sequence... Perform codebook inverse quantization and optical decoding to recover the original steganographic content. .

[0074] More than 95% of the semantic vectors can still be recovered after video recoding to H.264 (CRF=32).

[0075] Example 3: Audio Spectrum Domain Steganography. Based on the method described in Example 1, this example extends the steganography carrier to audio signals. By embedding semantic vectors in the time-frequency domain (spectrum) of the audio signal, cross-modal steganographic communication is achieved. The core modification of this example lies in the preprocessing of the audio carrier, the embedding position, and the adaptation of the network structure in steps 3 (steganography embedding) and 4 (steganography extraction). Steps 1, 2, and 5 remain consistent with Example 1.

[0076] Steps 1 and 2: Exactly the same as in Example 1. Generate the enhanced index sequence Z' to be embedded.

[0077] Step 3: Audio Spectrum Domain Steganographic Embedding of Host Audio Signal A host The process involves framing, windowing, and converting the data into a time-spectrum image S using a short-time Fourier transform (STFT). host Further converting it into a Mel spectrogram To better match the characteristics of human hearing perception. Among them, Let F be the set of real numbers, F be a Mel-band number, and T be a real number. a This represents the number of time frames.

[0078] Based on psychoacoustic theory, calculate M host The auditory masking threshold of each time frequency unit is determined. Regions with higher masking thresholds (i.e., "masking regions") are selected as embedding carriers to ensure that the embedded signal is not easily detected.

[0079] The index sequence Z' is passed through an adapted projection network Expand. audio (Z') is mapped to M host Dimension-compatible feature maps .

[0080] A lightweight convolutional neural network (CNN) is used as the audio steganography embedding network. M host and Feature-level fusion is performed to generate a Mimel spectrum M. stego The training loss function is similar to that in Example 1, but L perceptual It needs to be replaced with an audio quality-based metric (such as the difference between PESQ or SI-SNR).

[0081] For M stego Perform inverse Mel-Transform and inverse STFT to reconstruct the time-domain signal, generating the final encrypted audio A. stego .

[0082] Step 4: Audio Steganography Extraction This involves extracting the received, possibly compressed, steganographic audio data. Perform the same preprocessing as in step 3 to obtain the Mel spectrum to be analyzed. .

[0083] A CNN network symmetrical to the embedding network is used as the audio steganalysis network. ,from Extract the predicted index sequence .

[0084] Similar to Example 1, for Perform error correction decoding (ECC) -1 The recovered index sequence is obtained. .

[0085] Step 5: This is exactly the same as step 5 in Example 1. For the index sequence... Perform codebook inverse quantization and optical decoding to recover the original steganographic content. .

[0086] Secret Audio A stego Compared to the original audio A host Its sound quality degradation is minimal (PESQ decrease of less than 0.1, or perceived signal-to-noise ratio change of less than 0.4dB). The semantic vector can still be recovered after AAC compression, and the recovery rate of the semantic vector after AAC compression (128Kbps) remains above 91%.

[0087] Through the above multiple embodiments, the present invention achieves the following key technical capabilities: (1) Significantly improves the steganography capacity. After semantic vector compression, a single image can carry thousands of words of text, while traditional bit steganography is less than 1 / 10.

[0088] (2) High robustness: Through semantic vector quantization (discrete robustness), ECC error correction (anti-noise), etc., it can still recover semantic content under high damage and maintain stable recovery ability under multiple types of damage.

[0089] (3) It has strong cross-modal adaptability, and can use images, videos and audio as carriers, making it suitable for a variety of situations.

[0090] (4) The semantic recovery advantage brought by optical decoding: Compared with traditional steganography methods that can only recover bits, this invention can directly recover long text, formulas, tables and even image descriptions.

[0091] The comparison between this scheme and traditional steganography schemes is shown in Table 1 below: Table 1 Comparison between traditional steganography schemes and this scheme

[0092] The method described in this embodiment achieves significantly better performance than traditional blind watermarking and bit steganography methods on three types of datasets: images, videos, and audio. In particular, the improvement in semantic recovery accuracy is most significant under conditions of high compression, high noise, and recoding, effectively verifying the robustness and reliability of the optical semantic vector steganography framework in multimedia environments.

[0093] The proposed digital watermarking steganography method based on optical text compression achieves high-capacity embedding and high semantic fidelity recovery of long text information through an integrated framework of "optical encoding—vector quantization—deep embedding—robust extraction—optical decoding". The overall process is shown in Figure 1, including the following core steps: optical rendering, encoder encoding, vector quantization and error correction, deep steganography embedding, steganography extraction, and optical decoding.

[0094] This invention demonstrates significant advantages in various real-world media processing scenarios (social media compression, image editing, video recoding, and audio compression), particularly excelling in semantic-level steganography capacity, visual imperceptibility, and cross-modal robustness. It provides a key technological breakthrough for fields such as digital watermarking, content traceability, copyright identification, and covert communication, and has broad engineering application prospects in multimedia security and AI-driven content protection.

Claims

1. A digital watermarking steganography method based on optical text compression, characterized in that, Includes the following steps: Step 1: Perform payload generation and optical encoding, perform optical rendering on the steganographic text or image, and input it into the visual semantic encoder to obtain a set of semantic vectors; Step 2: Perform semantic vector discretization and robust enhancement; Step 3: Perform deep steganographic embedding; Step 4: Perform steganographic extraction and error correction recovery; Step 5: Perform semantic vector reconstruction and optical decoding content generation.

2. The digital watermarking steganography method based on optical text compression according to claim 1, characterized in that, Step 1 includes: Step 1.1, receiving steganographic content P input by the user, which may be in the form of one or more of the following: natural language long text, structured data, program code, mathematical formula, QR code image, and grayscale / color image containing semantic information; Step 1.2, for text content, applying a preset layout template T... render Perform optical rendering, using template parameters including font type, font size, line spacing, paragraph indentation, page margins, and alignment, to generate a high-resolution optical payload image I. secret For image-based content, normalization and resolution adjustment are performed first; Step 1.3, the optical payload image I... secret The input is a pre-trained visual semantic encoder, which is built based on a visual language model and convolutional layers. It extracts global and local semantic features of the image through a multi-layer Transformer structure using a combination of global attention and window attention. The output is a semantic vector sequence V = {v1, v2, ..., v...} T }, where T is the number of visual semantic vectors, v i This represents the i-th semantic vector, where each v i ∈R d R is a d-dimensional dense vector representing high-level semantic information of text or images. d Let R represent a d-dimensional real vector space, where R represents the set of real numbers.

3. The digital watermarking steganography method based on optical text compression according to claim 1, characterized in that, Step 2 includes: Step 2.1, quantizing and mapping the semantic vector set V based on the pre-trained vector quantization codebook, where Codebook = {e1, e2, ..., e...} K }, where K is the codebook size, which is obtained through K-means training, e j Let e ​​represent the j-th codebook vector. j ∈R d Step 2.2, for each semantic vector v i Calculate its Euclidean distance to all codebook vectors, and select the nearest neighbor codebook index: In the formula, z i It is an integer index that represents the position of the semantic vector v in the codebook. i The number of the closest codebook vector. The overall meaning is to retrieve the i-th semantic vector v. i With the j-th codebook vector e j The square of the Euclidean distance between them The index j of the minimum value is taken for each input semantic vector v. i Iterate through all vectors e1, e2, ..., e in the codebook K Calculate v i With each e j Find the distance with v i The nearest e j The index number j is assigned to z. i Thus, the index sequence Z = {z1, z2, ..., z} is obtained. T }, where T is the number of visual semantic vectors; Step 2.3, group the index sequence Z into groups of m indices and apply systematic error correction coding, using Reed-Solomon coding or BCH coding, to generate redundant check bits, resulting in an extended index sequence Z' with enhanced error resistance. The length of the extended index sequence Z' after coding is increased compared to the original index sequence Z.

4. The digital watermarking steganography method based on optical text compression according to claim 3, characterized in that, Step 2 further includes: Step 2.4, encrypting or scrambling the extended index sequence Z'.

5. The digital watermarking steganography method based on optical text compression according to claim 1, characterized in that, Step 3 includes: Step 3.1, Host media preparation: Receiving carrier image I host 1. Standardize and align the video keyframe sequence or audio spectrogram; 2. Feature projection and spatial alignment: Map the one-dimensional index sequence Z' to a two-dimensional spatial feature map P using the learnable feature expansion network Expand(Z'). map ∈R H×W×C Where H and W are spatial dimensions, and C is the number of channels, this two-dimensional spatial feature map P map The spatial dimension and number of channels are matched with the feature dimensions of the host media; Step 3.3, multi-scale fusion embedding: construct a steganalysis embedding network H based on U-Net or residual dense network. embed By combining host media characteristics with P map Deep fusion is performed at multiple scales, preserving host structural information through skip connections and enhancing the adaptability of semantic embedding regions through attention mechanisms; Step 3.4, Imperceptible Constraints: During the training phase, the following loss function L is jointly optimized: , where L perceptual To mitigate perceived loss, visual / auditory quality is ensured based on VGG or LPIPS metrics; L extract To extract accuracy loss, supervise the recoverability of embedded information; L adversarial To combat the loss, a discriminator network is used to enhance the indistinguishability between the steganographic content and the original media. α, β, and γ are weighting coefficients, and I... stego Indicates classified media, I host This represents the received carrier image, and Z' is the encoded extended index sequence to be embedded. Extracting network H from encrypted media through steganography extract The initial index sequence is recovered; Step 3.5, robust enhancement training: in the steganalytic embedding network H embed During the training phase, it is combined with the steganalysis extraction network H extract End-to-end joint training is performed, and during the training process, a noise perturbation module N is introduced. noise For steganography embedding networks H embed The output encrypted media is subjected to various simulated attacks, including JPEG compression, Gaussian noise, random cropping, scaling, rotation, color dithering, video recoding, and audio AAC compression. The attacked encrypted media is then input into the steganography extraction network H. extract Calculate the extraction accuracy loss L extract .

6. The digital watermarking steganography method based on optical text compression according to claim 1, characterized in that, Step 4 includes: Step 4.1, constructing a steganalysis extraction network H based on a convolutional neural network or Vision Transformer. extract Input is the encrypted media that may be attacked. stego The output is the index prediction probability distribution p(z) for each position. i =k); Step 4.2, obtain the preliminary index sequence through maximum likelihood estimation or sequence decoding. Step 4.3, execute error correction decoding (ECC). -1 By using redundant check bits to detect and correct erroneous indexes, a corrected index sequence is obtained. Step 4.4: In multi-frame scenarios such as video or audio, multi-frame voting or temporal fusion strategies can be adopted to further improve the robustness of extraction.

7. The digital watermarking steganography method based on optical text compression according to claim 2, characterized in that, Step 5 includes: Step 5.1, reverse lookup using the codebook to obtain the index sequence. Map back to a set of semantic vectors: ,in, This represents the set of index sequences recovered after error correction and decoding. The index of the i-th recovered element is an integer whose range is determined by the codebook size K; Codebook[·] represents the vector retrieved from the codebook using the index as the key; This is represented as the semantic vector reconstructed at the i-th position. This represents the set of semantic vectors to be reconstructed; step 5.2, ... The input is a structure-based optical decoder (Decoder), which, together with the visual semantic encoder (Encoder) in step 1, forms a symmetrical structure for the pre-trained visual language model. This decoder can reconstruct the original optical image or directly generate text descriptions based on semantic vectors. Step 5.3: If the output is an optical image, post-processing is used to improve readability. Step 5.4: The final output is the reconstructed content. This completes the full recovery link from semantic vectors to readable information.

8. A digital watermarking steganography system based on optical text compression, characterized in that, include: The payload generation module receives steganographic content and generates a set of semantic vectors through text rendering and a visual semantic encoder. The vector quantization module maps the semantic vectors to an index sequence and performs error correction coding to generate an extended index sequence. The steganography embedding module receives the host image and the extended index sequence and generates a stegated image. The steganography extraction module recovers the predicted preliminary index sequence from the stegated medium and performs error correction decoding. The optical decoding module dequantizes the index sequence recovered after error correction decoding and inputs it into the optical decoder for reconstruction, thereby generating the final reconstructed content.