Adaptive multi-level digital watermarking method based on coding process optimization

An adaptive multi-level digital watermarking method based on multimodal intelligent analysis and deep coding collaboration solves the problems of adaptability, coding collaboration, and robustness of watermarking technology in ultra-high-definition video, achieving efficient copyright protection and ensuring stable extraction and concealment of watermarks under compression, transcoding, and attack scenarios.

CN121940550AActive Publication Date: 2026-04-28HANGZHOU BAOMIHUA TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU BAOMIHUA TECH CO LTD
Filing Date
2026-03-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing digital watermarking technologies struggle to achieve a balance between adaptability, coding coordinability, robustness, and concealment in ultra-high-definition videos, resulting in poor copyright protection. In particular, watermarks are prone to distortion or are difficult to extract in video compression, transcoding, and attack scenarios.

Method used

An adaptive multi-level digital watermarking method that combines multimodal intelligent analysis with deep coding collaboration is adopted. High-dimensional feature vectors are generated through deep analysis of multimodal content. Combined with a reinforcement learning-driven dynamic decision model, the watermark strength, position and type are adaptively adjusted. Furthermore, it is deeply collaborative with video coding to achieve multi-level watermark embedding and extraction.

Benefits of technology

It significantly enhances the survivability of watermarks under various attack and signal processing scenarios, balances robustness and concealment, ensures the accurate extraction and integrity of copyright information, and does not affect the visual quality of the video, adapting to different scenarios with different content values ​​and protection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940550A_ABST
    Figure CN121940550A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive multi-level digital watermarking method based on coding process optimization. The method comprises the following steps: firstly, performing deep preprocessing on an input video stream through a pre-trained multi-modal deep learning model, extracting space complexity, time dynamics and content significance features, and generating a uniform feature vector; and constructing a dynamic decision model based on the feature vectors, adaptively determining the watermark intensity, the embedding position and the type, and realizing accurate matching of watermark parameters and video content characteristics. A multi-level watermark hierarchical embedding mechanism is adopted, a robust invisible watermark, a fragile invisible watermark and a dynamic visible watermark are embedded into a video, and the security is enhanced by combining chaotic encryption or Hash chain encryption. During copyright verification, watermark information of each layer is recovered through an adaptive extraction algorithm, and data verification and infringement traceability are completed in combination with a block chain evidence storage system. According to the method, a full-process copyright protection system of embedding, coding, extraction and verification is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital media content security technology, specifically to an adaptive multi-level digital watermarking method based on encoding process optimization. Background Technology

[0002] With the rapid development of digital media technology, the production and dissemination of ultra-high-definition video content such as 4K and 8K are becoming increasingly widespread. While the convenient circulation of digital content in the internet environment has promoted the development of the media industry, it has also brought serious copyright infringement problems. Unauthorized copying, alteration, and dissemination are frequent occurrences, urgently requiring reliable copyright protection technologies to safeguard the rights of content creators. Digital watermarking technology, as a core means of current digital content copyright protection, has become a key solution to address infringement issues due to its ability to embed hidden information or explicit markings into content.

[0003] Current mainstream digital watermarking technologies are mainly divided into two categories: visible watermarks and invisible watermarks. Visible watermarks, by overlaying logos, text, or copyright notices onto the video frame to explicitly declare copyright ownership, can have a direct deterrent effect, but they significantly impact the visual viewing experience and are easily destroyed through simple operations such as cropping, blurring, and overlaying, thus offering limited protection. Invisible watermarks, on the other hand, embed core data such as copyright ID and creator information hidden within the video data, possessing excellent concealment and avoiding impact on the viewing experience. However, their core technological challenge lies in robustness—that is, whether the watermark can still be accurately extracted after undergoing common signal processing or malicious attacks such as video compression, transcoding, noise interference, and partial cropping. Existing digital watermarking technologies still have many shortcomings in practical applications and are insufficient to meet the copyright protection needs of the ultra-high-definition video era. The main problems include:

[0004] (1) Existing adaptive watermarking technologies mostly adopt fixed adjustment strategies based on simple rules. They can only adjust watermark parameters according to single or a few features such as video brightness and basic texture, which cannot cope with the high complexity and dynamic changes of video content. For example, when faced with action videos with frequent scene changes or static and flat landscape videos, uniform parameter adjustment rules are difficult to take into account the stability of watermarks and visual quality in different scenes. In complex scenes, watermarks are often lost due to insufficient anti-interference ability, or in static scenes, the image quality is damaged due to excessive watermark intensity.

[0005] (2) Currently, most watermark embedding processes are independent of the video compression coding process and are out of touch with mainstream video coding standards. Video compression is essentially a lossy processing process, which will significantly damage the embedded watermark information. In particular, invisible watermarks often rely on video frequency domain coefficients or pixel bit information. Quantization, entropy coding and other operations during the compression process can easily lead to the distortion of these key data carrying the watermark, ultimately making it difficult to accurately extract the watermark after compression, which seriously affects the effectiveness of copyright tracking and evidence collection.

[0006] (3) Robustness and invisibility are the core contradictions of digital watermarking technology. To improve the watermark's resistance to attacks, it is usually necessary to increase the watermark embedding strength, but this will cause visually perceptible distortion in the video and sacrifice visual quality. If the watermark strength is reduced to ensure image quality, the watermark's anti-interference ability will be greatly reduced, making it difficult to resist conventional processing such as compression and transcoding. Existing technologies lack an effective dynamic balancing mechanism and cannot achieve optimal adaptation between the two.

[0007] (4) Most existing solutions use only a single type of watermark for protection, such as relying solely on invisible watermarks for copyright tracking or relying solely on visible watermarks for deterrence. This single-layer protection system is weak in its ability to withstand risks. Once it encounters a targeted attack, the entire copyright protection mechanism becomes completely ineffective and cannot meet the differentiated scenarios with different content values ​​and different protection needs.

[0008] In summary, existing digital watermarking technologies have significant shortcomings in these aspects, necessitating a new digital watermarking method that can be deeply integrated with modern video coding processes and driven by intelligent algorithms. Summary of the Invention

[0009] To address the shortcomings of existing digital watermarking technologies in terms of adaptability, coding collaboration, balance between robustness and concealment, and integrity of protection layers, this invention proposes an adaptive multi-layered digital watermarking method based on coding process optimization through innovative designs such as multimodal intelligent analysis, deep coding collaboration, and a multi-layered watermarking architecture. The method includes the following steps:

[0010] S1: Perform multimodal deep content analysis on the input video stream, extract spatial complexity, temporal dynamics, visual saliency and content security risk assessment features, and generate a unified high-dimensional feature vector through standardization and weighted fusion;

[0011] S2: Based on the high-dimensional feature vector, the watermark strength, embedding position, and type combination are adaptively determined through a reinforcement learning-driven dynamic decision-making model;

[0012] S3: Perform multi-level watermarking hierarchical embedding according to the decision results, and adopt corresponding embedding strategies for robust invisible watermarks, fragile invisible watermarks, semi-fragile invisible watermarks and dynamic visible watermarks respectively, and combine encryption technology to enhance security.

[0013] S4: The watermark embedding process is deeply integrated with video coding, and a differentiated optimization strategy is adopted in the quantization, prediction and entropy coding stages, while achieving dynamic bitrate control and edge node load balancing.

[0014] S5: Performs edge decoding preprocessing on the video stream to be verified, adaptively extracts watermark information from each layer, and combines the consortium blockchain to complete integrity verification, legality verification, and infringement tracing.

[0015] Preferably, S1 includes:

[0016] S11: Perform frame-level splitting on the video stream and dynamically adjust the window size and number of frames based on the scene switching frequency using an elastic window grouping strategy.

[0017] S12: Input the labeled frame group into the multi-branch enhanced multimodal deep learning model, which includes a Transformer branch, an attention mechanism 3D convolution branch, an improved Gaussian mixture model saliency detection subnetwork, and a content security risk assessment subnetwork;

[0018] S13: Generate a spatial complexity heatmap with the same resolution as the original frame by adaptive pyramid downsampling and core spatial parameter calculation;

[0019] S14: Integrate optical flow method with motion trajectory prediction model to extract temporal dynamic features and generate scoring matrix and motion stability index;

[0020] S15: Detect visually salient regions based on the characteristics of the human visual system and generate a salient region mask map;

[0021] S16: Combining content value, dissemination scenarios, and historical infringement data, a weighted formula is used to determine the content security risk level;

[0022] S17: Standard scores are used to process various features, and high-dimensional feature vectors are generated by weighted fusion.

[0023] Preferably, S2 includes:

[0024] S21: Construct a dynamic decision model for a deep reinforcement learning framework, using watermark anti-attack capability, video visual quality, and transmission efficiency as the joint reward function, integrating multi-source data support, and establishing an autonomous learning mechanism;

[0025] S22: Set the basic threshold range for watermark strength and perform real-time correction based on visual perception characteristics, content features and risk level;

[0026] S23: The embedding position is selected based on the saliency mask map and macroblock priority score. When the proportion of non-sensitive areas is insufficient, a complementary strategy of sparse embedding and macroblock grouping is adopted.

[0027] S24: Construct a multi-dimensional evaluation system for content value, security requirements, and transmission costs, and match the corresponding watermark type combination mode based on the evaluation results.

[0028] Preferably, S3 includes:

[0029] S31: Robust invisible watermark is embedded in low-frequency coefficients of video frames. The core copyright information is encrypted using a lattice-based homomorphic encryption algorithm. The embedding is achieved by modifying the coefficients through a quantization error minimization algorithm.

[0030] S32: The fragile invisible watermark uses an improved adaptive LSB algorithm to embed high-frequency coefficients, integrates frame hash values, timestamps and other information, and employs dual protection of hash chain encryption and digital signature;

[0031] S33: Semi-fragile invisible watermarks are embedded in the mid-frequency coefficients and the second-lowest pixel bit, and the tolerance and sensitivity are balanced by adaptive threshold and neighborhood mean compensation.

[0032] S34: The dynamic visible watermark contains information such as user ID and real-time timestamp. It uses a random path switching algorithm to adjust the position and adaptively adjusts the transparency and brightness parameters according to the screen brightness.

[0033] Preferably, S4 includes:

[0034] S41: Organize the watermark embedding parameters and transmit them to the edge node, adapt to the encoder, and realize the timing synchronization of watermark embedding and encoding;

[0035] S42: During the quantization stage, differentiated quantization parameters and quantization matrix optimization strategies are adopted for macroblocks corresponding to different watermark types;

[0036] S43: Intra-frame prediction traverses multiple modes to select a scheme that meets the watermark loss rate target, and inter-frame prediction optimizes motion compensation accuracy and extends the reference frame buffer time.

[0037] S44: Divide the watermarked data into high-priority categories and process them first using hierarchical coding and adaptive entropy coding algorithms;

[0038] S45: Real-time monitoring of bitrate and control through adjustment of parameters in non-watermarked areas; diverting non-core encoding tasks when edge node load exceeds limits.

[0039] Preferably, S5 includes:

[0040] S51: Separate the original video frames and encoding parameters through the decoder, call the corresponding decryption algorithm and initialize the adaptive extraction engine;

[0041] S52: After extracting the fragile watermark, verify its integrity through hash chain verification and digital signature verification. If the verification fails, locate the tampered area and extent.

[0042] S53: After extracting the semi-fragile watermark, compare the version number, query the consortium blockchain transmission log and verify the timestamp to complete the transmission legality verification;

[0043] S54: Robust watermarks are extracted using an iterative watermark detection algorithm and a transfer learning model, and the core copyright information is recovered after decryption;

[0044] S55: Based on the target detection algorithm, the dynamic visible watermark area is located, and key information is extracted and the infringing entity is initially identified through optical character recognition;

[0045] S56: Upload the verification data to the consortium blockchain, implement multi-node verification through a practical Byzantine fault-tolerant consensus mechanism, generate a copyright verification report, and initiate an infringement warning.

[0046] Preferably, the Transformer branch of the enhanced multimodal deep learning model supports multi-resolution spatial feature capture from 16×16 to 256×256 pixels; the attention mechanism 3D convolution branch uses a combination of 3×3×3 convolution kernels and spatial attention weights to enhance the perception of cross-frame motion information; the improved Gaussian mixture model saliency detection subnetwork introduces human visual contrast weight factors to optimize detection accuracy; and the content security risk assessment subnetwork integrates a historical infringement database interface to realize risk correlation analysis.

[0047] Preferably, the core spatial parameters include texture density entropy, average edge gradient magnitude, proportion of flat regions, and color distribution entropy; the texture density entropy is calculated using a 5×5 pixel window, the average edge gradient magnitude is calculated using the Sobel operator, flat regions are defined as regions with a grayscale difference ≤ 5, and the color distribution entropy is determined based on the histogram entropy value of the HSV color space; the pixel values ​​of the spatial complexity heatmap are negatively correlated with the visual sensitivity of the region, and the higher the pixel value, the more suitable it is for embedding watermarks.

[0048] Preferably, in S14, the temporal dynamic feature extraction adopts a fusion of the Lucas-Kanade optical flow method and the Kalman filter model to calculate the amplitude of the motion vector between adjacent frames, the distribution of multiple directions in the 0-360° interval, and the future motion trend; the inter-frame difference is calculated by the mean of the absolute value of the gray-level difference, and the scene switching probability is determined based on the proportion of frame pairs with an inter-frame structural similarity SSIM less than a threshold.

[0049] Preferably, the nodes of the consortium blockchain include copyright agencies, judicial agencies, and video platforms. A practical Byzantine fault-tolerant consensus mechanism is used to perform hash verification on the copyright verification data packets, and a hash value is generated and written into the block. The copyright verification report is synchronized to all consortium nodes, and the copyright verification report includes screenshots of infringement evidence, original watermark information, and data tracing the dissemination path.

[0050] The beneficial effects of this invention are as follows:

[0051] (1) This invention significantly enhances the survivability of watermarks under various attack and signal processing scenarios through multi-layered hierarchical watermark embedding and collaborative optimization of the encoding process. Robust invisible watermarks are embedded in low-frequency coefficients in the video frame frequency domain. Combined with chaotic encryption or lattice-based homomorphic encryption technology, they can effectively resist multiple compressions, transcodings, partial cropping, and noise interference of ultra-high-definition videos. Even after high-intensity compression, the accuracy of copyright information extraction can still maintain a high level. Vulnerable invisible watermarks rely on hash chain encryption and multi-check bit design, which are highly sensitive to any tampering operations. They can accurately locate the tampering area and degree, providing a reliable basis for content integrity verification. Dynamic visible watermarks effectively avoid being cropped or obscured through random path switching and adaptive brightness adjustment, forming an intuitive infringement deterrent. The three layers of watermarks work together to build a comprehensive risk resistance system.

[0052] (2) The multimodal content analysis of this invention enables adaptive dynamic adjustment of watermark parameters, perfectly balancing robustness and concealment. By extracting the spatial complexity, temporal dynamics, and content saliency features of the video, the watermark is preferentially embedded in non-saliency areas that are not sensitive to the human eye, and the watermark intensity is reduced in sensitive areas, so that the impact of watermark embedding on the visual quality of the video is controlled within a range imperceptible to the human eye. This achieves effective carrying of copyright information without affecting the normal viewing experience of the audience.

[0053] (3) This invention seamlessly integrates watermark parameters with mainstream coding standards. During the coding and quantization stage, a differentiated quantization strategy is adopted for macroblocks containing watermarks. Macroblocks with robust watermarks are embedded with smaller quantization parameters to reduce information loss, while regions with fragile watermarks have their quantization matrices optimized to enhance stability, thus reducing the damage to the watermark during compression at its source. During the intra-frame prediction and entropy coding stages, the prediction mode with the least watermark loss is prioritized, and watermarked data is divided into high-priority data for priority transmission, ensuring the integrity of the encoded watermark information. This collaborative mechanism significantly improves the survival rate of watermarks in ultra-high-definition video with high compression ratios.

[0054] (4) This invention relies on a pre-trained multimodal deep learning model to analyze the content features of different types of videos in real time, dynamically adjust the watermark strength, position and type, and achieve intelligent processing throughout the entire process without manual intervention. For high-value copyrighted content, the full three-layer watermark mode is automatically activated. For ordinary short videos, only the robust invisible watermark layer can be enabled to reduce resource consumption. For scenarios requiring real-time source tracing, the dynamic visible watermark layer is additionally activated.

[0055] In summary, this invention achieves breakthroughs in watermark robustness, concealment, coding coordination, scene adaptability, and copyright verification integrity, and can meet the full-chain needs of ultra-high-definition video from production and transmission to copyright protection. Attached Figure Description

[0056] Figure 1 This is a diagram illustrating the method steps of an embodiment of the present invention. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] This embodiment uses 8K resolution ultra-high-definition video as the core processing object, selects H.266 / VVC enhanced version as the target encoding standard, and integrates Transformer architecture and attention mechanism convolution to construct an enhanced multimodal deep learning model to achieve intelligent, highly secure, and low-latency adaptive digital watermarking processing throughout the entire process. Figure 1 As shown, it includes the following steps:

[0059] S1: Multimodal video content deep analysis and full-dimensional feature extraction. This step constructs a full-dimensional feature representation of the video content through frame-level splitting, multi-branch feature extraction, and feature fusion, providing data support for subsequent watermark parameter decisions. Specifically, it includes the following steps:

[0060] S11: Frame-level video stream splitting and intelligent flexible window grouping. The input 8K video stream is split frame by frame, and a scene complexity prediction model is simultaneously launched to analyze the dynamic characteristics of the scene in real time. When the scene switching frequency exceeds a preset threshold (≥3 times per second), a window size of 6 frames is used to ensure rapid capture of temporal changes in dynamic scenes. When a static scene is detected, such as a landscape or a fixed shot, the window size is expanded to 40 frames, improving processing efficiency while maintaining the integrity of temporal features. After grouping, frame group data with scene labels is generated, laying the foundation for dimensional adaptation in subsequent feature extraction.

[0061] S12: Input and Branch Initialization of Enhanced Multimodal Deep Learning Model. Each group of labeled frames is input into the pre-trained enhanced multimodal deep learning model, and the model initiates the coordinated operation of its four main functional branches:

[0062] (1) SWIN Transformer branch: used to accurately extract multi-scale spatial features, supporting multi-resolution feature capture from 16×16 to 256×256 pixels.

[0063] (2) Attention mechanism 3D convolution branch: used to capture long-term temporal features. It combines 3×3×3 convolution kernels with spatial attention weights to enhance the perception of cross-frame motion information.

[0064] (3) Improved Gaussian mixture model saliency detection subnetwork: used to locate the visual focus area, and introduces human visual contrast weight factor to optimize detection accuracy.

[0065] (4) Content security risk assessment sub-network: used to determine the video security risk level, and integrates content type classifier, dissemination scenario matching module and historical infringement database interface.

[0066] S13: Spatial Complexity Feature Extraction and Heatmap Generation. Spatial feature parsing is performed on a single-frame image. The specific process is as follows:

[0067] (1) Perform adaptive pyramid downsampling on a single frame image, with sampling levels ranging from 1 to 4 and dynamically adjusted according to the image resolution, preserving key details while reducing computational load.

[0068] (2) Calculate the core space parameters: including the texture density entropy value, which represents the texture richness, and the calculation window is 5×5 pixels; the average edge gradient magnitude, which is calculated using the Sobel operator with a step size of 1 pixel; the proportion of flat areas, the proportion of areas with a gray level difference ≤ 5; and the color distribution entropy based on the histogram entropy value of the HSV color space.

[0069] (3) Generate an ultra-high resolution spatial complexity heatmap based on the above parameters. The heatmap resolution is consistent with the original frame. The pixel value corresponds to the visual sensitivity of the region. The higher the value, the lower the sensitivity and the more suitable it is for embedding watermarks.

[0070] S14: Temporal dynamic feature extraction and quantification index generation. Temporal feature parsing is performed on frame groups. The specific process is as follows:

[0071] (1) The optical flow method, namely the Lucas-Kanade algorithm, and the motion trajectory prediction, namely the Kalman filter model, are combined to calculate the amplitude of the motion vector between adjacent frames, namely the pixel displacement, the direction distribution, namely the 0-360° interval divided into 8 directions, and the motion trend of the next 3 frames.

[0072] (2) Calculate the mean absolute value of the inter-frame difference, i.e. the gray-scale difference, and the scene switching probability, i.e. the proportion of frame pairs with inter-frame structural similarity SSIM≤0.7.

[0073] (3) Output the time dynamics score matrix and motion stability index. Each row of the time dynamics score matrix corresponds to 1 frame and each column corresponds to motion amplitude, continuity and switching probability. The motion stability index is stable if the motion vector standard deviation is ≤2 and violent if >5.

[0074] S15: Visually salient region detection and labeling. Sensitive region localization is performed based on the characteristics of the human visual system. The specific process is as follows:

[0075] (1) Combining the contrast sensitivity characteristics, i.e., the high-frequency region has a higher contrast weight, the visual attention mechanism, i.e., the central region has a higher attention weight, and the color sensitivity model, i.e., the red / green channel is more sensitive than the blue channel, calculate the saliency score of each pixel.

[0076] (2) Mark the regions with a significance score ≥ the threshold as sensitive regions. The threshold is 1.2 times the global score mean. Sensitive regions include human faces, the core of bright objects, etc., and the rest of the regions are marked as non-sensitive regions.

[0077] (3) Generate a salient region mask map for subsequent watermark embedding location selection.

[0078] S16: Content Security Risk Assessment and Level Determination. Risk level classification is performed based on multi-dimensional data. The specific process is as follows:

[0079] (1) Content type classification: Videos are classified into high-value content, medium-value content and low-value content through image classification model. High-value content includes movies, sports events and exclusive documentaries. Medium-value content includes ordinary long videos and high-quality short videos. Low-value content includes daily record short videos.

[0080] (2) Matching of dissemination scenarios: Match the risk coefficient according to the input dissemination platform. The coefficient for professional video platforms is 0.6, the coefficient for social platforms is 1.0, and the coefficient for private domain sharing is 1.2.

[0081] (3) Historical infringement data query: Call the database to query the historical infringement frequency of similar content. A frequency of ≥5 times is considered a high-risk association.

[0082] (4) Comprehensive assessment of risk level: The risk value is calculated by weighted formula. The weighted formula is: content value weight 0.4 + dissemination scenario coefficient 0.3 + historical infringement weight 0.3. The risk value <0.7 is low risk, 0.7-1.0 is medium risk, and >1.0 is high risk.

[0083] S17: Feature Standardization and High-Dimensional Feature Vector Generation. The extracted spatial, temporal, significance, and risk assessment features are processed uniformly.

[0084] (1) Feature standardization: Each feature parameter is standardized using the standard score Z-Score. After standardization, the mean of the feature is 0 and the standard deviation is 1, thus eliminating the difference in dimensions.

[0085] (2) Weighted feature fusion: The standardized features are weighted and summed through a trained weighted network. The weights of spatial features are 0.35, temporal features are 0.3, significance features are 0.2, and risk assessment is 0.15.

[0086] (3) Generate a 1024-dimensional high-dimensional feature vector and transmit it to the adaptive watermark parameter decision module.

[0087] S2: Reinforcement Learning-Driven Adaptive Dynamic Decision-Making of Watermark Parameters. This step, based on the high-dimensional feature vector generated in S1, uses a reinforcement learning model to achieve intelligent decision-making regarding watermark strength, location, and type. Specifically, it includes:

[0088] S21: Construction of a multi-dimensional dynamic decision-making model.

[0089] (1) Model architecture: The deep reinforcement learning framework, namely the DRL framework, is adopted. The joint reward function is watermark anti-attack capability, video visual quality and transmission efficiency. The state space is the high-dimensional feature vector generated by S1, and the action space is the combination of watermark strength, position and type parameters.

[0090] (2) Data support: It integrates historical attack defense database, real-time network bandwidth monitoring data and user profile data. The historical attack defense database contains defense strategies for 20 types of attack scenarios such as compression, transcoding, cropping and noise. The real-time network bandwidth monitoring data is divided into low bandwidth <10Mbps, medium bandwidth 10-50Mbps, and high bandwidth >50Mbps. The user profile data contains creators' priority requirements for copyright protection.

[0091] (3) Autonomous learning mechanism: After each watermark embedding-extraction verification is completed, the result is fed back to the model, and the decision-making strategy is optimized through gradient descent to improve the accuracy of subsequent decisions.

[0092] S22: Dynamic decision-making on watermark intensity based on visual perception and risk dual-driven approach.

[0093] (1) Setting the basic threshold range: Based on the H.266 / VVC encoding characteristics, the basic range of watermark strength is set to 0.1-0.8. The higher the value, the stronger the anti-attack ability and the greater the visual impact.

[0094] (2) Real-time intensity correction logic. Non-sensitive areas + high spatial complexity + high temporal dynamics + high risk level: Intensity is increased to the upper limit of the threshold of 0.7 to 0.8 to enhance resistance to compression and cropping. Non-sensitive areas + medium features + medium risk level: Intensity is adjusted to 0.4 to 0.6 to balance anti-attack capabilities and visual quality. Sensitive areas + low spatial complexity + static + low risk level: Intensity is decreased to the lower limit of the threshold of 0.1 to 0.3 to avoid image quality distortion. Medium feature areas: Combined with real-time bandwidth dynamic fine-tuning, the intensity is reduced by 0.1 to 0.2 when bandwidth is low to reduce data volume.

[0095] S23: Dynamic decision-making for watermark location based on saliency-priority dual guidance.

[0096] (1) Priority location selection: Based on the saliency mask map of S15 and the macroblock priority scoring system, the embedding location is selected. The macroblock priority calculation formula is spatial complexity score × 0.5 + visual attention score × 0.3 + motion stability score × 0.2. Priority is given to macroblocks with non-sensitive areas + macroblock priority ≥ 0.6 + user visual attention < 0.3.

[0097] (2) Intelligent distributed embedding strategy, which is activated when the proportion of non-sensitive areas is <30%. Sparse embedding is performed in low-perception areas 5 to 10 pixels outside the outline of sensitive areas, such as the human figure, with an embedding density not exceeding 10% of the total number of macroblocks in each frame. A macroblock grouping and complementary mechanism is adopted: four adjacent macroblocks are divided into one group, and each group embeds part of the watermark information. The integrity of the watermark is ensured by cross-validation within the group, reducing the impact of single macroblock distortion on vision.

[0098] S24: Dynamic decision-making on watermark type based on a two-dimensional value-cost approach.

[0099] (1) Construction of a multi-dimensional evaluation system: The system is constructed with content value, security requirements and transmission cost as evaluation dimensions. The weight of content value is 0.4, the weight of security requirements is 0.3, the weight of transmission cost is 0.3, and the generation type selection score is generated.

[0100] (2) Watermark combination mode matching. High value + high security requirements + high bandwidth, such as movies and sports events: activate the four-layer watermark mode, namely robust invisible watermark + fragile invisible watermark + semi-fragile invisible watermark + dynamic visible watermark. Medium to low value + medium security requirements + medium to low bandwidth, such as ordinary short videos: activate the two-layer watermark mode, namely robust invisible watermark + semi-fragile invisible watermark. Real-time traceability is required, such as live content + medium to high security requirements: on the basis of the two-layer or four-layer mode, activate the dynamic visible watermark to strengthen the deterrent against infringement.

[0101] S3: Four-layer watermarking hierarchical embedding and advanced encryption enhancement. This step performs four-layer watermarking embedding according to decision parameters, combined with encryption technology to improve security. It consists of four sub-steps:

[0102] S31: Robust invisible watermark embedding for core copyright tracking.

[0103] (1) Macroblock preprocessing: Perform integer DCT transformation on the macroblock selected by S23. The transformation block size is 8×8. Extract the low-frequency coefficients as the embedding carrier. The low-frequency coefficients are the first 16 coefficients and have strong anti-compression ability.

[0104] (2) Watermark data encryption: The copyright ID, creator information and authorization validity period are integrated into the original data. The copyright ID is 128 bits, the creator information is 64 bits, and the authorization validity period is 32 bits. The total length of the original data is 224 bits. It is encrypted by a lattice-based homomorphic encryption algorithm, namely the NTRU algorithm variant, and converted into a quantum computing resistant binary watermark sequence. The sequence length is extended to 512 bits.

[0105] (3) Coefficient modification and embedding: The low-frequency coefficients are modified by using the quantization error minimization algorithm and coefficient stability constraint. The quantization error minimization algorithm ensures that the quantization error of the modified coefficient is ≤2, and the coefficient stability constraint ensures that the difference in modification amount between adjacent coefficients is ≤1. The least significant bit of the low-frequency coefficients is the LSB embedded watermark sequence.

[0106] (4) Encoding compatibility verification: Real-time monitoring of whether the modified coefficients conform to the quantization range of H.266 / VVC Enhanced Edition, which is -2048 to 2047. If it exceeds the range, the modification range will be adjusted back.

[0107] S32: Fragile invisible watermark embedding for content integrity verification.

[0108] (1) Watermark data generation: Integrate content integrity verification information, timestamp, frame sequence number and device fingerprint to generate fragile watermark data. The content integrity verification information is a frame hash value based on SHA-256 with a length of 256 bits, the timestamp is 32 bits, the frame sequence number is 16 bits, the device fingerprint is 64 bits, and the total length of the fragile watermark data is 368 bits.

[0109] (2) Embedding algorithm selection: An improved adaptive LSB algorithm is adopted to dynamically select the embedding bit according to the pixel value distribution. When the pixel value is even, the second LSB is selected, and when it is odd, the first LSB is selected to reduce the impact of embedding on pixel brightness.

[0110] (3) Verification and positioning mark: Multiple verification bits and tampering positioning mark are set in the high frequency coefficients of the video frame. The high frequency coefficients are the last 8 coefficients after DCT transformation. The multiple verification bits are 16 bits and each 8-bit watermark corresponds to 1 verification bit. The tampering positioning mark is 32 bits and the macroblock coordinates of the embedded position are marked.

[0111] (4) Double encryption: The double encryption method combines hash chain encryption and digital signature. The hash chain encryption uses the watermark hash value of the previous frame as the encryption key for the next frame. The digital signature uses the creator's private key to ensure that the watermark cannot be tampered with. Once the content is modified, the watermark hash verification will immediately become invalid.

[0112] S33: Semi-fragile invisible watermark embedding for distinguishing between legal and illegal processing.

[0113] (1) Watermark data generation: integrate the content version number, transmission node ID, and transmission timestamp to generate semi-fragile watermark data. The content version number is 16 bits, the transmission node ID is 32 bits per node and supports up to 8 nodes, the transmission timestamp is 32 bits, and the total length of the semi-fragile watermark data is 288 bits.

[0114] (2) Embedding carrier selection: The intermediate frequency coefficients of the video frame and the second lowest pixel are selected as the joint carrier. The intermediate frequency coefficients are 17-32 coefficients after DCT transformation, taking into account both tolerance and sensitivity. The second lowest pixel is the 3rd LSB.

[0115] (3) Tolerance optimization: An adaptive threshold is used for the embedding part of the mid-frequency coefficient. When the absolute value of the coefficient is >10, the embedding is used to ensure that the watermark is not destroyed by slight compression. The embedding part of the second lowest pixel is compensated by the neighborhood mean. After embedding, the difference between the pixel value and the neighborhood mean is ≤3.

[0116] (4) Sensitivity protection: Increase the embedding density for sensitive areas that are maliciously tampered with, such as cropping or graffiti, with an embedding density of ≥200 pixels per frame, to ensure that the watermark extraction error rate is ≥80% after tampering.

[0117] S34: Dynamically visible watermark embedding for infringement deterrence.

[0118] (1) Watermark content design: The watermark content includes user ID, real-time timestamp, copyright statement abbreviation, and dynamic verification code. The user ID is 32-bit and converted to a decimal string. The real-time timestamp is accurate to the second and the format is YYYY-MM-DD HH:MM:SS. The copyright statement abbreviation is such as ©2024 ContentOwner. The dynamic verification code is updated once every 30 seconds and is a 4-digit number.

[0119] (2) Position and deformation adjustment: The watermark position and deformation are adjusted by a random path dynamic switching algorithm. The algorithm is based on the motion trend prediction of S14. The watermark movement path is consistent with the movement direction of the screen. The watermark position is adjusted in each frame and the offset is ≤10 pixels to avoid the fixed position being cropped.

[0120] (3) Adaptive transparency and brightness: In strong light scenes, i.e., the average brightness of the screen is ≥200, the transparency is set to 20% to 25%, the brightness is increased by 10%, ensuring visibility without being dazzling; In weak light scenes, i.e., the average brightness of the screen is ≤80, the transparency is set to 10% to 15%, the brightness is reduced by 5%, and the dark details are avoided; In medium brightness scenes, the transparency is set to 15% to 20%, the brightness is kept at the default, and the blending with the screen is ensured.

[0121] S4: Edge-Coordinated Watermark Embedding and Video Coding Depth Optimization. This step achieves deep coordination between the watermark and H.266 / VVC coding, reducing the damage to the watermark caused by coding. Specifically, it includes the following steps:

[0122] S41: Watermark embedding parameter transmission and edge coding co-initialization.

[0123] (1) Parameter organization and transmission: The target macroblock coordinates, DCT coefficient quantization adjustment values, watermark type markers, and encryption key indexes of S2 decision are integrated into a parameter package and transmitted to the edge node through a low-latency dedicated data interface. The data interface adopts a variant of the UDP protocol and the latency does not exceed 50ms.

[0124] (2) Encoder adaptation: The H.266 / VVC enhanced encoder of the edge node starts the watermarking collaborative mode and loads the encoding configuration file that matches the watermark type, such as the low-loss quantization configuration corresponding to the robust watermark.

[0125] (3) Cooperative timing synchronization: Ensure that the timing of watermark embedding and encoding frame processing is consistent. After the embedding of one frame is completed, the encoding of the frame is started immediately to avoid parameter misalignment caused by frame buffering.

[0126] S42: Intelligent hierarchical optimization during the encoding and quantization stage.

[0127] (1) Layered quantization strategy: Robust watermark macroblock: Use small quantization parameters, i.e., QP of 22 to 26, and QP of 28 to 32 for regular macroblocks to reduce the loss of information in low and medium frequency coefficients; Fragile watermark region: Optimize the quantization matrix, reduce the quantization step size of high frequency coefficients by 20%, and enhance the stability of LSB bits; Semi-fragile watermark region: Use medium quantization precision, i.e., QP of 26 to 28, to balance tolerance and sensitivity; Ordinary macroblock: Quantize according to the regular target bit rate. The target bit rate for 8K video is 50-80Mbps, and 10% bit rate redundancy is reserved for watermark region optimization.

[0128] (2) 8K video adaptation algorithm: Introducing a macroblock-level quantization parameter adaptive adjustment algorithm, dynamically fine-tuning the QP value according to the macroblock watermark carrying capacity, content complexity and real-time bandwidth. When the macroblock watermark carrying capacity is ≥512 bits, the QP decreases by 2 to 3. When the content complexity is high, the QP increases by 1 to 2. When the bandwidth is insufficient, the QP of ordinary macroblocks increases by 3 to 4 while the watermark macroblock remains unchanged.

[0129] S43: Watermark awareness optimization in intra-frame and inter-frame prediction stages.

[0130] (1) Intra-frame prediction optimization: Start the watermark-aware optimal prediction mode selection algorithm, and traverse the 35 intra-frame prediction modes of H.266 / VVC, including Planar, DC, angle mode, etc. Calculate the loss rate of watermark information under each mode. The loss rate is the ratio of the number of watermark bit errors after prediction to the total number of bits. Select the mode with a loss rate ≤5% to perform prediction.

[0131] (2) Inter-frame prediction optimization: The bidirectional prediction and motion compensation optimization are integrated. The bidirectional prediction combines the forward reference frame and the backward reference frame. The compensation accuracy is adjusted according to the motion vector of the watermark embedding position, i.e., the S14 output. The motion vector error is ≤1 pixel. For watermarked macroblocks, the reference frame buffer time is extended from 3 frames to 5 frames to avoid watermark distortion caused by frequent changes in reference frames.

[0132] S44: High-priority data processing during the entropy coding stage.

[0133] (1) Data layering: Macroblock data containing watermark information is divided into high-priority data, which includes watermark embedding correlation coefficient and check bit. Ordinary macroblock data is divided into low-priority data.

[0134] (2) Encoding order adjustment: The layered encoding and context-adaptive binary arithmetic encoding, namely CABAC, are combined to prioritize the encoding and transmission of high-priority data, ensuring that watermark data occupies bandwidth first.

[0135] (3) Efficiency optimization: An adaptive entropy coding optimization algorithm is introduced to dynamically adjust the context model according to the binary distribution of the watermark data, i.e. the proportion of 0 / 1, thereby improving the coding efficiency by 10% to 15% and avoiding the code rate exceeding the standard due to the increase of watermark data.

[0136] S45: Dynamic bitrate control and edge node load balancing.

[0137] (1) Real-time bitrate monitoring: The output bitrate is monitored through the bitrate statistics module, with a sampling interval of 100ms, and compared with the target bitrate, such as 50Mbps for 8K video. If the bitrate exceeds the target bitrate by 10%, the QP and intra-frame prediction mode in the non-watermarked area are fine-tuned, the QP is increased by 2 to 3, the intra-frame prediction mode with a higher compression ratio is selected, and the parameters of the watermarked area are not adjusted. If the bitrate is insufficient, that is, lower than the target bitrate by 5%, the entropy coding efficiency of the watermarked area is optimized, the frequency of context model updates is increased, and the amount of effective data is increased.

[0138] (2) Load balancing: Real-time monitoring of CPU utilization of edge nodes, with a threshold of 80%. If the utilization exceeds the threshold, non-core coding tasks such as inter-frame prediction of ordinary macroblocks will be automatically diverted to adjacent edge nodes. The node with the lowest load will be selected based on the node load heat map to ensure that the coding delay is ≤100ms to meet the 8K real-time transmission requirements.

[0139] S5: Intelligent extraction and consortium blockchain-based copyright verification throughout the entire process. This step enables watermark extraction, verification, and infringement tracing, constructing a closed-loop protection system. Specifically, it includes the following steps:

[0140] S51: Video stream edge decoding preprocessing and decryption algorithm call.

[0141] (1) Edge decoding: The video stream to be verified is transmitted to the edge node, and the H.266 / VVC decoder is started to decode, separating the original video frames and encoding parameters. The encoding parameters include QP value, prediction mode, etc.

[0142] (2) Parameter matching: According to the watermark type marker in the encoding parameters, the corresponding decryption algorithm is called from the key management library. For example, robust watermark corresponds to NTRU decryption, and fragile watermark corresponds to hash chain decryption.

[0143] (3) Extraction engine initialization: Start the adaptive extraction engine, load the feature templates that match the embedding time, such as the space complexity heatmap template generated by S1, to improve the extraction accuracy.

[0144] S52: Extraction and integrity verification of fragile, invisible watermarks.

[0145] (1) Watermark extraction: Perform high-frequency coefficient inverse DCT transform on the decoded video frame to extract multiple check bits and tamper location markers; recover fragile watermark data through adaptive LSB extraction algorithm, and the extraction algorithm is consistent with the bit selection logic during embedding.

[0146] (2) Integrity verification. Hash chain verification: Calculate the hash value of the watermark in the current frame and compare it with the hash value after decryption of the previous frame. If they match, the verification passes. Digital signature verification: Verify the watermark signature using the creator's public key to confirm the legality of the watermark's source.

[0147] (3) Tampering location: If the verification fails, the tampered area is located according to the tampering location mark, i.e. macroblock coordinates and block matching algorithm. The block matching algorithm compares the current frame macroblock with the original frame macroblock. The error does not exceed 1 macroblock. The output tampering range is such as frame 100-120, macroblock (20,30)-(40,50) and tampering degree such as 30% cropping and 15% graffiti coverage.

[0148] S53: Extraction and legality verification of semi-fragile invisible watermarks.

[0149] (1) Watermark extraction: Perform inverse DCT transformation of intermediate frequency coefficients and parsing of the second lowest pixel of the video frame to recover the content version number, transmission node ID and timestamp.

[0150] (2) Legality Verification. Version Consistency: Compare the extracted version number with the version number filed by the copyright holder to confirm whether it is an authorized version. Transmission Link Verification: Query the transmission log in the consortium blockchain based on the transmission node ID to verify whether the node is an authorized transmission node, such as whether it is in the list of CDN nodes allowed by the copyright holder. Timestamp Verification: Confirm that the transmission timestamp is within the authorization validity period to avoid the spread of expired content.

[0151] S54: Robust extraction of invisible watermarks and restoration of core copyright information.

[0152] (1) Coefficient recovery: Perform integer DCT inverse transform on the video frame to extract the low- and mid-frequency coefficients. The coefficient positions are consistent with those during embedding.

[0153] (2) Watermark extraction: The iterative watermark detection algorithm and the transfer learning model are used to extract the binary watermark sequence from the coefficients. The iterative watermark detection algorithm iterates 5 to 8 times and corrects the coefficient error in each iteration. The transfer learning model is pre-trained in 20 attack scenarios to improve anti-interference ability.

[0154] (3) Information restoration: Decrypt the watermark sequence using the NTRU decryption algorithm to restore the copyright ID, creator information and authorization validity period.

[0155] (4) Fault tolerance guarantee: Even if the video undergoes 3 H.266 / VVC compressions, 1 10% cropping, or Gaussian noise interference, the 3 compression bitrates are reduced from 80Mbps to 20Mbps, and the Gaussian noise interference variance is 0.01, the core information extraction accuracy is still guaranteed to be ≥95%.

[0156] S55: Dynamically visible watermark extraction and preliminary identification of infringing entities.

[0157] (1) Watermark recognition: The dynamic watermark detection model based on YOLOv8 (a deep learning object detection algorithm) is used to locate the visible watermark area in the video frame, with a detection accuracy of ≥98%.

[0158] (2) Information parsing: User ID and timestamp are extracted by optical character recognition (OCR), and dynamic verification codes are parsed by verification code recognition model.

[0159] (3) Subject identification: The user ID is compared with the user registration information in the consortium blockchain to initially identify possible infringing subjects such as uploading accounts and devices; the time period of the infringement is confirmed by combining the timestamp to provide preliminary clues for tracing the source.

[0160] S56: Closed loop of consortium blockchain collaborative verification and infringement early warning.

[0161] (1) Data on the blockchain: The extracted four-layer watermark information, integrity verification results, tampered location data, and transmission logs are integrated into a copyright verification data packet and transmitted to the copyright alliance chain. The alliance nodes include copyright agencies, judicial agencies, and video platforms.

[0162] (2) Multi-node consensus: The consortium blockchain uses the Practical Byzantine Fault Tolerance consensus mechanism, namely PBFT consensus mechanism, to perform hash verification on data packets, generate SHA-256 hash values ​​and write them into blocks. The consensus delay is ≤3 seconds, ensuring that the data cannot be tampered with.

[0163] (3) Verification report generation: The consortium blockchain automatically generates a copyright verification report, which includes the verification result (i.e., whether there is infringement or not), evidence of infringement (i.e., tampered screenshots), watermark information, and subject information (i.e., the infringer initially identified), and is synchronized to all consortium nodes.

[0164] (4) Infringement warning: If infringement is determined, the real-time warning module is activated and the warning information is pushed to the copyright holder via SMS and API interface. The information includes the source of infringement, i.e., IP address, platform account, dissemination path, i.e., the list of platforms that have been disseminated, and the tampering situation. The copyright holder can initiate rights protection actions based on the report to complete the closed loop of the whole process of extraction-verification-warning-rights protection.

[0165] This embodiment achieves deep collaboration between 8K ultra-high-definition video watermark embedding and encoding through the above steps. While ensuring video visual quality, it significantly improves the watermark's resistance to compression and tampering, achieving a subjective quality score (MOS) ≥ 4.2 out of 5. Simultaneously, it utilizes a consortium blockchain to ensure compliance and traceability of copyright verification, meeting the end-to-end copyright protection needs of ultra-high-definition video from production to distribution.

[0166] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An adaptive multi-level digital watermarking method based on encoding process optimization, characterized in that, Includes the following steps: S1: Perform multimodal deep content analysis on the input video stream, extract spatial complexity, temporal dynamics, visual saliency and content security risk assessment features, and generate a unified high-dimensional feature vector through standardization and weighted fusion; S2: Based on the high-dimensional feature vector, the watermark strength, embedding position, and type combination are adaptively determined through a reinforcement learning-driven dynamic decision-making model; S3: Perform multi-level watermarking hierarchical embedding according to the decision results, and adopt corresponding embedding strategies for robust invisible watermarks, fragile invisible watermarks, semi-fragile invisible watermarks and dynamic visible watermarks respectively, and combine encryption technology to enhance security. S4: The watermark embedding process is deeply integrated with video coding, and a differentiated optimization strategy is adopted in the quantization, prediction and entropy coding stages, while achieving dynamic bitrate control and edge node load balancing. S5: Performs edge decoding preprocessing on the video stream to be verified, adaptively extracts watermark information from each layer, and combines the consortium blockchain to complete integrity verification, legality verification, and infringement tracing.

2. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 1, characterized in that, S1 includes: S11: Perform frame-level splitting on the video stream and dynamically adjust the window size and number of frames based on the scene switching frequency using an elastic window grouping strategy. S12: Input the labeled frame group into the multi-branch enhanced multimodal deep learning model, which includes a Transformer branch, an attention mechanism 3D convolution branch, an improved Gaussian mixture model saliency detection subnetwork, and a content security risk assessment subnetwork; S13: Generate a spatial complexity heatmap with the same resolution as the original frame by adaptive pyramid downsampling and core spatial parameter calculation; S14: Integrate optical flow method with motion trajectory prediction model to extract temporal dynamic features and generate scoring matrix and motion stability index; S15: Detect visually salient regions based on the characteristics of the human visual system and generate a salient region mask map; S16: Combining content value, dissemination scenarios, and historical infringement data, a weighted formula is used to determine the content security risk level; S17: Standard scores are used to process various features, and high-dimensional feature vectors are generated by weighted fusion.

3. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 1, characterized in that, S2 includes: S21: Construct a dynamic decision model for a deep reinforcement learning framework, using watermark anti-attack capability, video visual quality, and transmission efficiency as the joint reward function, integrating multi-source data support, and establishing an autonomous learning mechanism; S22: Set the basic threshold range for watermark strength and perform real-time correction based on visual perception characteristics, content features and risk level; S23: The embedding position is selected based on the saliency mask map and macroblock priority score. When the proportion of non-sensitive areas is insufficient, a complementary strategy of sparse embedding and macroblock grouping is adopted. S24: Construct a multi-dimensional evaluation system for content value, security requirements, and transmission costs, and match the corresponding watermark type combination mode based on the evaluation results.

4. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 1, characterized in that, S3 includes: S31: Robust invisible watermark is embedded in low-frequency coefficients of video frames. The core copyright information is encrypted using a lattice-based homomorphic encryption algorithm. The embedding is achieved by modifying the coefficients through a quantization error minimization algorithm. S32: The fragile invisible watermark uses an improved adaptive LSB algorithm to embed high-frequency coefficients, integrates frame hash values, timestamps and other information, and employs dual protection of hash chain encryption and digital signature; S33: Semi-fragile invisible watermarks are embedded in the mid-frequency coefficients and the second-lowest pixel bit, and the tolerance and sensitivity are balanced by adaptive threshold and neighborhood mean compensation. S34: The dynamic visible watermark contains information such as user ID and real-time timestamp. It uses a random path switching algorithm to adjust the position and adaptively adjusts the transparency and brightness parameters according to the screen brightness.

5. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 1, characterized in that, S4 includes: S41: Organize the watermark embedding parameters and transmit them to the edge node, adapt to the encoder, and realize the timing synchronization of watermark embedding and encoding; S42: During the quantization stage, differentiated quantization parameters and quantization matrix optimization strategies are adopted for macroblocks corresponding to different watermark types; S43: Intra-frame prediction traverses multiple modes to select a scheme that meets the watermark loss rate target, and inter-frame prediction optimizes motion compensation accuracy and extends the reference frame buffer time. S44: Divide the watermarked data into high-priority categories and process them first using hierarchical coding and adaptive entropy coding algorithms; S45: Real-time monitoring of bitrate and control through adjustment of parameters in non-watermarked areas; diverting non-core encoding tasks when edge node load exceeds limits.

6. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 1, characterized in that, S5 includes: S51: Separate the original video frames and encoding parameters through the decoder, call the corresponding decryption algorithm and initialize the adaptive extraction engine; S52: After extracting the fragile watermark, verify its integrity through hash chain verification and digital signature verification. If the verification fails, locate the tampered area and extent. S53: After extracting the semi-fragile watermark, compare the version number, query the consortium blockchain transmission log and verify the timestamp to complete the transmission legality verification; S54: Robust watermarks are extracted using an iterative watermark detection algorithm and a transfer learning model, and the core copyright information is recovered after decryption; S55: Based on the target detection algorithm, the dynamic visible watermark area is located, and key information is extracted and the infringing entity is initially identified through optical character recognition; S56: Upload the verification data to the consortium blockchain, implement multi-node verification through a practical Byzantine fault-tolerant consensus mechanism, generate a copyright verification report, and initiate an infringement warning.

7. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 2, characterized in that, The Transformer branch of the enhanced multimodal deep learning model supports multi-resolution spatial feature capture from 16×16 to 256×256 pixels; The attention mechanism 3D convolution branch combines a 3×3×3 convolution kernel with spatial attention weights to enhance the perception of cross-frame motion information; the improved Gaussian mixture model saliency detection subnetwork introduces human visual contrast weight factors to optimize detection accuracy; and the content security risk assessment subnetwork integrates a historical infringement database interface to realize risk correlation analysis.

8. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 2, characterized in that, The core spatial parameters include texture density entropy, average edge gradient magnitude, proportion of flat regions, and color distribution entropy. The texture density entropy is calculated using a 5×5 pixel window, the average edge gradient magnitude is calculated using the Sobel operator, flat regions are defined as regions with a grayscale difference ≤ 5, and the color distribution entropy is determined based on the histogram entropy value of the HSV color space. The pixel values ​​of the spatial complexity heatmap are negatively correlated with the visual sensitivity of the region, and the higher the pixel value, the more suitable it is for embedding watermarks.

9. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 2, characterized in that, The temporal dynamic feature extraction in S14 adopts a fusion of Lucas-Kanade optical flow method and Kalman filter model to calculate the amplitude of motion vectors between adjacent frames, the distribution of multiple directions in the 0-360° interval, and the future motion trend. Inter-frame difference is calculated using the mean of the absolute values ​​of grayscale differences, and the scene switching probability is determined based on the proportion of frame pairs with an inter-frame structural similarity (SSIM) less than a threshold.

10. The adaptive multi-level digital watermarking method based on encoding process optimization according to claim 6, characterized in that, The nodes of the consortium blockchain include copyright agencies, judicial institutions, and video platforms. It uses a practical Byzantine fault-tolerant consensus mechanism to perform hash verification on copyright verification data packets, generate hash values, and write them into blocks. The copyright verification report is synchronized to all consortium nodes. The copyright verification report includes screenshots of infringement evidence, original watermark information, and data tracing the dissemination path.

Citation Information

Patent Citations

  • Video monitoring system and video data authentication method thereof

    CN101668185A

  • Method for embedding multiple robust watermarks in video based on three-dimensional Discrete Fourier Transformation (DFT)

    CN102510521A

  • H.264 video watermark embedding and extraction method based on mixed coding / decoding

    CN103152578A

  • Video copyright protection method and system based on block chain and digital watermarking technology

    CN114003871A

  • Transform-based soft fusion robust image watermarking method

    CN115880125A

Cited By

  • A digital watermark embedding method and device, electronic equipment and storage medium

    CN122179633A

  • Watermark adaptive embedding method based on dataset features

    CN122333432A

  • An Adaptive Watermark Embedding Method Based on Dataset Features

    CN122333432B