Big data watermarking method and device based on artificial intelligence, equipment and medium

Through the method of combining federated learning and generative adversarial networks, watermark signals consistent with the distribution of big data are generated, and time-sensitive reinforcement learning and cross-modal autoencoder are used for dynamic adjustment and verification, solving the robustness and concealment problems of multimodal big data watermarks, realizing trusted management of watermarks in distributed environments.

CN120509015AInactive Publication Date: 2025-08-19GUILIN UNIV OF TECH AT NANNING
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510619955.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional watermarking technology is difficult to adapt to the characteristic differences of multimodal big data, resulting in data distortion or insufficient robustness after watermark embedding, unable to cope with the dynamic changes of real-time streaming data, and faces the risks of data leakage and illegal tampering in a distributed environment.

Method used

Cross-modal correlation analysis is performed through the federated learning framework, the data is divided into real-time streaming data blocks and static data blocks, and the watermark signal consistent with the data distribution is generated using the generative adversarial network, and the embedding parameters are dynamically adjusted in combination with the time-sensitive reinforcement learning model, the watermark signal is extracted using a cross-modal autoencoder, the anti-attack verification model is used for integrity verification, and the watermark key and parameters are recorded through the blockchain.

Benefits of technology

It improves the extraction accuracy and completeness verification efficiency of multimodal big data watermarks, enhances the robustness and concealment of watermarks in a distributed environment, and realizes the trusted management of watermarks throughout the life cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509015A_ABST
    Figure CN120509015A_ABST
Patent Text Reader

Abstract

The invention relates to a big data watermarking method and device based on artificial intelligence, equipment and a medium. The method comprises the following steps: dividing distributed node data into a real-time stream data block and a static data block, and performing watermark embedding priority marking on the real-time stream data block and the static data block; generating a first watermark signal and a second watermark signal based on a generative adversarial network, dynamically adjusting the embedding parameters of the watermark signals through a time-sensitive reinforcement learning model, and generating a watermark key bound with the embedding parameters; extracting an embedded watermark signal through a preset cross-modal auto-encoder, and verifying the embedded watermark signal through an anti-attack verification model; and performing dynamic watermarking according to a verification result, and recording a watermark key and a corresponding embedding parameter through a block chain. According to the method, the robustness, the security and the management efficiency of the big data watermark are further improved through the technical means of data classification and priority marking, adversarial watermark generation and dynamic embedding, cross-modal feature extraction and anti-attack verification, block chain recording and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and in particular relates to a big data watermarking method, device, equipment and medium based on artificial intelligence. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, data sharing and collaborative computing in distributed environments across organizations and platforms have become the norm. Federated learning, a new distributed machine learning framework, allows different nodes to jointly model data without leaking the original data, effectively solving the data silo problem. However, data faces security risks such as data leakage, illegal tampering, and copyright disputes during distributed transmission, storage, and multi-party collaborative processing. Watermarking technology is a key technical means to address these issues.

[0003] However, traditional digital watermarking technologies are mostly designed for single-modal data, such as images and audio, embedding imperceptible identifying information within the data to achieve copyright protection and content authentication. However, in big data scenarios, data exhibits multi-source, multi-modal, and dynamic characteristics, making traditional watermarking technologies difficult to meet the needs of complex data environments. For example, single-modal watermark embedding strategies cannot adapt to the characteristic differences of multimodal data, which may lead to data distortion or insufficient watermark robustness after embedding. Furthermore, fixed embedding parameters are difficult to cope with the dynamic changes of real-time streaming data, and can easily cause watermark loss when the data is updated or attacked. Summary of the Invention

[0004] Based on this, it is necessary to provide artificial intelligence-based big data watermarking methods, devices, equipment and media to address the above technical problems, so as to improve the extraction accuracy and integrity verification efficiency of multimodal data watermarks, and enhance the robustness and concealment of watermarks in distributed environments.

[0005] In the first aspect, this application provides a big data watermarking method based on artificial intelligence, including:

[0006] Performing cross-modal correlation analysis on multiple distributed node data through a federated learning framework to obtain a global feature distribution, dividing the distributed node data into real-time streaming data blocks and static data blocks, and embedding watermarks and priority markings on the real-time streaming data blocks and the static data blocks to obtain a marked first real-time streaming data block set and a first static data block set;

[0007] Based on a generative adversarial network, a first watermark signal and a second watermark signal are generated that are consistent with the data distribution of the first real-time stream data block set and the first static data block set, respectively. Embedding parameters of the first watermark signal and the second watermark signal are dynamically adjusted through a time-sensitive reinforcement learning model to generate a second real-time stream data block set and a second static data block set containing watermarks, and a watermark key bound to the embedding parameters is generated.

[0008] By using a preset cross-modal autoencoder, the embedded watermark signal in the second real-time stream data block set and the second static data block set is extracted, and the integrity of the embedded watermark signal is verified by using an anti-attack verification model, and the verification result is output;

[0009] According to the verification result, the second real-time stream data block set and the second static data block set are dynamically watermarked to generate an updated second real-time stream data block set and an updated second static data block set, and the watermark key and the corresponding embedding parameters are recorded through the blockchain.

[0010] In one embodiment, cross-modal correlation analysis is performed on multiple distributed node data through a federated learning framework to obtain a global feature distribution, the distributed node data is divided into real-time streaming data blocks and static data blocks, and watermarks are embedded in the real-time streaming data blocks and the static data blocks to mark their priority, thereby obtaining a marked first real-time streaming data block set and a first static data block set, including:

[0011] Each distributed node performs standardized calculation and encryption on the mean, variance, and sparsity of local data to obtain the encrypted features of each distributed node data, and uploads the encrypted features to the central server;

[0012] The central server decrypts the encrypted features and aggregates them to generate a global feature distribution;

[0013] A cross-modal association graph is constructed based on the global feature distribution. The nodes of the cross-modal association graph include image modality nodes, text modality nodes, and time series modality nodes. A shared feature set with redundancy higher than a preset threshold is extracted from the cross-modal association graph and embedded as a watermark.

[0014] Distributed node data is divided into real-time streaming data blocks and static data blocks according to data flow rate and data modality type. Real-time streaming data blocks include image modality and time series modality data, and static data blocks include text modality data.

[0015] Based on the shared feature set and the reinforcement learning model, watermark embedding priorities are assigned to the real-time streaming data blocks and the static data blocks to generate a first real-time streaming data block set and a first static data block set.

[0016] In one embodiment, generating a second real-time stream data block set containing a watermark and a second static data block set includes:

[0017] Generate a first watermark signal and a second watermark signal respectively consistent with the data distribution of the first real-time stream data block set and the first static data block set by a generative adversarial network;

[0018] Based on the initial embedding strength parameter and the initial embedding position coordinates, embedding the first watermark signal and the second watermark signal into the first real-time stream data block set and the first static data block set respectively to generate a temporary watermarked data block set;

[0019] A time-sensitive reinforcement learning model is constructed by defining a state space, an action space, and a reward function, wherein the state space includes the data block flow rate, sensitivity, current watermark capacity, and watermark embedding priorities of a temporary watermarked data block set and a first real-time stream data block set and a first static data block set; the action space includes an embedding strength parameter and a set of embedding position coordinates; and the reward function includes a stealth score and a robustness score. The stealth score is calculated by the confidence output value of the discriminator of the generative adversarial network for the temporary watermarked data block set, and the robustness score is calculated by the watermark extraction success rate after simulating a resampling attack on the temporary watermarked data block set.

[0020] The proximal strategy optimization algorithm is used to train the time-sensitive reinforcement learning model. The embedding strength parameters and embedding position coordinate sets of the first and second watermark signals are dynamically adjusted with the watermark embedding priority as the constraint condition to obtain the optimized embedding strength parameters and the optimized embedding position coordinate set.

[0021] According to the optimized embedding strength parameter and the optimized embedding position coordinate set, corresponding watermark signals are embedded in the corresponding coordinate positions of the first real-time stream data block set and the first static data block set to generate a second real-time stream data block set and a second static data block set.

[0022] In one embodiment, extracting embedded watermark signals from the second real-time stream data block set and the second static data block set by presetting a cross-modal autoencoder includes:

[0023] By presetting a cross-modal autoencoder, the following operations are performed on each data block in the second real-time stream data block set and the second static data block according to the data type:

[0024] Perform multi-scale convolution operations on image data blocks to generate deep feature maps;

[0025] Use a multi-head attention mechanism to process text data blocks and generate semantic feature vectors;

[0026] Performing bidirectional LSTM processing on the time series data blocks in the second real-time stream data block set and the second static data block set to generate a time series feature sequence;

[0027] The deep feature map, semantic feature vector and temporal feature sequence are mapped to the shared feature space through a preset learnable weight matrix to generate fusion features;

[0028] Based on the fusion features, a watermark mask is generated through a gated convolutional network, and the area in the watermark mask where the element value is greater than the preset mask threshold is determined as the potential watermark area;

[0029] The features corresponding to the potential watermark area are sparsely constrained to extract the watermark feature vector, which is then input into the pre-trained adversarial decoder to output the embedded watermark signal.

[0030] In one embodiment, integrity verification is performed on the embedded watermark signal using an anti-attack verification model, and the verification result is output, including:

[0031] The following simulated attack operations are performed on the watermarked data in the second real-time stream data block set and the second static data block set:

[0032] Perform random downsampling and bicubic interpolation restoration on the image data block to generate a perturbed image data block;

[0033] Add Gaussian noise and salt and pepper noise to the time series data block to generate a disturbed time series data block;

[0034] Compress the text data block from 32-bit floating point to 8-bit integer to generate a quantized text data block;

[0035] Based on the embedded watermark signal, the structural similarity index value and peak signal-to-noise ratio value of the perturbed image data block and the perturbed time series data block with the corresponding original data block are calculated, and the watermark signal bit error rate is calculated for the quantized text data block;

[0036] The verification threshold corresponding to the structural similarity index value, peak signal-to-noise ratio value and watermark signal bit error rate is dynamically adjusted according to the data block type to generate the verification result.

[0037] In one embodiment, dynamically updating the second real-time stream data block set and the second static data block set based on the verification result to generate an updated second real-time stream data block set and an updated second static data block set includes:

[0038] When the verification result meets the preset conditions, the invalid data block index list is output and the watermark update is triggered;

[0039] Calculate the dynamic attenuation factor based on the historical watermark extraction success rate;

[0040] Generate a set of corrected embedding strength parameters and corrected embedding position coordinates based on the current data distribution and dynamic attenuation factor through a time-sensitive reinforcement learning model;

[0041] Based on the invalid data block index list, performing an inverse embedding operation on the invalid data blocks in the second real-time stream data block set and the second static data block set to remove the original watermark signal;

[0042] According to the shared feature set, an anti-attack enhanced watermark signal is generated, and according to the corrected embedding strength parameter and the corrected embedding position coordinate set, the anti-attack enhanced watermark signal is embedded in the corresponding coordinate position of the invalid data block to generate an updated second real-time stream data block set and an updated second static data block set.

[0043] In one embodiment, after obtaining the first watermark signal and the second watermark signal, the method further includes: generating a noise signal through a Laplace mechanism according to a preset privacy budget, wherein the noise signal includes image noise, text noise, and time series noise;

[0044] Superimposing the noise signal onto the first watermark signal and the second watermark signal to generate a first privacy-preserving watermark and a second privacy-preserving watermark respectively;

[0045] The privacy embedding parameters of the first privacy-preserving watermark and the second privacy-preserving watermark are dynamically adjusted through a time-sensitive reinforcement learning model, and the privacy embedding parameters are encrypted using the Paillier algorithm to obtain encrypted parameters, which are then distributed and stored in each distributed node.

[0046] During the back-propagation process of federated learning, a random mask matrix is added to the watermark generation gradient matrix of the generative adversarial network and the watermark embedding gradient matrix of the time-sensitive reinforcement learning model to generate obfuscated gradients.

[0047] The obfuscated gradients are uploaded to a central server for aggregation, and the parameters of the generative adversarial network and time-sensitive reinforcement learning model are updated based on the aggregation results.

[0048] Secondly, this application also provides a big data watermarking device based on artificial intelligence, including:

[0049] A feature analysis and dynamic segmentation module is used to perform cross-modal correlation analysis based on multiple distributed node data through a federated learning framework to obtain a global feature distribution, divide the distributed node data into real-time streaming data blocks and static data blocks, and embed watermarks and priority tags on the real-time streaming data blocks and static data blocks to obtain a marked first set of real-time streaming data blocks and a first set of static data blocks;

[0050] a watermark generation and adaptive embedding module, configured to generate, based on a generative adversarial network, a first watermark signal and a second watermark signal that are consistent with the data distribution of the first real-time stream data block set and the first static data block set, respectively; dynamically adjust embedding parameters of the first watermark signal and the second watermark signal through a time-sensitive reinforcement learning model; generate a second real-time stream data block set and a second static data block set containing the watermark; and generate a watermark key bound to the embedding parameters;

[0051] a watermark extraction and robustness verification module, configured to extract the embedded watermark signal from the second real-time stream data block set and the second static data block set through a preset cross-modal autoencoder, perform integrity verification on the embedded watermark signal through an anti-attack verification model, and output a verification result;

[0052] The dynamic watermark update and evidence storage module is used to dynamically update the watermark of the second real-time stream data block set and the second static data block set based on the verification results, generate an updated second real-time stream data block set and an updated second static data block set, and record the watermark key and corresponding embedding parameters through the blockchain.

[0053] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the first aspect when executing the computer program.

[0054] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the first aspect when executed by a processor.

[0055] The aforementioned AI-based big data watermarking method, apparatus, device, and medium utilize a federated learning framework to perform cross-modal correlation analysis on data from multiple distributed nodes, obtaining a global feature distribution. Based on this, the data is divided into real-time streaming data blocks and static data blocks, while also marking watermark embedding priorities. This provides a structured, prioritized, and high-quality data foundation for subsequent watermark processing. Secondly, a generative adversarial network is used to generate a watermark signal consistent with the original data distribution, generating watermarked data blocks and bound watermark keys. This ensures a high degree of compatibility between the watermark signal and the original data. Furthermore, a reinforcement learning model adaptively optimizes embedding parameters based on the dynamic characteristics of the data and the watermarking requirements, achieving a balance between watermark concealment and robustness, and improving the watermark's viability in complex data environments. Furthermore, a pre-configured cross-modal autoencoder accurately extracts watermark signals based on the characteristics of multimodal data. An anti-attack verification model comprehensively assesses the integrity of the watermark by simulating various attack scenarios, providing a rigorous verification mechanism for its reliability.

[0056] Finally, the data block is dynamically watermarked based on the verification results, ensuring that the watermark is always valid during data updates and version iterations. Blockchain technology provides tamper-proof and trusted storage for watermark keys and embedded parameters, ensuring data traceability and copyright management in a distributed environment, and realizing closed-loop management of the entire process from watermark generation, embedding, extraction and verification to updating and storing evidence.

[0057] Compared with traditional big data watermarking methods, this method significantly improves the extraction accuracy of multimodal big data watermarks through technical means such as federated learning-driven data preprocessing, watermark generation and embedding that integrates generative adversarial and reinforcement learning, and cross-modal adaptive watermark extraction and verification. It also enhances the robustness and concealment of watermarks in complex attack environments, and realizes the full life cycle trusted management of watermarks in distributed data scenarios, providing a more scientific, efficient and reliable solution for the secure sharing, copyright protection and content traceability of big data. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 A flowchart of a big data watermarking method based on artificial intelligence is provided as an exemplary embodiment of the present invention;

[0060] Figure 2 A schematic structural diagram of a big data watermarking device based on artificial intelligence is provided as an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0062] In one embodiment, Figure 1 As shown, a big data watermarking method based on artificial intelligence is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0063] S101: Perform cross-modal correlation analysis on multiple distributed node data through a federated learning framework to obtain a global feature distribution, divide the distributed node data into real-time streaming data blocks and static data blocks, and perform watermark embedding priority marking on the real-time streaming data blocks and the static data blocks to obtain a marked first real-time streaming data block set and a first static data block set.

[0064] Specifically, federated learning is a distributed machine learning method that combines data from multiple distributed nodes for model training and analysis while protecting data privacy. This data can be in different modalities, such as text, images, audio, and video. Correlation analysis can uncover the inherent connections and global feature distributions between data in different modalities, thereby providing a more comprehensive understanding of the overall characteristics of the data. Based on this, distributed node data can be divided into real-time streaming data blocks and static data blocks. Real-time streaming data blocks refer to continuously generated and updated data, such as network traffic data and real-time sensor monitoring data. They are characterized by high timeliness, large data volume, and continuous data. Static data blocks, on the other hand, are relatively stable and infrequently updated, such as user registration information and historical transaction records. Furthermore, considering the varying importance and urgency of different types of data when embedding watermarks, real-time streaming data blocks and static data blocks can be prioritized for watermark embedding, thereby more effectively allocating watermark embedding resources and improving the efficiency and effectiveness of watermark embedding.

[0065] S102: Based on the generative adversarial network, a first watermark signal and a second watermark signal are generated respectively, which are consistent with the data distribution of the first real-time stream data block set and the first static data block set. The embedding parameters of the first watermark signal and the second watermark signal are dynamically adjusted through the time-sensitive reinforcement learning model to generate a second real-time stream data block set and a second static data block set containing watermarks, and a watermark key bound to the embedding parameters is generated.

[0066] Specifically, a generative adversarial network (GAN) consists of a generator and a discriminator. The generator generates a watermark signal, while the discriminator determines whether the generated watermark signal resembles the distribution of the real data. Through this network, the generator continuously optimizes the generated watermark signal to match its statistical characteristics with the original data set, thereby improving the stealth and imperceptibility of the watermark embedding. Reinforcement learning, on the other hand, is a method in which an intelligent agent learns through trial and error in an environment to maximize its cumulative reward. Time-sensitive reinforcement learning considers the impact of time on rewards. During the watermark embedding process, the time-sensitive reinforcement learning model dynamically adjusts watermark embedding parameters, such as watermark strength, embedding location, and embedding method, based on the real-time performance of the watermark, dynamic changes in the data, and external environmental factors. This allows the model to optimally balance robustness, imperceptibility, and security requirements. Ultimately, the model generates a second set of real-time stream data blocks containing the watermark, and a second set of static data blocks. The watermark key, a crucial basis for subsequent watermark extraction and verification, is bound to the embedding parameters, ensuring the traceability and security of the watermark embedding process.

[0067] S103: By using a preset cross-modal autoencoder, the embedded watermark signal in the second real-time stream data block set and the second static data block set is extracted, and the integrity of the embedded watermark signal is verified by an anti-attack verification model, and the verification result is output.

[0068] Specifically, an autoencoder is an unsupervised neural network model used to learn a compressed representation of data. A cross-modal autoencoder can process data from different modalities and extract common features. This cross-modal autoencoder can accurately extract the watermark signal from watermarked data. Furthermore, an anti-attack verification model can be used to verify the extracted watermark signal by simulating various attack methods, such as noise interference, data compression, and format conversion. This model checks whether the watermark remains intact under these attacks and outputs a verification result. This result can be used to determine the quality of the watermark embedding and the integrity of the data, providing a basis for subsequent watermark updating and recording.

[0069] S104: According to the verification result, the second real-time stream data block set and the second static data block set are dynamically watermarked to generate an updated second real-time stream data block set and an updated second static data block set, and the watermark key and the corresponding embedding parameters are recorded through the blockchain.

[0070] Specifically, if the verification results indicate that the watermark signal has integrity issues or has been attacked, the data needs to be watermarked and re-embedded to ensure copyright protection and data integrity. Dynamic watermarking can generate an updated second set of real-time stream data blocks and an updated second set of static data blocks. Simultaneously, the watermark key and corresponding embedding parameters can be recorded on the blockchain. Blockchain is a distributed ledger technology characterized by decentralization, immutability, and traceability. By recording the watermark key and embedding parameters on the blockchain, critical information can be securely stored and reliably traceable, providing strong support for watermark verification and copyright protection.

[0071] In the above method, a federated learning framework is used to perform cross-modal correlation analysis on data from multiple distributed nodes, and the data types are divided and watermark embedding priorities are marked. This not only achieves the effective integration and structured processing of multi-source heterogeneous data, but also provides scientific priority guidance for subsequent watermark embedding, improving the pertinence and efficiency of watermark processing. Secondly, a watermark signal consistent with the original data distribution is generated based on a generative adversarial network, and the embedding parameters are dynamically adjusted using a time-sensitive reinforcement learning model to generate watermarked data blocks and bound watermark keys. This not only ensures the concealment of the watermark, but also effectively balances the concealment and robustness of the watermark, improving the overall performance of the watermark. Furthermore, by pre-setting a cross-modal autoencoder to extract the watermark signal, the watermark signal can be accurately extracted, and the anti-attack verification model can comprehensively evaluate the integrity of the watermark by simulating multiple attack scenarios, providing a rigorous verification mechanism for the reliability of the watermark. Finally, the data block is dynamically watermarked based on the verification results to ensure the continued effectiveness of the watermark during data updates and version iterations. The tamper-proof nature of blockchain technology provides trusted storage for watermark keys and embedded parameters, enabling traceable management of the watermark throughout its life cycle and improving the security and manageability of the watermark in a distributed environment.

[0072] In one embodiment, cross-modal correlation analysis is performed on multiple distributed node data through a federated learning framework to obtain a global feature distribution, the distributed node data is divided into real-time streaming data blocks and static data blocks, and watermarks are embedded in the real-time streaming data blocks and the static data blocks to mark their priority, thereby obtaining a marked first real-time streaming data block set and a first static data block set, including:

[0073] Each distributed node performs standardized calculation and encryption on the mean, variance, and sparsity of local data to obtain the encrypted features of each distributed node data, and uploads the encrypted features to the central server;

[0074] The central server decrypts the encrypted features and aggregates them to generate a global feature distribution. Based on this global feature distribution, a cross-modal association graph is constructed. The nodes of the cross-modal association graph include image modality nodes, text modality nodes, and time series modality nodes. A shared feature set with redundancy higher than a preset threshold is extracted from the cross-modal association graph and embedded as a watermark.

[0075] Distributed node data is divided into real-time streaming data blocks and static data blocks according to data flow rate and data modality type. Real-time streaming data blocks include image modality and time series modality data, and static data blocks include text modality data.

[0076] Based on the shared feature set and the reinforcement learning model, watermark embedding priorities are assigned to the real-time streaming data blocks and the static data blocks to generate a first real-time streaming data block set and a first static data block set.

[0077] Specifically, the mean reflects the central tendency of the data, the variance indicates the degree of dispersion of the data, and the sparsity describes the proportion of zero or near-zero values in the data. Through standardized calculations, the data features of different nodes can be converted to a unified scale, which is convenient for subsequent comparison and analysis. By encrypting it, the security of the data during transmission can be guaranteed. The central server can then decrypt it and aggregate it taking into account the weights and correlations of the data at different nodes to generate a global feature distribution. Schematically, a cross-modal association graph is a graph structure in which the nodes represent data of different modalities, such as image modal nodes, text modal nodes, and time series modal nodes, and the edges between the nodes represent the association relationship between data of different modalities. This association relationship can be determined by calculating the similarity, correlation or other statistical indicators between the modalities. By constructing a cross-modal association graph through global feature distribution, the mutual connection between data of different modalities can be intuitively displayed, providing a basis for subsequent feature extraction and watermark embedding.

[0078] Specifically, redundancy refers to the degree to which a feature recurs in data of different modalities. By extracting a shared feature set with redundancy higher than a preset threshold from the cross-modal association graph, common features in data of different modalities can be found. This feature has high stability and representativeness in the data and can better carry watermark information. Data flow rate refers to the speed at which data is generated and the frequency of updates. Real-time streaming data usually has a high flow rate and needs to be processed and analyzed in a timely manner, while static data is relatively stable and has a low update frequency. Based on the shared feature set and the reinforcement learning model, watermark embedding priorities can be assigned to real-time streaming data blocks and static data blocks. In principle, when assigning priorities, factors such as the importance of the data, the sensitivity of the data, the frequency of data use, and the security requirements of the data can be comprehensively considered. Finally, a marked first set of real-time streaming data blocks and a first set of static data blocks are generated, which provides clear guidance for subsequent watermark embedding operations.

[0079] In one embodiment, generating a second real-time stream data block set containing a watermark and a second static data block set includes:

[0080] Generate a first watermark signal and a second watermark signal respectively consistent with the data distribution of the first real-time stream data block set and the first static data block set by a generative adversarial network;

[0081] Based on the initial embedding strength parameter and the initial embedding position coordinates, embedding the first watermark signal and the second watermark signal into the first real-time stream data block set and the first static data block set respectively to generate a temporary watermarked data block set;

[0082] A time-sensitive reinforcement learning model is constructed by defining a state space, an action space, and a reward function, wherein the state space includes the data block flow rate, sensitivity, current watermark capacity, and watermark embedding priorities of a temporary watermarked data block set and a first real-time stream data block set and a first static data block set; the action space includes an embedding strength parameter and a set of embedding position coordinates; and the reward function includes a stealth score and a robustness score. The stealth score is calculated by the confidence output value of the discriminator of the generative adversarial network for the temporary watermarked data block set, and the robustness score is calculated by the watermark extraction success rate after simulating a resampling attack on the temporary watermarked data block set.

[0083] The proximal strategy optimization algorithm is used to train the time-sensitive reinforcement learning model. The embedding strength parameters and embedding position coordinate sets of the first and second watermark signals are dynamically adjusted with the watermark embedding priority as the constraint condition to obtain the optimized embedding strength parameters and the optimized embedding position coordinate set.

[0084] According to the optimized embedding strength parameter and the optimized embedding position coordinate set, corresponding watermark signals are embedded in the corresponding coordinate positions of the first real-time stream data block set and the first static data block set to generate a second real-time stream data block set and a second static data block set.

[0085] Specifically, the generator in the generative adversarial network can generate a watermark signal with similar statistical characteristics based on the input noise signal and the characteristics of the target data distribution. For example, if the first set of real-time stream data blocks is image data, the generator will generate a watermark signal with similar texture, color distribution and edge characteristics. Through this network, the generated watermark signal can be better embedded in the original data while maintaining a high degree of imperceptibility and concealment. The initial embedding strength parameter determines the strength of the watermark signal, that is, the prominence of the watermark signal in the original data. The initial embedding position coordinates specify the specific position of the watermark signal in the data block. For example, in image data, the embedding position can be a specific area or pixel position of the image. Subsequently, a temporary set of watermarked data blocks can be generated by embedding the watermark signal under the initial embedding strength parameter and the initial embedding position coordinates.

[0086] Specifically, the time-sensitive reinforcement learning model is a reinforcement learning model that takes time into account and is suitable for decision-making optimization problems in dynamic environments. The data block flow rate reflects the dynamic rate of data change, sensitivity indicates the data's sensitivity to watermark embedding, and the current watermark capacity refers to the maximum number of watermarks that can be embedded in a data block. The watermark embedding priority guides the order and importance of watermark embedding. Furthermore, the reward function is a key metric for evaluating the effectiveness of watermark embedding. A higher confidence output value of the discriminator indicates a more difficult-to-distinguish watermark signal, meaning a higher stealth score and better stealth. The robustness score can be calculated by simulating the watermark extraction success rate after a resampling attack on a set of temporary watermarked data blocks. This resampling attack is a common attack method that attempts to destroy the embedded watermark signal by sampling and reconstructing the data. A high robustness score indicates that the watermark signal can still be successfully extracted after the attack, indicating good robustness.

[0087] Schematically, the proximal policy optimization algorithm is an efficient reinforcement learning algorithm that improves learning efficiency while ensuring policy update stability. This algorithm continuously tries different actions, i.e., embedding parameters, and gradually learns the optimal embedding parameters based on the feedback of the reward function. Finally, based on the optimized embedding strength parameter and the optimized embedding position coordinate set, a second set of real-time stream data blocks and a second set of static data blocks are generated.

[0088] In one embodiment, extracting embedded watermark signals from the second real-time stream data block set and the second static data block set by presetting a cross-modal autoencoder includes:

[0089] By presetting a cross-modal autoencoder, the following operations are performed on each data block in the second real-time stream data block set and the second static data block according to the data type:

[0090] Perform multi-scale convolution operations on image data blocks to generate deep feature maps;

[0091] Use a multi-head attention mechanism to process text data blocks and generate semantic feature vectors;

[0092] Performing bidirectional LSTM processing on the time series data blocks in the second real-time stream data block set and the second static data block set to generate a time series feature sequence;

[0093] The deep feature map, semantic feature vector and temporal feature sequence are mapped to the shared feature space through a preset learnable weight matrix to generate fusion features;

[0094] Based on the fusion features, a watermark mask is generated through a gated convolutional network, and the area in the watermark mask where the element value is greater than the preset mask threshold is determined as the potential watermark area;

[0095] The features corresponding to the potential watermark area are sparsely constrained to extract the watermark feature vector, which is then input into the pre-trained adversarial decoder to output the embedded watermark signal.

[0096] Specifically, multi-scale convolution is an image processing technique in which multiple convolution kernels of different sizes slide across an image, extracting features from different local regions and generating deep feature maps. The multi-head attention mechanism is a powerful text processing technique that can simultaneously focus on multiple locations within a text, capturing long-range dependencies and semantic information. This mechanism divides the text into multiple subspaces, independently calculates attention weights in each subspace, and concatenates and linearly transforms the attention results from these subspaces to generate a semantic feature vector. This vector effectively represents the semantic content of the text. The bidirectional LSTM (Long Short-Term Memory) is a neural network architecture for processing time series data that can simultaneously consider both past and future information. When processing a block of time series data, the bidirectional LSTM scans the data in both forward and backward directions, capturing forward and backward dependencies, respectively, ultimately generating a time series feature sequence. This sequence contains the feature representation of the time series data in the temporal dimension. A pre-set learnable weight matrix learns the correlation and importance between features from different modalities, performing a weighted fusion of the three features to generate a fused feature. The fused feature contains not only the visual information of the image, the semantic information of the text, but also the temporal information of the time series data.

[0097] Specifically, a gated convolutional network is a network structure that can adaptively adjust convolution operations by introducing a gating mechanism to control the activation level of the convolution kernel. In this embodiment, the gated convolutional network uses fused features as input to generate a watermark mask. The watermark mask is a matrix with the same shape as the data block, and its element values represent the probability or intensity of the presence of the watermark signal at the corresponding position. By setting a preset mask threshold, the area in the watermark mask where the element value is greater than the threshold can be determined as a potential watermark area. The potential watermark area is an area that may contain a watermark signal and provides a target area for further watermark extraction. Sparse constraint is a feature extraction technique that limits the sparsity of the feature vector so that only a few elements in the feature vector have large values, thereby extracting the most representative features. For example, the watermark feature vector can be extracted by minimizing the L1 norm of the feature vector. By inputting the watermark feature vector into a pre-trained adversarial decoder, the embedded watermark signal can be recovered by learning the distribution characteristics of the watermark signal.

[0098] In one embodiment, integrity verification is performed on the embedded watermark signal using an anti-attack verification model, and the verification result is output, including:

[0099] The following simulated attack operations are performed on the watermarked data in the second real-time stream data block set and the second static data block set:

[0100] Perform random downsampling and bicubic interpolation restoration on the image data block to generate a perturbed image data block;

[0101] Add Gaussian noise and salt and pepper noise to the time series data block to generate a disturbed time series data block;

[0102] Compress the text data block from 32-bit floating point to 8-bit integer to generate a quantized text data block;

[0103] Based on the embedded watermark signal, the structural similarity index value and peak signal-to-noise ratio value of the perturbed image data block and the perturbed time series data block with the corresponding original data block are calculated, and the watermark signal bit error rate is calculated for the quantized text data block;

[0104] The verification threshold corresponding to the structural similarity index value, peak signal-to-noise ratio value and watermark signal bit error rate is dynamically adjusted according to the data block type to generate the verification result.

[0105] Specifically, random downsampling is an image attack method that reduces image resolution by randomly discarding some pixels. For example, a certain percentage of pixels can be randomly selected from an image for sampling, generating a low-resolution image. The low-resolution image is then restored using bicubic interpolation, generating perturbed image data blocks of the same size as the original image. This process simulates the compression and restoration operations that images may experience during transmission or storage. Furthermore, adding Gaussian noise to time series data can simulate random interference during transmission. Salt and pepper noise is a type of impulse noise that randomly inserts extreme points into time series data to simulate sudden interference that may occur during transmission. Furthermore, 32-bit floating-point representation has higher precision, while 8-bit integer representation has lower precision. By compressing text data from high-precision to low-precision, a certain amount of error can be introduced, thereby testing the integrity and robustness of the embedded watermark signal under this quantization operation.

[0106] Furthermore, the structural similarity index (SSI) is a metric that measures the structural similarity between images or data blocks. It takes into account the brightness, contrast, and structural information of the data blocks, and can more comprehensively reflect the similarity between data blocks. The closer this value is to 1, the more structurally similar the perturbed data block is to the original data block, and the more likely the embedded watermark signal is to maintain integrity under attacks. The peak signal-to-noise ratio (PSNR) is a metric that measures the error between data blocks and can be calculated by calculating the mean squared error between the data blocks. A higher SSI value indicates a smaller error between the perturbed data block and the original data block, and the more likely the embedded watermark signal is to maintain integrity under attacks. The bit error rate (BER) refers to the ratio of erroneous bits to the total number of bits in the watermark signal after quantization. A lower BER indicates a more robust integrity of the embedded watermark signal under quantization. Since different types of attacks have varying degrees of impact on the embedded watermark signal, the verification threshold can be dynamically adjusted based on the data block type. For example, for image data blocks, higher SSI and PSNR thresholds can be set to ensure that the watermark signal maintains high integrity under image attacks. If the measured structural similarity index value and peak signal-to-noise ratio value are higher than the corresponding threshold, and the bit error rate is lower than the corresponding threshold, it can be considered that the embedded watermark signal maintains its integrity under the corresponding attack; otherwise, it is considered that the watermark signal has been damaged. The verification result can be used for subsequent watermark renewal and data integrity verification operations.

[0107] In one embodiment, based on the verification result, dynamically updating the second real-time stream data block set and the second static data block set to generate an updated second real-time stream data block set and an updated second static data block set includes:

[0108] When the verification result meets the preset conditions, the invalid data block index list is output and the watermark update is triggered;

[0109] Calculate the dynamic attenuation factor based on the historical watermark extraction success rate;

[0110] Generate a set of corrected embedding strength parameters and corrected embedding position coordinates based on the current data distribution and dynamic attenuation factor through a time-sensitive reinforcement learning model;

[0111] Based on the invalid data block index list, performing an inverse embedding operation on the invalid data blocks in the second real-time stream data block set and the second static data block set to remove the original watermark signal;

[0112] According to the shared feature set, an anti-attack enhanced watermark signal is generated, and according to the corrected embedding strength parameter and the corrected embedding position coordinate set, the anti-attack enhanced watermark signal is embedded in the corresponding coordinate position of the invalid data block to generate an updated second real-time stream data block set and an updated second static data block set.

[0113] Specifically, the preset conditions can be set according to actual application requirements to determine whether the watermark signal maintains sufficient integrity and robustness under attack. The invalid data block index list contains the index information of all data blocks where the watermark signal has failed, which is used for subsequent watermark reprinting operations. The historical watermark extraction success rate refers to the ratio of the number of successful watermark signal extraction attempts to the total number of attempts during the historical watermark embedding and extraction process. If the historical watermark extraction success rate is high, it means that the current watermark embedding strategy is relatively effective, and the dynamic attenuation factor can be set to a smaller value to maintain the current embedding strength. By designing a dynamic attenuation factor, the watermark embedding strategy can be dynamically adjusted based on historical data, improving the adaptability and reliability of watermark embedding. Subsequently, the time-sensitive reinforcement learning model can select an action based on the current data distribution and the dynamic attenuation factor, and observe the reward after executing the action. Through continuous trial and learning, the optimal modified embedding parameters, namely the modified embedding strength parameters and the modified embedding position coordinate set, are found to adapt to changes in the current data distribution and the historical watermark extraction success rate.

[0114] Specifically, the inverse embedding operation refers to removing the embedded watermark signal from the data block in order to re-embed the new watermark signal. When performing the inverse embedding operation, the invalid data block can be located according to the invalid data block index list, and then the original watermark signal can be removed using the opposite operation of the embedding process. For example, if the watermark signal is embedded using a specific embedding algorithm, the inverse embedding operation can use the corresponding inverse algorithm to remove the watermark signal. Subsequently, an attack-resistant enhanced watermark signal is generated based on the shared feature set, which can ensure that the watermark signal is better integrated with the characteristics of the data block and improve the robustness of the watermark signal. Finally, by correcting the embedding strength parameter and the embedding position coordinate set, the attack-resistant enhanced watermark signal can be embedded in the corresponding coordinate position of the invalid data block, and then an updated second real-time stream data block set and an updated second static data block set are generated. The updated data block set contains a new and more robust watermark signal that can better resist various attacks and protect the copyright and integrity of the data.

[0115] In one embodiment, after obtaining the first watermark signal and the second watermark signal, the method further includes:

[0116] According to the preset privacy budget, a noise signal is generated through the Laplace mechanism, where the noise signal includes image noise, text noise, and time series noise;

[0117] Superimposing the noise signal onto the first watermark signal and the second watermark signal to generate a first privacy-preserving watermark and a second privacy-preserving watermark respectively;

[0118] The privacy embedding parameters of the first privacy-preserving watermark and the second privacy-preserving watermark are dynamically adjusted through a time-sensitive reinforcement learning model, and the privacy embedding parameters are encrypted using the Paillier algorithm to obtain encrypted parameters, which are then distributed and stored in each distributed node.

[0119] During the back-propagation process of federated learning, a random mask matrix is added to the watermark generation gradient matrix of the generative adversarial network and the watermark embedding gradient matrix of the time-sensitive reinforcement learning model to generate obfuscated gradients.

[0120] The obfuscated gradients are uploaded to a central server for aggregation, and the parameters of the generative adversarial network and time-sensitive reinforcement learning model are updated based on the aggregation results.

[0121] Specifically, the privacy budget is used to quantify the degree of privacy protection. The Laplace mechanism is a differential privacy protection method that protects privacy by adding Laplace-distributed noise to the data. Therefore, based on the preset privacy budget, the appropriate noise intensity can be calculated to generate image noise, text noise, and time series noise respectively. By superimposing these noise signals with the watermark signal of the corresponding data type, the privacy protection of the watermark signal can be enhanced. Furthermore, through the time-sensitive reinforcement learning model, the privacy embedding parameters can be dynamically adjusted based on the dynamic characteristics of the current data and privacy requirements. These privacy embedding parameters can include noise intensity, embedding location, embedding method, etc. For example, if the current data flow rate is high, the model can adjust the privacy embedding parameters to increase the noise intensity, thereby better protecting privacy.

[0122] Furthermore, the Paillier algorithm is an additive homomorphic encryption algorithm that can perform addition and multiplication operations on encrypted data. Encrypting the private embedding parameters using the Paillier algorithm ensures the security of the parameters during transmission and storage. The encrypted private embedding parameters are distributed across distributed nodes, with each node storing only a portion of the encrypted parameters. This distributed storage further enhances privacy protection. Even if a node is attacked, the attacker cannot obtain the complete private embedding parameters. Backpropagation in federated learning is a crucial step in model training, used to update model parameters. During backpropagation, a random mask matrix is added to the watermark generation gradient matrix of the generative adversarial network and the watermark embedding gradient matrix of the time-sensitive reinforcement learning model. This random mask matrix is a randomly generated matrix whose element values follow a random distribution. For example, the random mask matrix can be added to the gradient matrix to generate obfuscated gradients. This process effectively hides the true gradient information, preventing attackers from inferring model parameters or data information from the gradient information. Furthermore, the central server can receive obfuscated gradients from each distributed node and perform aggregation operations such as summation or weighted averaging on these gradients. Based on the aggregated gradient information, the parameters of the generative adversarial network and time-sensitive reinforcement learning model are updated. Because the gradient information has been obfuscated, the central server cannot obtain the true gradient information, thus protecting data privacy and the security of model parameters.

[0123] Based on the same inventive concept, Figure 2 As shown, the embodiment of the present application also provides an artificial intelligence-based big data watermarking device 200 for implementing the artificial intelligence-based big data watermarking method involved above. The implementation solution provided by the device is similar to the implementation solution described in the above method, so the specific limitations of one or more artificial intelligence-based big data watermarking device embodiments provided below can be found in the limitations of the various method embodiments above, and will not be repeated here. The device includes:

[0124] The feature analysis and dynamic segmentation module 201 is configured to perform cross-modal correlation analysis based on multiple distributed node data through a federated learning framework to obtain a global feature distribution, divide the distributed node data into real-time streaming data blocks and static data blocks, and perform watermark embedding priority marking on the real-time streaming data blocks and the static data blocks to obtain a marked first set of real-time streaming data blocks and a first set of static data blocks;

[0125] The watermark generation and adaptive embedding module 202 is configured to generate, based on a generative adversarial network, a first watermark signal and a second watermark signal that are consistent with the data distribution of the first set of real-time stream data blocks and the first set of static data blocks, respectively, dynamically adjust embedding parameters of the first watermark signal and the second watermark signal through a time-sensitive reinforcement learning model, generate a second set of real-time stream data blocks containing watermarks and a second set of static data blocks, and generate a watermark key bound to the embedding parameters;

[0126] The watermark extraction and robustness verification module 203 is configured to extract the embedded watermark signal from the second real-time stream data block set and the second static data block set using a preset cross-modal autoencoder, perform integrity verification on the embedded watermark signal using an anti-attack verification model, and output a verification result;

[0127] The dynamic watermark update and evidence storage module 204 is used to dynamically update the watermark of the second real-time stream data block set and the second static data block set based on the verification result, generate an updated second real-time stream data block set and an updated second static data block set, and record the watermark key and corresponding embedding parameters through the blockchain.

[0128] In the above-mentioned device, the feature analysis and dynamic segmentation module 201 performs cross-modal correlation analysis on multiple distributed node data through a federated learning framework, effectively integrating multi-source heterogeneous data and dividing it into real-time streaming data blocks and static data blocks, marking the watermark embedding priority, and providing a structured and orderly data foundation for subsequent watermark processing, avoiding the blindness of watermark embedding and improving the rationality and effectiveness of watermark embedding. The watermark generation and adaptive embedding module 202 generates a watermark signal consistent with the original data distribution based on a generative adversarial network, ensuring a high degree of adaptability between the watermark signal and the original data. At the same time, it uses a time-sensitive reinforcement learning model to dynamically adjust the embedding parameters, and can perform adaptive optimization based on the dynamic characteristics of the data and the watermark requirements, achieving a balance between watermark concealment and robustness, enhancing the watermark's ability to survive in complex data environments, and improving the quality and effectiveness of watermark embedding.

[0129] The watermark extraction and robustness verification module 203 extracts the watermark signal through a preset cross-modal autoencoder, which can accurately extract the watermark signal based on the characteristics of multimodal data. It also uses an anti-attack verification model to perform integrity verification on the embedded watermark signal, simulating multiple attack scenarios, and can comprehensively evaluate the integrity of the watermark, providing a strict verification mechanism for the reliability of the watermark, thereby improving the credibility and security of the watermark. The dynamic watermark update and evidence storage module 204 dynamically updates the watermark of the data block based on the verification results, ensuring the effectiveness of the watermark during data updates and version iterations. In addition, the watermark key and the corresponding embedded parameters are recorded through the blockchain, providing tamper-proof and trusted storage for the watermark key and embedded parameters, realizing the traceability and copyright management of the watermark in a distributed environment, ensuring the security and traceability of the data, and improving the reliability and standardization of watermark management.

[0130] In an exemplary embodiment, the present invention further provides a computer device comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the AI-based big data watermarking method of the present application. A multi-core processor is preferred to improve the system's parallel processing capabilities. The memory provides sufficient temporary storage space to support program execution and data processing. The memory capacity should be large enough to accommodate large amounts of supply information and computing tasks.

[0131] In an exemplary embodiment, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the artificial intelligence-based big data watermarking method of the present application.

[0132] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.

Claims

1. A big data watermarking method based on artificial intelligence, characterized in that: The method comprises: Performing cross-modal correlation analysis on multiple distributed node data through a federated learning framework to obtain a global feature distribution, dividing the distributed node data into real-time streaming data blocks and static data blocks, and embedding watermarks and priority markings on the real-time streaming data blocks and the static data blocks to obtain a marked first real-time streaming data block set and a first static data block set; Based on a generative adversarial network, a first watermark signal and a second watermark signal are generated, each of which is consistent with the data distribution of the first real-time stream data block set and the first static data block set. Embedding parameters of the first watermark signal and the second watermark signal are dynamically adjusted through a time-sensitive reinforcement learning model to generate a second real-time stream data block set and a second static data block set containing watermarks, and a watermark key bound to the embedding parameters is generated. extracting embedded watermark signals from the second real-time stream data block set and the second static data block set by using a preset cross-modal autoencoder, performing integrity verification on the embedded watermark signals using an anti-attack verification model, and outputting a verification result; According to the verification result, the second real-time stream data block set and the second static data block set are dynamically watermarked to generate an updated second real-time stream data block set and an updated second static data block set, and the watermark key and the corresponding embedding parameters are recorded through the blockchain.

2. The method according to claim 1, characterized in that The method performs cross-modal correlation analysis on the data of multiple distributed nodes through a federated learning framework to obtain a global feature distribution, divides the distributed node data into real-time stream data blocks and static data blocks, and performs watermark embedding priority marking on the real-time stream data blocks and the static data blocks to obtain a marked first real-time stream data block set and a first static data block set, including: Each distributed node performs standardized calculation and encryption on the mean, variance, and sparsity of the local data to obtain the encrypted features of the distributed node data, and uploads the encrypted features to the central server; The central server decrypts the encrypted features and aggregates them to generate the global feature distribution; Constructing a cross-modal association graph based on the global feature distribution, wherein the nodes of the cross-modal association graph include image modality nodes, text modality nodes, and time sequence modality nodes, and extracting a shared feature set with redundancy higher than a preset threshold from the cross-modal association graph, and embedding the shared feature set into a carrier as a watermark; Dividing the distributed node data into the real-time stream data block and the static data block according to the data flow rate and the data modality type, the real-time stream data block includes image modality and time series modality data, and the static data block includes text modality data; Based on the shared feature set and the reinforcement learning model, watermark embedding priorities are assigned to the real-time streaming data blocks and the static data blocks to generate the first real-time streaming data block set and the first static data block set.

3. The method according to claim 1, characterized in that The generating of the second real-time stream data block set containing the watermark and the second static data block set comprises: Generate the first watermark signal and the second watermark signal respectively consistent with the data distribution of the first real-time stream data block set and the first static data block set by the generative adversarial network; Based on an initial embedding strength parameter and an initial embedding position coordinate, embedding the first watermark signal and the second watermark signal in the first real-time stream data block set and the first static data block set respectively to generate a temporary watermarked data block set; The time-sensitive reinforcement learning model is constructed by defining a state space, an action space, and a reward function, wherein the state space includes a data block flow rate, sensitivity, and current watermark capacity of the temporary watermarked data block set, and watermark embedding priorities of the first real-time stream data block set and the first static data block set; the action space includes an embedding strength parameter and a set of embedding position coordinates; and the reward function includes a stealth score and a robustness score, wherein the stealth score is calculated by a confidence output value of a discriminator of the generative adversarial network for the temporary watermarked data block set; and the robustness score is calculated by a watermark extraction success rate after a simulated resampling attack on the temporary watermarked data block set; The time-sensitive reinforcement learning model is trained using a proximal strategy optimization algorithm, and the embedding strength parameters and the embedding position coordinate sets of the first watermark signal and the second watermark signal are dynamically adjusted with the watermark embedding priority as a constraint condition to obtain optimized embedding strength parameters and optimized embedding position coordinate sets; According to the optimized embedding strength parameter and the optimized embedding position coordinate set, corresponding watermark signals are embedded in corresponding coordinate positions in the first real-time stream data block set and the first static data block set, respectively, to generate the second real-time stream data block set and the second static data block set.

4. The method according to claim 1, wherein The extracting the embedded watermark signal from the second real-time stream data block set and the second static data block set by presetting a cross-modal autoencoder includes: By using the preset cross-modal autoencoder, the following operations are performed on each data block in the second real-time stream data block set and the second static data block according to the data type: Perform multi-scale convolution operations on image data blocks to generate deep feature maps; Use a multi-head attention mechanism to process text data blocks and generate semantic feature vectors; Performing bidirectional LSTM processing on the time series data blocks in the second real-time stream data block set and the second static data block set to generate a time series feature sequence; Mapping the depth feature map, the semantic feature vector, and the temporal feature sequence to a shared feature space through a preset learnable weight matrix to generate a fusion feature; Based on the fusion features, a watermark mask is generated through a gated convolutional network, and areas in the watermark mask whose element values are greater than a preset mask threshold are determined as potential watermark areas; Sparse constraints are performed on features corresponding to the potential watermark area to extract a watermark feature vector, and the watermark feature vector is input into a pre-trained adversarial decoder to output the embedded watermark signal.

5. The method according to claim 4, characterized in that The performing integrity check on the embedded watermark signal by using an anti-attack verification model and outputting the check result includes: Perform the following simulated attack operations on the watermarked data in the second real-time stream data block set and the second static data block set: Performing random downsampling and bicubic interpolation restoration on the image data block to generate a disturbed image data block; Adding Gaussian noise and salt and pepper noise to the time series data block to generate a disturbed time series data block; Compressing the text data block from 32-bit floating point to 8-bit integer to generate a quantized text data block; Based on the embedded watermark signal, calculating the structural similarity index value and peak signal-to-noise ratio value of the disturbed image data block and the disturbed time series data block with the corresponding original data block, and calculating the watermark signal bit error rate for the quantized text data block; The verification threshold corresponding to the structural similarity index value, the peak signal-to-noise ratio value, and the watermark signal bit error rate is dynamically adjusted according to the data block type to generate the verification result.

6. The method according to claim 1, characterized in that The method of dynamically updating the second real-time stream data block set and the second static data block set according to the verification result to generate an updated second real-time stream data block set and an updated second static data block set includes: When the verification result meets the preset conditions, the invalid data block index list is output and the watermark update is triggered; Calculate the dynamic attenuation factor based on the historical watermark extraction success rate; generating, by the time-sensitive reinforcement learning model, a revised embedding strength parameter and a revised embedding position coordinate set based on the current data distribution and the dynamic attenuation factor; Based on the invalid data block index list, performing an inverse embedding operation on the invalid data blocks in the second real-time stream data block set and the second static data block set to remove the original watermark signal; Based on the shared feature set, an anti-attack enhanced watermark signal is generated, and based on the corrected embedding strength parameter and the corrected embedding position coordinate set, the anti-attack enhanced watermark signal is embedded in the corresponding coordinate position of the invalid data block to generate the updated second real-time stream data block set and the updated second static data block set.

7. The method according to claim 1, characterized in that After obtaining the first watermark signal and the second watermark signal, the method further includes: Generate a noise signal through a Laplace mechanism according to a preset privacy budget, where the noise signal includes image noise, text noise, and time series noise; superimposing the noise signal onto the first watermark signal and the second watermark signal to generate a first privacy-preserving watermark and a second privacy-preserving watermark, respectively; Dynamically adjust the privacy embedding parameters of the first privacy-preserving watermark and the second privacy-preserving watermark through the time-sensitive reinforcement learning model, encrypt the privacy embedding parameters using the Paillier algorithm to obtain encrypted parameters, and distribute and store the encrypted parameters in each of the distributed nodes; During the back-propagation process of federated learning, a random mask matrix is added to the watermark generation gradient matrix of the generative adversarial network and the watermark embedding gradient matrix of the time-sensitive reinforcement learning model to generate a confusion gradient; The confused gradients are uploaded to the central server for aggregation, and the parameters of the generative adversarial network and the time-sensitive reinforcement learning model are updated based on the aggregation results.

8. A big data watermarking device based on artificial intelligence, characterized in that: The device comprises: A feature analysis and dynamic segmentation module is configured to perform cross-modal correlation analysis on multiple distributed node data through a federated learning framework to obtain a global feature distribution, divide the distributed node data into real-time streaming data blocks and static data blocks, and perform watermark embedding priority marking on the real-time streaming data blocks and the static data blocks to obtain a marked first set of real-time streaming data blocks and a first set of static data blocks; a watermark generation and adaptive embedding module, configured to generate, based on a generative adversarial network, a first watermark signal and a second watermark signal, respectively, that are consistent with the data distribution of the first set of real-time stream data blocks and the first set of static data blocks; dynamically adjust embedding parameters of the first watermark signal and the second watermark signal through a time-sensitive reinforcement learning model; generate a second set of real-time stream data blocks and a second set of static data blocks containing watermarks; and generate a watermark key bound to the embedding parameters; a watermark extraction and robustness verification module, configured to extract embedded watermark signals from the second set of real-time stream data blocks and the second set of static data blocks using a preset cross-modal autoencoder, perform integrity verification on the embedded watermark signals using an anti-attack verification model, and output a verification result; A dynamic watermark update and evidence storage module is used to dynamically update the second real-time stream data block set and the second static data block set based on the verification result, generate an updated second real-time stream data block set and an updated second static data block set, and record the watermark key and the corresponding embedding parameters through the blockchain.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Photovoltaic power generation data encryption protection method based on local differential privacy

    CN121659342A