Data Embedding Network for Secure Big Data Watermarking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data watermarking techniques for big data often damage or alter the original data, making it unsuitable for machine learning, deep learning, or reinforced learning, and fail to ensure that the marked data is recognized as similar to the original data by computer models.
Innovation Solution
A data embedding network is developed to integrate original data with mark data, generating marked data that is distinct from the original data to humans but recognized as similar by machine learning models, thereby supporting data trading and sharing in big data markets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional watermarking techniques are used to identify data origin and prevent unauthorized distribution, then data identification and protection are improved, but the original data is damaged or altered making it unsuitable for machine learning
Solution Approach 1:
The patent introduces an embedding network as an intermediary that processes both the original data and mark data. This network generates marked data that contains embedded identification information while preserving the essential characteristics needed for machine learning tasks, thus mediating between data protection requirements and data utility requirements
Solution Approach 2:
The embedding network transforms the data by changing its parameters - integrating mark data with original data in a way that modifies the data structure while maintaining its functional properties for machine learning. This allows the data to simultaneously serve identification and learning purposes
2Reliability
If mark data is integrated with original data to generate marked data, then data identification is improved, but the marked data must be recognized as similar to original data by machine learning models
Solution Approach 1:
The embedding network serves as an intermediary that carefully integrates mark data with original data, ensuring that the integration process preserves the essential information content while embedding identification markers. This mediator approach ensures both identification capability and information preservation
Solution Approach 2:
The patent applies mark data locally within the overall data structure rather than uniformly throughout. This allows specific regions or features to carry identification information while other regions maintain their original characteristics for machine learning processing, preserving local quality where needed
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for learning a data embedding network is provided. The method includes steps of: a learning device acquiring and inputting original training data and mark training data into the data embedding network which integrates them and generates marked training data; inputting the marked training data into a learning network which applies a network operation to them and generates 1-st characteristic information, and inputting the original training data into the learning network which applies a network operation to them and generates 2-nd characteristic information; learning the data embedding network such that a data error is minimized, by referring to part of errors referring to the 1-st and the 2-nd characteristic information and errors referring to task specific outputs and their ground truths, and a marked data score is maximized, and learning a discriminator such that a original data score is maximized and the marked data score is minimized.