Data Embedding Network for Secure Big Data Watermarking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data watermarking techniques for big data often damage or alter the original data, making it unsuitable for machine learning, deep learning, or reinforced learning, and fail to ensure that the marked data is recognized as similar to the original data by computer models.

Innovation Solution

A data embedding network is developed to integrate original data with mark data, generating marked data that is distinct from the original data to humans but recognized as similar by machine learning models, thereby supporting data trading and sharing in big data markets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional watermarking techniques are used to identify data origin and prevent unauthorized distribution, then data identification and protection are improved, but the original data is damaged or altered making it unsuitable for machine learning

Engineering Contradiction:
Improvedata identification and protectionVSAvoiddata quality for machine learning
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent introduces an embedding network as an intermediary that processes both the original data and mark data. This network generates marked data that contains embedded identification information while preserving the essential characteristics needed for machine learning tasks, thus mediating between data protection requirements and data utility requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The embedding network transforms the data by changing its parameters - integrating mark data with original data in a way that modifies the data structure while maintaining its functional properties for machine learning. This allows the data to simultaneously serve identification and learning purposes

Inventive Principle:
Principle #35Parameter changes

2Reliability

If mark data is integrated with original data to generate marked data, then data identification is improved, but the marked data must be recognized as similar to original data by machine learning models

Engineering Contradiction:
Improvedata origin identificationVSAvoidinformation similarity for machine learning
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The embedding network serves as an intermediary that carefully integrates mark data with original data, ensuring that the integration process preserves the essential information content while embedding identification markers. This mediator approach ensures both identification capability and information preservation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies mark data locally within the overall data structure rather than uniformly throughout. This allows specific regions or features to carry identification information while other regions maintain their original characteristics for machine learning processing, preserving local quality where needed

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3834107B1Method for training and testing data embedding network to generate marked data by integrating original data with mark data, and training device and testing device using the same
Publication Date: 2025.05.21 DEEPING SOURCE INC
  • EP3834107B1 patent drawingFigure 1
  • EP3834107B1 patent drawingFigure 2
  • EP3834107B1 patent drawingFigure 3

AI summary

A method for learning a data embedding network is provided. The method includes steps of: a learning device acquiring and inputting original training data and mark training data into the data embedding network which integrates them and generates marked training data; inputting the marked training data into a learning network which applies a network operation to them and generates 1-st characteristic information, and inputting the original training data into the learning network which applies a network operation to them and generates 2-nd characteristic information; learning the data embedding network such that a data error is minimized, by referring to part of errors referring to the 1-st and the 2-nd characteristic information and errors referring to task specific outputs and their ground truths, and a marked data score is maximized, and learning a discriminator such that a original data score is maximized and the marked data score is minimized.