Multi-Modality Table Data Generation Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing table data generation methods are limited to handling numerical and category data, resulting in insufficient information for characterizing real scenarios, and fail to balance data sharing with data privacy, leading to challenges in data utilization and model training.
Innovation Solution
A method that generates multi-modality table data by encoding text and image data, combining table data with multi-modality learning to improve characterization capability and reduce model configuration costs, while ensuring data privacy through progressive privacy-preserving data publishing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing table data generation methods are used, then the process is simple, but the characterization capability of generated data is insufficient
Solution Approach 1:
The patent merges table data generation with multi-modality learning by combining table data with text and image data. The model integrates multiple data types (table data, text data, image data) into a unified generation framework, allowing the system to leverage complementary information from different modalities to enhance characterization capability while maintaining coherent table structure generation.
Solution Approach 2:
The patent creates a universal data generation model that can handle multiple data types simultaneously. The multi-modality table data generation model serves multiple functions: generating table data, incorporating text descriptions, and integrating image information, making the system adaptable to various data characterization needs without requiring separate specialized models for each data type.
2Loss of information
If multi-modality learning is integrated with table data generation, then the characterization capability improves, but the computational and storage resources increase
Solution Approach 1:
The patent segments the multi-modality data generation process into distinct modules: table data encoding, text data encoding, image data encoding, and a generation model that processes these segmented inputs. This modular segmentation allows the system to handle each data type efficiently and combine them systematically, reducing overall computational burden compared to a monolithic approach.
Solution Approach 2:
The patent implements a nested structure where table data, text data, and image data are encoded into embedding spaces that can be processed hierarchically. The generation model takes these nested embeddings and produces coordinated multi-modality table data, allowing efficient resource utilization by processing data at different levels of abstraction simultaneously.
3Adaptability or versatility
If data sharing is implemented, then information resources are fully utilized, but data privacy protection becomes challenging
Solution Approach 1:
The patent generates synthetic table data that copies the statistical properties, patterns, and relationships of real data without actually sharing the real data itself. The multi-modality generation model creates artificial data instances that preserve essential characteristics for analysis while eliminating sensitive information, allowing data utilization without direct data sharing and thus protecting privacy.
Data Source
AI summary
Embodiments of the present disclosure relate to a method, a device, and a computer program product for generating data. The method includes obtaining a multi-modality embedding by encoding multi-modality data, the multi-modality data comprising text data and image data. The method further includes obtaining a table embedding by encoding table data associated with the multi-modality data. The method further includes generating a condition embedding based on the multi-modality embedding and the table embedding. The method further includes generating multi-modality table data based on the condition embedding. In this way, it is possible to combine table data generation with multi-modality learning, which improves the characterization capability of the generated data and makes the generated data have sufficient information to describe a real scenario. At the same time, it is possible to reduce the model configuration cost and save computational and storage resources.


