Data Mesh Structuring Unstructured Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current decentralized data architectures, such as data meshes, face challenges in structuring unstructured data to meet user-specific access requirements, leading to inefficiencies in data access and organization across complex networks.
Innovation Solution
Implementing a data mesh with a data ingestion engine that utilizes artificial intelligence and machine learning to structure data, process unstructured datasets, and distribute them according to their level of structure, enabling users to access data at their desired level of organization through tokenization and encryption, and optimizing data storage across multiple silos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If unstructured data is stored in a decentralized data mesh, then data accessibility and distribution are improved, but data organization and structure are worsened
Solution Approach 1:
The patent segments unstructured data into multiple structured data sets based on different structure levels (e.g., fully structured, semi-structured, unstructured portions). Each segment is stored in appropriate locations within the decentralized mesh, allowing users to access data at their desired level of structure while maintaining overall organization through the segmentation framework.
Solution Approach 2:
The patent introduces an additional dimension of data structure by creating multiple versions of the same data at different structure levels. This dimensional approach allows the system to provide both highly structured data for organized access and unstructured data for flexibility, resolving the contradiction between structure and accessibility.
2Adaptability or versatility
If data is structured to meet specific user access requirements, then user-specific access is improved, but system complexity increases
Solution Approach 1:
The patent implements a dynamic data structuring system where data is automatically structured based on user requirements and access patterns. The system adapts the level of structuring dynamically rather than maintaining fixed structures, allowing it to serve different user needs without requiring manual reconfiguration of the entire system architecture.
Solution Approach 2:
The system performs automatic data structuring and organization without requiring manual intervention for each user's specific needs. The data mesh automatically identifies, structures, and distributes data according to predefined rules and user profiles, reducing the operational complexity while maintaining high adaptability to different user requirements.
3Reliability
If unstructured data is processed and structured, then data quality and searchability are improved, but processing time and resources increase
Solution Approach 1:
The patent applies preliminary structuring actions to unstructured data during the data ingestion phase, creating intermediate structured representations before full processing is required. This preliminary action reduces the processing burden later while still improving data quality and searchability, as the data is partially structured in advance rather than requiring complete restructuring at access time.
Data Source
AI summary
An apparatus and method for using a data mesh to structure unstructured data is provided. The apparatus and method involves transmitting an unstructured dataset from a first node to a data mesh, transmitting receiving the unstructured dataset at a data ingestion engine, analyzing data included in the unstructured dataset, splitting the data included in the dataset into a plurality of data fragments, assigning a level of structure to each data fragment aggregating the data fragments to form a plurality of data segments, transmitting each data segment that is assigned a level of structure to a corresponding data silo, storing each data segment in the corresponding data silo, creating a data structure map of where each data segment is stored, receiving a data request at the data ingestion engine from a second node, using the data structure map to locate the plurality of data segments of the dataset, determining if one or more of the data segments has an assigned level of structure that is less than an assigned level of structure of the second node, and restructuring based on the assigned level of structure of data.


