Schema-less Data Processing System with Natural Language Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data processing systems face challenges in handling data from multiple sources with different schemas, requiring costly schema definition, updates, and complex logic for data analysis, leading to increased costs and difficulties in merging and searching data across various data sources.
Innovation Solution
A data processing system that uses a schema-independent data structure with natural language extraction and annotation information, allowing for data processing without predefined schema, utilizing a processor and storage device to execute programs that search and merge data based on attribute information and annotation associations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If schema is defined for each data source to enable accurate mechanical reading and saving of data, then data processing accuracy is improved, but human cost, economic cost, and time cost increase
Solution Approach 1:
The patent extracts the schema definition requirement from the data processing system by introducing a schema-less data structure. Data is stored with embedded metadata (data type, format, encoding) within each record itself, eliminating the need for external schema definitions while maintaining processing accuracy through self-describing data structures.
Solution Approach 2:
The patent introduces an intermediary conversion layer that automatically transforms data from various sources into a unified internal representation. This mediator handles format conversion, encoding normalization, and structure alignment without requiring explicit schema definitions, reducing both time cost and human intervention.
2Manufacturing precision
If schema is updated to reflect new data requirements, then data structure accuracy is improved, but the cost of updating schema and past data increases
Solution Approach 1:
The patent implements a dynamic data structure where schemas are not fixed but adapt automatically to new data types and formats. The system evolves its data structure through versioning and incremental updates, allowing new data elements to be incorporated without requiring comprehensive re-schematization of existing data, thereby maintaining accuracy while improving update efficiency.
Solution Approach 2:
The patent incorporates forward-compatible design principles where the data structure anticipates future data types and formats. By预留ing extensibility mechanisms and using flexible typing systems, the system can accommodate new data requirements without requiring structural changes to existing data, eliminating the need for costly updates to past data.
3Loss of information
If different schemas are defined for different data sources to maintain data integrity, then data meaning accuracy is improved, but the difficulty of matching between schemas increases
Solution Approach 1:
The patent implements a universal data structure that can represent multiple data types and formats within a single unified schema. This multi-functional structure accommodates diverse data sources (tabular, hierarchical, semi-structured) using common elements like flexible field types, nested structures, and metadata tags, eliminating the need for complex schema matching while preserving data meaning through context-rich representation.
Solution Approach 2:
The patent uses parameter-based data representation where data characteristics (type, format, encoding, relationships) are expressed as modifiable parameters rather than fixed schema constraints. This allows the system to adapt to different data sources by changing parameters dynamically, maintaining data meaning accuracy through parameter validation while avoiding complex schema matching procedures.
4Quantity of substance
If multiple data sources with different schemas are merged using union or join operations, then data completeness is improved, but the cost of schema association and logic updates increases
Solution Approach 1:
The patent segments data into independent, self-contained records with embedded metadata that describe their structure and relationships. This segmentation allows data from multiple sources to be merged without requiring complex join operations, as each record carries its own structural information. Data completeness is achieved by combining segmented records while maintaining their individual integrity, reducing logic complexity significantly.
Solution Approach 2:
The patent implements a merging mechanism that combines data from multiple sources into a unified structure without requiring pre-established schema associations. The system merges data by aligning records based on their intrinsic metadata and relationships rather than external schema definitions, achieving data completeness while avoiding the complexity of coordinated schema updates across multiple data sources.
Data Source
AI summary
A data processing system executes data processing by accessing a database. The database has a data structure including extraction target data of natural language from a data source and a search target data that is associated with the extraction target data and that can be interpreted in the data processing, and the search target data includes first attribute information of natural language indicating attribute of the extraction target data and annotation information by associating a noun phrase of natural language indicating annotation related to the extraction target data and second attribute information of natural language indicating an attribute of the annotation, the first attribute information is information searched with a first search character string specific to the data processing when an input character string is given, and the annotation information is information searched based on the input character string to the data processing when the input character string is given.


