Schema-less Data Processing System with Natural Language Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data processing systems face challenges in handling data from multiple sources with different schemas, requiring costly schema definition, updates, and complex logic for data analysis, leading to increased costs and difficulties in merging and searching data across various data sources.

Innovation Solution

A data processing system that uses a schema-independent data structure with natural language extraction and annotation information, allowing for data processing without predefined schema, utilizing a processor and storage device to execute programs that search and merge data based on attribute information and annotation associations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If schema is defined for each data source to enable accurate mechanical reading and saving of data, then data processing accuracy is improved, but human cost, economic cost, and time cost increase

Engineering Contradiction:
Improvedata processing accuracyVSAvoidtime cost
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts the schema definition requirement from the data processing system by introducing a schema-less data structure. Data is stored with embedded metadata (data type, format, encoding) within each record itself, eliminating the need for external schema definitions while maintaining processing accuracy through self-describing data structures.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary conversion layer that automatically transforms data from various sources into a unified internal representation. This mediator handles format conversion, encoding normalization, and structure alignment without requiring explicit schema definitions, reducing both time cost and human intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If schema is updated to reflect new data requirements, then data structure accuracy is improved, but the cost of updating schema and past data increases

Engineering Contradiction:
Improvedata structure accuracyVSAvoidupdate efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements a dynamic data structure where schemas are not fixed but adapt automatically to new data types and formats. The system evolves its data structure through versioning and incremental updates, allowing new data elements to be incorporated without requiring comprehensive re-schematization of existing data, thereby maintaining accuracy while improving update efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates forward-compatible design principles where the data structure anticipates future data types and formats. By预留ing extensibility mechanisms and using flexible typing systems, the system can accommodate new data requirements without requiring structural changes to existing data, eliminating the need for costly updates to past data.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If different schemas are defined for different data sources to maintain data integrity, then data meaning accuracy is improved, but the difficulty of matching between schemas increases

Engineering Contradiction:
Improvedata meaning accuracyVSAvoidschema matching complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements a universal data structure that can represent multiple data types and formats within a single unified schema. This multi-functional structure accommodates diverse data sources (tabular, hierarchical, semi-structured) using common elements like flexible field types, nested structures, and metadata tags, eliminating the need for complex schema matching while preserving data meaning through context-rich representation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter-based data representation where data characteristics (type, format, encoding, relationships) are expressed as modifiable parameters rather than fixed schema constraints. This allows the system to adapt to different data sources by changing parameters dynamically, maintaining data meaning accuracy through parameter validation while avoiding complex schema matching procedures.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If multiple data sources with different schemas are merged using union or join operations, then data completeness is improved, but the cost of schema association and logic updates increases

Engineering Contradiction:
Improvedata completenessVSAvoidlogic complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments data into independent, self-contained records with embedded metadata that describe their structure and relationships. This segmentation allows data from multiple sources to be merged without requiring complex join operations, as each record carries its own structural information. Data completeness is achieved by combining segmented records while maintaining their individual integrity, reducing logic complexity significantly.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a merging mechanism that combines data from multiple sources into a unified structure without requiring pre-established schema associations. The system merges data by aligning records based on their intrinsic metadata and relationships rather than external schema definitions, achieving data completeness while avoiding the complexity of coordinated schema updates across multiple data sources.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10896227B2Data processing system, data processing method, and data structure
Publication Date: 2021.01.19 HITACHI LTD
  • US10896227B2 patent drawing
  • US10896227B2 patent drawing
  • US10896227B2 patent drawing

AI summary

A data processing system executes data processing by accessing a database. The database has a data structure including extraction target data of natural language from a data source and a search target data that is associated with the extraction target data and that can be interpreted in the data processing, and the search target data includes first attribute information of natural language indicating attribute of the extraction target data and annotation information by associating a noun phrase of natural language indicating annotation related to the extraction target data and second attribute information of natural language indicating an attribute of the annotation, the first attribute information is information searched with a first search character string specific to the data processing when an input character string is given, and the annotation information is information searched based on the input character string to the data processing when the input character string is given.