Semi-Structured Data Categorization via Hybrid ML Algorithms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inconsistency in user-entered data across different formats, languages, and structures hinders effective search and categorization, leading to missed relevant information, especially in semi-structured data like recipes where misspellings and synonyms are common.

Innovation Solution

A computer system with data interpretative algorithms, including hybrid Maximum Entropy and LDA models, and term frequency-inverse document frequency, is used to categorize and organize semi-structured data, enabling accurate search and retrieval by selecting and analyzing subsets of data entries based on topical fields and cosine similarity analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is collected from multiple users in different formats and languages, then the quantity and diversity of data increases, but the consistency and reliability of data decreases

Engineering Contradiction:
Improvequantity of dataVSAvoiddata consistency
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent introduces an intermediary processing layer between data collection and storage that automatically standardizes semi-structured data. This intermediary system uses natural language processing and machine learning models to normalize different data formats, resolve synonyms, and correct misspellings, thereby maintaining data consistency while accepting diverse user inputs in various formats and languages.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically adjusts data processing parameters based on the input data characteristics. It uses adaptive algorithms that change the level of standardization applied to different data types, selecting appropriate normalization strategies for each data entry to maintain consistency across the entire dataset while preserving the diversity of original inputs.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If data is stored in semi-structured format to maintain flexibility, then the adaptability of data storage increases, but the precision of data retrieval decreases

Engineering Contradiction:
Improvestorage flexibilityVSAvoidretrieval accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary processing actions to semi-structured data during the ingestion phase. It pre-categorizes, tags, and structures data before storage using machine learning models that predict the most appropriate classification. This preliminary structuring enables precise retrieval operations later while the system maintains the underlying flexibility of semi-structured storage formats.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If traditional search methods are used on unstructured data, then the simplicity of search operation is maintained, but the completeness of search results decreases

Engineering Contradiction:
Improvesearch simplicityVSAvoidsearch completeness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent replaces traditional mechanical keyword-matching search methods with intelligent information retrieval systems that use natural language processing and semantic analysis. These systems substitute simple text comparison with advanced algorithms that understand context, synonyms, and relationships between terms, thereby maintaining user-friendly search operations while dramatically improving the completeness of retrieved results.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9477777B2Method for analyzing and categorizing semi-structured data
Publication Date: 2016.10.25 RAKUTEN GROUP INC
  • US9477777B2 patent drawing
  • US9477777B2 patent drawing
  • US9477777B2 patent drawing

AI summary

A computer system interconnected to a community of users having a data processor input module programmed to receive communications from said users including one or more inputs regarding food recipes and store said inputs in accessible memory. A data processor determining module programmed to access stored data and to apply a data interpretative algorithm to said data to unify and organize disparate data inputs into a cohesive database relating to recipes. Also, a search entry module connected to the recipe database to permit access to the database to support a search algorithm applied to the database.