Machine Learning Model for Structured Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high cost and time-consuming process of generating structured data for search engines, particularly for multiple domains like sports, wrestling, football, and basketball, due to the need for manual effort and subscription fees from structured data providers, limits the availability and flexibility of user experiences.

Innovation Solution

Extracting and processing unstructured information from articles and other sources using a trained machine learning model to identify and store relevant entities as structured content, allowing for the generation of user interfaces without relying on a single structured data provider, thus avoiding manual creation and subscription fees.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If structured data is obtained from structured data providers, then data quality and reliability are improved, but cost and dependency increase

Engineering Contradiction:
Improvedata qualityVSAvoiddependency on providers
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system creates structured data copies from unstructured content by training machine learning models on available unstructured data sources. The models learn to extract and structure information independently, producing structured data that mirrors the quality of provider-based data without the dependency on specific providers.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service by automatically generating structured data from unstructured sources through machine learning models. Instead of relying on external providers to create and maintain structured data, the system autonomously performs extraction, validation, and structuring of information from multiple unstructured sources.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual effort is used to create structured content, then data accuracy is improved, but time and expense increase

Engineering Contradiction:
Improvedata accuracyVSAvoidcontent generation speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system replaces manual mechanical processes with automated machine learning models. The models are trained on unstructured data and automatically perform extraction, validation, and structuring tasks that previously required human reviewers, thereby maintaining accuracy while dramatically increasing productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary action by pre-training machine learning models on large volumes of unstructured data before deployment. This pre-training enables the models to automatically recognize patterns and extract structured information with high accuracy without requiring manual intervention for each data point.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple data sources are utilized, then data availability is improved, but processing complexity increases

Engineering Contradiction:
Improvedata availabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements universality by designing machine learning models that can process multiple types of unstructured data sources (articles, transcripts, social media posts, etc.) through a unified framework. The models are trained to handle diverse formats and sources simultaneously, reducing processing complexity while maintaining data availability across multiple sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11797590B2Generating structured data for rich experiences from unstructured data streams
Publication Date: 2023.10.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11797590B2 patent drawing
  • US11797590B2 patent drawing
  • US11797590B2 patent drawing

AI summary

Aspects of the present disclosure are directed to providing a rich content experience based on information received from unstructured content. A plurality of information items may be obtained from a plurality of data source, where each information item includes unstructured content. The plurality of information items may be provided to a trained machine learning model, where the model is trained with training data that includes information items and corresponding labeled entities for a plurality of historical events. In examples, a formatted request may be received, where the formatted request is associated with one or more labeled entities associated with the trained machine learning model. The trained machine learning model may identify multiple entities from the unstructured content based on the formatted request associated with the one or more labeled entities. In examples, each identified entity of the multiple identified entities is stored as structured content responsive to the formatted request.