Schema Augmentation for Exploratory Research via Semantic Proximity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current internet research tools are inadequate for exploratory research, which involves complex information gathering and organization across various sources, as they fail to provide insight into the relevance of collected information and lack schema organization, leading to fragmentation between content collection and organization.

Innovation Solution

A schema augmentation system using deep neural transformer models for natural language understanding (NLU) to process and categorize exploratory research content, determining semantic proximity and providing capabilities for building content-based organizational structures, enabling users to organize and discover relevant information units dynamically during research tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If current internet research tools are used for exploratory research, then basic search functionality is provided, but the tools fail to provide insight into relevance of collected information and lack schema organization

Engineering Contradiction:
Improveloss of semantic meaning and organizationVSAvoidcomplexity of schema augmentation system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary schema augmentation system that sits between content collection and organization functions. This system uses machine learning models to automatically extract schemas from collected content and generate organizational structures, mediating the gap between raw information gathering and structured knowledge representation without requiring direct user intervention in the complex process

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If content collection and organization are performed separately, then each function can be optimized independently, but fragmentation occurs between content collection and organization

Engineering Contradiction:
Improveease of independent function optimizationVSAvoidfragmentation of research context
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges content collection and organization functions by integrating schema extraction and organizational structure generation directly into the content processing pipeline. The system collects content and simultaneously extracts schemas and generates organizational structures from the same processing operations, eliminating the fragmentation between separate collection and organization phases while maintaining the ability to optimize each function independently through modular architecture

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If deep neural transformer models are applied for NLU of text clippings, then semantic understanding and schema extraction are improved, but computational resources and processing time increase

Engineering Contradiction:
Improveprecision of semantic understandingVSAvoidprocessing time for schema extraction
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training the deep neural transformer models on large corpora of text data before deployment. The models are pre-trained to recognize common schemas, entities, and relationships across diverse domains. During actual schema extraction tasks, the pre-trained models can quickly process text clippings with high precision without requiring extensive computation time for learning patterns from scratch, thus reducing processing time while maintaining measurement precision

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240104405A1Schema augmentation system for exploratory research
Publication Date: 2024.03.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240104405A1 patent drawing
  • US20240104405A1 patent drawing
  • US20240104405A1 patent drawing

AI summary

In examples, a schema augmentation system for exploratory research leverages intelligence from a machine learning model to augment such tasks by leveraging intelligence derived from machine learning capabilities. Augmenting tasks include schematization of content, such as information units and groupings of information units. Based on the schematization of such content, semantic proximities for information units are determined. The semantic proximities may be used to identify and present potentially relevant information units, for example to accelerate the exploratory research task at hand. As such, users engaged in consumption of heterogenous content (e.g., across client applications and/or content sources), may receive machine-augmented support to find potential information units. To optimize machine training, user input may be received, such that the system may intelligently augment the user's exploratory research task based on the semantic coherence of the content processed from information units and associated user behavior.