Language Model Pre-training Using Structured Database Expressions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional language models struggle with identifying novel information and scaling across domains due to their reliance on large, human-annotated datasets and the loss of structural information during text conversion, limiting their effectiveness in new applications.

Innovation Solution

The proposed solution involves maintaining hierarchical and structural relationships during plain text conversion for training language models, allowing them to leverage structured databases and preserve relational information, enabling improved search results without the need for extensive human annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If language models are trained using large human-annotated datasets, then they achieve effective performance on specific domains, but they become difficult to scale and require extensive manual effort

Engineering Contradiction:
Improvemodel performanceVSAvoidscaling capability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent creates synthetic training data by converting structured database records into natural language expressions that mimic human-annotated data. This copying approach generates large volumes of training examples automatically, eliminating the need for manual annotation while preserving domain-specific knowledge and relationships from the structured data sources.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service by allowing the training data to be generated from existing structured databases without requiring human annotators. The conversion process automatically creates natural language expressions from structured records, making the training data generation process autonomous and scalable to multiple domains simultaneously.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If structured databases are converted to plain text for training, then the data becomes suitable for language model training, but hierarchical and structural relationships are lost

Engineering Contradiction:
Improvetraining data format compatibilityVSAvoidstructural information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent adds a new dimension to plain text by incorporating structured relationship indicators within the natural language expressions. The conversion process embeds hierarchical and relational information as explicit linguistic structures (such as parent-child relationships, categorizations, and connections) within the text, allowing the language model to learn both natural language patterns and structural relationships simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The training data becomes a composite of plain text and structured relationship information. Each natural language expression contains embedded structural elements that preserve the hierarchical and relational properties of the original database records, creating a rich training corpus that combines the benefits of natural language processing with structured data integrity.

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If language models are trained on domain-specific data, then they achieve good results within that domain, but they perform poorly on novel searches or searches outside the domain

Engineering Contradiction:
Improvedomain-specific accuracyVSAvoidcross-domain capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal training approach where structured databases from multiple domains are converted into natural language expressions using consistent conversion rules. This produces training data that teaches the language model domain-specific knowledge while maintaining a unified understanding of relationships and structures across different domains, enabling the model to transfer learning effectively to novel domains.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary conversion of structured data from multiple domains into natural language expressions before training the language model. This pre-processing creates a comprehensive corpus that exposes the model to diverse domain-specific patterns and relationships in advance, preparing it to handle both domain-specific queries and novel searches by recognizing underlying structural patterns across domains.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230394232A1Pre-training language models using natural language expressions extracted from structured databases
Publication Date: 2023.12.07 NVIDIA CORP
  • US20230394232A1 patent drawing
  • US20230394232A1 patent drawing
  • US20230394232A1 patent drawing

AI summary

Systems and methods provide for training a language model on the relationships present in a structural database. Information within a structural data is processed and converted into plain text such that the relationships within the database, such as hierarchical relationships, relations, etc. are maintained and represented in a plain text format. This information may be used as training data for a language model to provide pre-training for one or more domains. The language model may then be leveraged with natural language searching in order to identify results within a search domain response to an input query.