Domain-Specific Ontology Annotations for LLM Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing large language models (LLMs) face challenges in understanding and generating domain-specific technical content due to ambiguities in language and the lack of sufficient training data in specific engineering domains.
Innovation Solution
The method involves generating domain-specific training data by using a domain-specific ontology and processing it through a natural language processing pipeline or a code parsing pipeline, which annotates the data with domain-specific ontology annotations, thereby improving the model's understanding and generation capabilities in specific engineering domains.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If LLMs are trained on general language data, then they can understand common language, but they cannot achieve domain-specific precision and conciseness required in engineering contexts
Solution Approach 1:
The patent applies local quality by creating domain-specific training data with specialized characteristics tailored to engineering contexts. The system processes domain-specific information through pipelines that annotate with domain ontologies, ensuring the training data has the precise quality needed for engineering applications rather than using generic data
Solution Approach 2:
The patent implements preliminary action by pre-processing and structuring domain-specific information before it is used for training. The system includes data processing pipelines that organize, annotate, and prepare domain ontologies and information in advance, transforming raw domain data into structured training formats
2Manufacturing precision
If LLMs use general training data, then training is simpler and faster, but the model generates ambiguous responses that lack conciseness and relevance for specific engineering domains
Solution Approach 1:
The patent applies segmentation by dividing the data processing into distinct pipelines: a natural language processing pipeline for text documents and a code parsing pipeline for source code. Each pipeline segments the processing into specific stages (tokenization, parsing, semantic analysis, ontology annotation) to manage complexity while improving generation precision
Solution Approach 2:
The patent uses domain ontologies as intermediary structures that mediate between raw domain-specific information and the LLM training process. These ontologies serve as structured intermediaries that organize domain knowledge and enable precise, concise generation without requiring direct complex processing of all raw domain data
3Reliability
If LLMs are trained on extensive domain-specific data, then domain expertise improves, but the data processing and structuring becomes more complex and time-consuming
Solution Approach 1:
The patent implements self-service through automated data processing pipelines that independently process, structure, and annotate domain-specific information. The system automatically transforms raw domain data into training-ready formats without requiring manual intervention, reducing time loss while maintaining high domain expertise through comprehensive processing
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a computer-implemented method for generating domain (DMN) specific training data (TRD) for a large language model (LLM). It is proposed, said method comprising the steps of: providing a domain (DMN) specific ontology (OTG) relating to said domain (DMN), providing domain (DMN) specific information (SCI) relating to said domain (DMN), processing said domain (DMN) specific information (SCI) in a data processing-pipeline (DDL) for structuring data for training of said large language model (LLM), wherein said domain (DMN) specific ontology (OTG) is provided as a recognition pattern (RCP) in a step of said data processing-pipeline (DPL), such that said structured training data (STD) comprises domain (DMN) specific ontology (OTG) annotations (ANT).