Domain-Specific Ontology Annotations for LLM Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing large language models (LLMs) face challenges in understanding and generating domain-specific technical content due to ambiguities in language and the lack of sufficient training data in specific engineering domains.

Innovation Solution

The method involves generating domain-specific training data by using a domain-specific ontology and processing it through a natural language processing pipeline or a code parsing pipeline, which annotates the data with domain-specific ontology annotations, thereby improving the model's understanding and generation capabilities in specific engineering domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If LLMs are trained on general language data, then they can understand common language, but they cannot achieve domain-specific precision and conciseness required in engineering contexts

Engineering Contradiction:
Improvedomain-specific precisionVSAvoidtraining data quantity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies local quality by creating domain-specific training data with specialized characteristics tailored to engineering contexts. The system processes domain-specific information through pipelines that annotate with domain ontologies, ensuring the training data has the precise quality needed for engineering applications rather than using generic data

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements preliminary action by pre-processing and structuring domain-specific information before it is used for training. The system includes data processing pipelines that organize, annotate, and prepare domain ontologies and information in advance, transforming raw domain data into structured training formats

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If LLMs use general training data, then training is simpler and faster, but the model generates ambiguous responses that lack conciseness and relevance for specific engineering domains

Engineering Contradiction:
Improvegeneration concisenessVSAvoiddata processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the data processing into distinct pipelines: a natural language processing pipeline for text documents and a code parsing pipeline for source code. Each pipeline segments the processing into specific stages (tokenization, parsing, semantic analysis, ontology annotation) to manage complexity while improving generation precision

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses domain ontologies as intermediary structures that mediate between raw domain-specific information and the LLM training process. These ontologies serve as structured intermediaries that organize domain knowledge and enable precise, concise generation without requiring direct complex processing of all raw domain data

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If LLMs are trained on extensive domain-specific data, then domain expertise improves, but the data processing and structuring becomes more complex and time-consuming

Engineering Contradiction:
Improvedomain expertiseVSAvoiddata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements self-service through automated data processing pipelines that independently process, structure, and annotate domain-specific information. The system automatically transforms raw domain data into training-ready formats without requiring manual intervention, reducing time loss while maintaining high domain expertise through comprehensive processing

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4553697A1Computer-implemented method for generating domain specific training data for a large language model
Publication Date: 2025.05.14 SIEMENS IND SOFTWARE NV
  • EP4553697A1 patent drawingFigure 1
  • EP4553697A1 patent drawingFigure 2
  • EP4553697A1 patent drawingFigure 3

AI summary

The invention relates to a computer-implemented method for generating domain (DMN) specific training data (TRD) for a large language model (LLM). It is proposed, said method comprising the steps of: providing a domain (DMN) specific ontology (OTG) relating to said domain (DMN), providing domain (DMN) specific information (SCI) relating to said domain (DMN), processing said domain (DMN) specific information (SCI) in a data processing-pipeline (DDL) for structuring data for training of said large language model (LLM), wherein said domain (DMN) specific ontology (OTG) is provided as a recognition pattern (RCP) in a step of said data processing-pipeline (DPL), such that said structured training data (STD) comprises domain (DMN) specific ontology (OTG) annotations (ANT).