LLM-Validated Multimodal Graphs for Audio-Text Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing knowledge graphs (KGs) rely heavily on human annotations, which are laborious and expensive, limiting their scalability, and generative large language models (LLMs) suffer from hallucinations and lack of controllable prompts for multimodal applications.

Innovation Solution

Utilize predefined semantic frames and LLMs to automatically construct multimodal knowledge graphs from audio datasets, leveraging LLMs for categorization and relation verification to mitigate hallucinations and reduce the need for human annotations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human annotations are used to construct knowledge graphs, then the quality and accuracy of the knowledge graph is improved, but the labor cost and time consumption increase significantly

Engineering Contradiction:
Improveknowledge graph accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by using LLMs to pre-process and generate initial knowledge graph structures, relationships, and annotations before human review. This allows the system to perform the labor-intensive categorization and relationship extraction work in advance, reducing the time humans need to spend on manual annotation while maintaining quality through subsequent verification steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces LLMs as an intermediary between raw data and final knowledge graphs. The LLM acts as a mediator that performs initial categorization, relationship extraction, and annotation generation, then passes this processed information to humans for verification. This intermediary layer significantly reduces the direct human labor required while maintaining accuracy through the human-in-the-loop verification process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If generative LLMs are used for knowledge graph construction, then productivity is improved, but hallucination effects reduce reliability

Engineering Contradiction:
Improveknowledge graph construction speedVSAvoidfact accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where LLM-generated knowledge graph elements are verified against ground truth data and existing knowledge bases. The system continuously refines its outputs by comparing generated content with verified information, correcting hallucinations, and improving accuracy over time. This feedback loop maintains reliability while preserving the productivity benefits of automated generation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary anti-action by implementing verification and validation steps that proactively counteract potential hallucinations before they propagate through the knowledge graph. The system uses multiple checks, including consistency verification, fact-checking against reliable sources, and cross-validation, to prevent inaccurate information from being incorporated into the final knowledge graph structure.

Inventive Principle:
Principle #9Preliminary anti-action

3Reliability

If manual verification of LLM-generated knowledge graphs is performed, then reliability is improved, but productivity decreases

Engineering Contradiction:
Improveknowledge graph accuracyVSAvoidconstruction speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies partial action by implementing selective verification where only certain high-risk or uncertain portions of LLM-generated knowledge graphs undergo manual review. Rather than verifying every single element, the system uses confidence scores and uncertainty estimation to identify which parts need human attention, thus maintaining high reliability for critical elements while preserving overall productivity through automated processing of lower-risk content.

Inventive Principle:
Principle #16Partial or excessive action

4Adaptability or versatility

If automated methods are used to reduce human annotations, then scalability is improved, but measurement precision may worsen

Engineering Contradiction:
Improveknowledge graph scalabilityVSAvoidannotation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies universality by designing a hybrid system that can adapt its verification depth based on the specific application domain and data characteristics. The same automated LLM-based construction pipeline can be used across different knowledge graph domains, with adjustable levels of human verification applied depending on the criticality and complexity of each domain, thus achieving both scalability and maintained precision across diverse applications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250335705A1System and method for knowledge-based audio-text modeling via automatic multimodal graph construction
Publication Date: 2025.10.30 ROBERT BOSCH GMBH
  • US20250335705A1 patent drawing
  • US20250335705A1 patent drawing
  • US20250335705A1 patent drawing

AI summary

Knowledge-based audio-text modeling via automatic multimodal graph construction is performed. An audio dataset is received, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip of the audio data. Graph nodes of interest are identified from a sematic network, the graph nodes being descriptive of semantics of the knowledge domain of the contents of the audio dataset. A large language model (LLM) is utilized for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph. The extracted knowledge graph is validated utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data.