LLM Document Synthesis for Low-Latency Tacit Knowledge Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for tacit knowledge extraction or generation are manual, tedious, prone to human error, and suffer from high latency, lack of contextuality, and accuracy, especially in mixed scenarios of domain-specific and general-purpose languages.

Innovation Solution

A processor-implemented method and system utilizing high-performance computing to preprocess, tokenize, and vectorize knowledge data, compute similarity measures, and leverage a large language model (LLM) with reinforcement learning and human feedback to generate tacit knowledge aligned to user preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual methods are used for tacit knowledge extraction, then human alignment and contextual understanding are improved, but productivity and time efficiency deteriorate

Engineering Contradiction:
Improvehuman alignmentVSAvoidknowledge generation speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system segments the knowledge generation process into distinct modules: data preprocessing, ontology generation, tokenization, vector embedding, similarity search, and LLM-based generation. This segmentation allows automated processing at scale while maintaining quality control at each stage, resolving the contradiction between manual alignment quality and automated productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing layers (ontology generation, vector embeddings, similarity search) between raw data and final knowledge output. These intermediaries enable automated systems to achieve human-aligned results by structuring data in ways that preserve contextual relationships, thus maintaining reliability while enabling automated high-speed processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If conventional LLM inferencing is used, then simplicity of implementation is improved, but latency and response time deteriorate

Engineering Contradiction:
Improveimplementation simplicityVSAvoidoutput latency
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing data, generating ontologies, creating tokenized representations, and computing vector embeddings before the actual knowledge generation query is executed. This pre-computation stores processed data in optimized formats, enabling rapid retrieval and significantly reducing latency during actual inferencing while maintaining implementation simplicity through modular architecture.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated systems are used for tacit knowledge generation, then productivity is improved, but contextual understanding and human alignment deteriorate

Engineering Contradiction:
Improveautomation efficiencyVSAvoidcontextual accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the system continuously refines its processing based on similarity search results and contextual analysis. The LLM-based framework receives feedback from multiple sources including vector similarity scores, ontology relationships, and preprocessed data quality metrics, enabling automated systems to maintain high contextual accuracy while achieving scalable productivity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system replaces manual mechanical processes with computational mechanisms: manual data review is substituted with automated ontology generation and vector embedding; manual contextual analysis is replaced with similarity search algorithms; and manual knowledge synthesis is substituted with LLM-based generative processes. This substitution maintains or improves contextual accuracy while enabling automated high-speed processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Loss of information

If data from multiple sources is integrated, then knowledge completeness is improved, but data processing complexity and time deteriorate

Engineering Contradiction:
Improveknowledge completenessVSAvoidprocessing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges multiple data sources into a unified processing framework. Different data types (text, structured data, unstructured data) from multiple sources are combined through a common preprocessing pipeline, ontology model, and vector embedding space. This merging approach maintains knowledge completeness from diverse sources while managing complexity through standardized processing stages and modular architecture.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250299018A1Methods and systems for tacit knowledge generation using high performance computing in document synthesis
Publication Date: 2025.09.25 TATA CONSULTANCY SERVICES LTD
  • US20250299018A1 patent drawing
  • US20250299018A1 patent drawing
  • US20250299018A1 patent drawing

AI summary

The present disclosure herein addresses the problem of synthesizing a series of documents and extracting or summarizing meaningful information or content embedded as tacit knowledge in the series of documents. The embodiment of the present disclosure provides a system and method for tacit knowledge generation using large language model (LLM) in document synthesis. The method of the present disclosure performs intelligent document generation orchestrating a generative artificial intelligence solution workflow. In the present disclosure, tacit knowledge of subject matter experts in a knowledge base or in a series of documents is extracted. Further a content capturing the tacit knowledge is generated leveraging a large language models (LLMs) framework as the underlying architecture. The system of the present disclosure is artificial intelligence (AI) accelerated, cloud agnostic, latency defined, and security enabled.