Knowledge Base Generation Using NLP and Prompted LLM Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing knowledge base creation processes are manual, time-consuming, prone to errors, and struggle to keep pace with the growing volume of data, while machine learning models lack the ability to determine actionable solutions.

Innovation Solution

A method and system that autonomously generates a knowledge base by applying natural language processing to content items, determining their structure and meaning, and using a language model to identify problem descriptions, troubleshooting steps, and resolutions, which are then incorporated into a knowledge base article using a template.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual processes are used to create knowledge bases, then analysts can analyze information and create structured articles, but the process is time-consuming and cannot keep pace with growing data volume

Engineering Contradiction:
Improveknowledge base creation speedVSAvoidtime for manual analysis and article creation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of analysts reading, analyzing, and creating knowledge base articles with an automated system using natural language processing and large language models. The system automatically ingests content items, extracts problem descriptions, troubleshooting steps, and resolutions, then generates structured knowledge base articles without human intervention, dramatically increasing productivity and eliminating time loss associated with manual processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically processing content items and generating knowledge base articles without requiring analyst intervention. The automated pipeline independently performs information extraction, article generation, and quality validation, allowing the knowledge base to continuously update itself as new content becomes available, thus resolving the contradiction between manual effort and processing speed

Inventive Principle:
Principle #25Self-service

2Extent of automation

If machine learning models are used to summarize text, then text processing is automated, but the models lack the ability to determine actionable solutions needed for knowledge base articles

Engineering Contradiction:
Improvetext processing automationVSAvoidability to identify actionable troubleshooting solutions
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent introduces prompt engineering as an intermediary layer between the large language model and the knowledge base generation task. Carefully crafted prompts guide the model to extract specific actionable information (problem descriptions, troubleshooting steps, resolutions) from unstructured content, bridging the gap between general text summarization capabilities and the specific requirements for creating reliable, actionable knowledge base articles

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary action by using prompt engineering to pre-configure the large language model with specific extraction tasks before processing content. The prompts are designed in advance to guide the model toward identifying actionable solutions, ensuring that the automation not only processes text but also reliably extracts the specific information needed for knowledge base articles

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If manual knowledge base creation is used, then articles can be created with proper structure, but the process is prone to errors and inconsistencies

Engineering Contradiction:
Improveknowledge base article structure consistencyVSAvoiderror-free article creation
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent replaces the manual mechanical process of structuring knowledge base articles with an automated system that applies consistent templates and validation rules. The system automatically structures extracted information into standardized article formats, eliminating human errors and inconsistencies while maintaining high manufacturing precision through programmatic enforcement of structural requirements

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If analysts manually create knowledge bases, then they can ensure data quality, but they struggle to keep pace with continuously-growing data volume

Engineering Contradiction:
Improvedata quality assuranceVSAvoidability to process growing data volume
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically validating and quality-checking extracted information without requiring analyst review. The automated pipeline includes validation steps that ensure data quality metrics are met, allowing the system to process large volumes of continuously growing data while maintaining measurement precision through programmatic validation rules rather than manual inspection

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260057258A1Systems and methods for generating a knowledge base
Publication Date: 2026.02.26 VIRTUSA CORP
  • US20260057258A1 patent drawing
  • US20260057258A1 patent drawing
  • US20260057258A1 patent drawing

AI summary

Systems and methods for autonomously generating a knowledge base. The methods include receiving at an interface a plurality of content items containing natural language text, and applying natural language processing, using one or more processors executing instructions stored on memory, to the received plurality of content items to determine a structure of each of the plurality of content items and meaning of each of the plurality of content items to determine a prompt for a language model. The methods further include identifying, by supplying the prompt to the language model, at least one of a problem description associated with one or more of the plurality of content items, a troubleshooting step associated with one or more of the plurality of content items, or a resolution associated with one or more of the plurality of content items; and incorporating the identified problem description, troubleshooting step, or resolution into a knowledge base article using a knowledge base template specifying a format for generating the knowledge base article.