LLM Entity Extraction With Tailored Prompts and Explanations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data extraction models are limited to a single document type, require extensive training and expertise, and lack explanations for extracted values, making them inefficient and costly for users without technical knowledge.

Innovation Solution

A system and method for extracting entities from documents using large language models, allowing users to create tailored inputs for machine learning models, including prompt engineering and validation, to accurately extract and explain data from various document types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current data extraction models are used, then extraction accuracy for specific document types is achieved, but the models are limited to single document type and require extensive training data and expertise

Engineering Contradiction:
Improveextraction accuracyVSAvoiddocument type flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by creating a single data extraction platform that can handle multiple document types through a unified interface. The system uses a general-purpose machine learning model that accepts tailored inputs for different document types without requiring separate specialized models for each document type, thus achieving both accuracy and versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system achieves adaptability by changing input parameters rather than retraining the model. Users can provide tailored inputs with different parameters for different document types, and the same base model processes these variations to extract data accurately for each specific document type without requiring model retraining.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If specialized models are developed for each document type, then extraction accuracy improves, but development time and cost increase significantly

Engineering Contradiction:
Improveextraction accuracyVSAvoidmodel development time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-training a single machine learning model on diverse data. This pre-trained model serves as a universal foundation that can be quickly adapted to different document types through tailored inputs, eliminating the need for time-consuming retraining for each specific document type while maintaining high extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of creating entirely new models for each document type, the system uses copying by adapting the pre-trained model through tailored inputs. The same base model is reused for different document types with modified input parameters, significantly reducing development time and resources compared to training separate models for each document type.

Inventive Principle:
Principle #26Copying

3Extent of automation

If current extraction models are used, then data extraction is automated, but no explanations are provided for extracted values making debugging difficult

Engineering Contradiction:
Improveextraction automationVSAvoidextraction explanation
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The system implements feedback by providing explanations for extracted values as part of the output. This feedback mechanism allows users to understand why certain values were extracted, enabling them to verify accuracy and debug issues more effectively while maintaining full automation of the extraction process.

Inventive Principle:
Principle #23Feedback

4Ease of operation

If models are developed without technical expertise, then accessibility improves, but technical accuracy and parameter configuration may be compromised

Engineering Contradiction:
Improveuser accessibilityVSAvoidparameter configuration accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system enables self-service by allowing users to create tailored inputs for different document types without requiring technical expertise. The intuitive interface guides users through the process of configuring extraction parameters for their specific needs, maintaining both accessibility for non-technical users and precision for parameter configuration through structured input options.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250232114A1Document entity extraction platform based on large language models
Publication Date: 2025.07.17 FIDELITY INFORMATION SERVICES LLC
  • US20250232114A1 patent drawing
  • US20250232114A1 patent drawing
  • US20250232114A1 patent drawing

AI summary

Systems and methods are provided for extracting entities from a body of text, using large language models. An example method comprises receiving, from a user, a first input comprising a body of text to be processed for information and a second input comprising a set of at least one element, wherein each of the at least one element comprises information associated with an entity, and wherein each of the at least one entity is data to be extracted from the first input. The example method further comprises creating a tailored input for a machine learning model based on the second input, sending the tailored input to the machine learning model, receiving an output from the machine learning model, processing the output, and providing a processed interactive output to the user.