NLP Text Analytics for Audit Document Prioritization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Institutions face difficulties in efficiently reviewing and analyzing diverse documents related to transactions, user feedback, or regulatory filings due to their varied formats and content, making it time-consuming to prioritize and select relevant documents for audit purposes.

Innovation Solution

A system utilizing natural language processing and text analytics, comprising processors that apply text extraction and conversion techniques to generate machine-readable documents, create word embeddings, and establish models to determine document weights and importance scores, allowing for the prioritization and selection of documents based on audit requests.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual review and analysis of diverse documents is performed, then comprehensive audit coverage is achieved, but time consumption and effort increase significantly

Engineering Contradiction:
Improveaudit coverageVSAvoiddocument review time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical review processes with automated natural language processing systems. The NLP system automatically extracts entities, relationships, and insights from diverse document formats (PDFs, Word documents, spreadsheets, images), eliminating the need for manual reading and analysis while maintaining comprehensive audit coverage through systematic processing of all documents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary NLP processing layer between the diverse document sources and the audit analysis system. This intermediary automatically converts various document formats into structured data with extracted entities, relationships, and contextual information, enabling efficient downstream analysis without manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple text extraction and conversion techniques are applied to diverse document formats, then machine-readable document generation is improved, but processing complexity increases

Engineering Contradiction:
Improvedocument format handlingVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal document processing system that handles multiple document formats (PDF, Word, spreadsheets, images) through a single integrated NLP pipeline. The system automatically detects document types and applies appropriate extraction techniques, providing multi-functional capability without requiring separate processing systems for each format.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent dynamically adjusts processing parameters based on document characteristics. The system analyzes document format, size, and content type to automatically select and configure appropriate extraction and conversion techniques, optimizing processing efficiency while adapting to diverse input requirements.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated NLP processing is implemented, then document processing speed increases, but system complexity and computational resources required increase

Engineering Contradiction:
Improvedocument processing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the NLP processing system into modular components: document ingestion module, format detection module, text extraction module, entity recognition module, relationship extraction module, and output generation module. Each module performs a specific function, allowing independent optimization and reducing overall system complexity while maintaining high processing speed.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240386738A1Natural language processing and text analytics for audit testing with documentation prioritization and selection
Publication Date: 2024.11.21 WELLS FARGO BANK NA
  • US20240386738A1 patent drawing
  • US20240386738A1 patent drawing
  • US20240386738A1 patent drawing

AI summary

Disclosed are systems, methods, and computer readable media for natural language processing and text analytics of audit documentation for prioritization and selection. Text extraction and conversion techniques can analyze documents corresponding to an audit request to generate a dataset. A two-layer model can produce word embeddings to reconstruct linguistic contexts of words in the dataset. An embedding layer can map each word, and a classifier layer can generate a similarity score for each word. A three-layer model can determine weights of documents in the dataset. A ranking layer can obtain a document rank value for each document. An initial layer and successive layers can receive feature vectors and document rank values to assign weights to the documents. Based on the document weights and the audit request, the natural language processing and text analytics can determine an audit likelihood for each document to prioritize and select subsets of the documents.