NLP System for Pharmaceutical Facility Risk Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack a centralized data source for assessing risks in pharmaceutical production facilities and supply chains, making it difficult to collate and analyze information from various documents, leading to blind spots in manufacturing risks for critical products.

Innovation Solution

A natural language processing system that extracts and synthesizes data from raw text documents using a trained machine learning model to classify production issues and generate risk scores for facilities and supply chains, leveraging a pre-defined lexicon of pharmaceutical terminology and MedDRA preferred terms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data is collected from multiple document sources, then information completeness is improved, but data collation complexity increases

Engineering Contradiction:
Improveinformation completenessVSAvoiddata collation complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the complex data collation task into distinct processing stages: document ingestion, text extraction, entity recognition, relationship extraction, and risk scoring. Each stage handles a specific aspect of data processing, making the overall complex task manageable and systematic.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary components including a standardized data model that acts as a mediator between diverse document sources and the risk assessment system. This intermediary layer harmonizes different document formats and structures into a unified representation, reducing collation complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If information is aggregated from scattered sources, then risk assessment accuracy is improved, but processing time increases

Engineering Contradiction:
Improverisk assessment accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing documents during ingestion, extracting and storing key entities and relationships in advance. This preparation work is done before actual risk assessment queries, reducing processing time when assessments are needed while maintaining comprehensive data aggregation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual or mechanical information aggregation with automated natural language processing and machine learning systems. These intelligent systems efficiently process and synthesize information from scattered sources, improving both accuracy and speed compared to traditional manual methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Difficulty of detecting and measuring

If natural language processing is applied to raw text documents, then data extraction capability is improved, but system complexity increases

Engineering Contradiction:
Improvedata extraction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The NLP system is designed to be self-service, with automated entity recognition, relationship extraction, and risk scoring that operate without extensive manual intervention. The system self-adjusts and learns from the data, reducing the operational complexity burden despite the sophisticated processing capabilities.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11544464B1Method for assessing facility risks with natural language processing
Publication Date: 2023.01.03 PHARM 3R LLC
  • US11544464B1 patent drawing
  • US11544464B1 patent drawing
  • US11544464B1 patent drawing

AI summary

The present technology pertains to a method and system for assessing risks associated with facilities, based on using natural language processing. For example, a method can include receiving a natural language input comprising at least one raw text document associated with a facility and generating a plurality of segmented sentences from the raw text documents. The plurality of segmented sentences can be provided as inputs to a machine learning model trained to classify an input segmented sentence over a pre-defined lexicon of pharmaceutical terminology. Each segmented sentence can be classified into one or more classes given by the pre-defined lexicon of pharmaceutical terminology. A secondary classification can be performed for each classified segmented sentence to generate a production issue label based on an analysis of the classified segmented sentence. From the secondary classifications for the classified segmented sentences, at least one production category score for the facility can be generated.