Vaccination Data Extraction via NLP and ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face difficulties in efficiently extracting and verifying vaccination data from unstructured documents, which often lack well-defined formats, leading to resource wastage and incorrect information extraction, especially with handwritten documents.

Innovation Solution

A method utilizing machine learning and natural language processing to process structured and unstructured documents, convert them into a homogeneous format, extract vaccination data, and verify it against a registration authority, thereby reducing resource wastage and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processing methods are used for unstructured vaccination documents, then flexibility in handling various document formats is maintained, but resource wastage and incorrect information extraction occur

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidresource wastage
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent replaces manual mechanical processing of documents with an automated system combining optical character recognition (OCR), natural language processing (NLP), and machine learning models. This substitution eliminates human labor while improving extraction accuracy and reducing resource wastage through automated validation and verification processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary processing layer that includes document formatting standardization, OCR conversion, NLP-based information extraction, and machine learning validation. This intermediary system acts as a mediator between raw unstructured documents and the final verification process, ensuring accurate data extraction while reducing resource consumption through intelligent processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated processing is implemented for unstructured documents, then resource efficiency improves, but handling of heterogeneous document formats becomes difficult

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddocument format compatibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies homogeneity by converting all heterogeneous document formats into a standardized intermediate format through OCR and document structure normalization. This standardization enables the automated processing system to handle diverse vaccination documents (certificates, cards, forms) uniformly while maintaining high processing efficiency and adaptability across different document types.

Inventive Principle:
Principle #33Homogeneity

3Loss of energy

If traditional verification methods are used, then simplicity of the verification process is maintained, but computing resources and human resources are consumed inefficiently

Engineering Contradiction:
Improvecomputing resource consumptionVSAvoidverification system complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-processing documents through OCR, formatting standardization, and structured data extraction before verification. This preliminary processing organizes unstructured information into machine-readable formats, enabling efficient automated verification while reducing computing resource consumption during the actual verification process and minimizing human resource involvement.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4040311A1Utilizing machine learning and natural language processing to extract and verify vaccination data
Publication Date: 2022.08.10 ACCENTURE GLOBAL SOLUTIONS LTD
  • EP4040311A1 patent drawingFigure 1A
  • EP4040311A1 patent drawingFigure 1B
  • EP4040311A1 patent drawingFigure 1C

AI summary

A device may receive, based on a request, document data identifying structured and unstructured documents associated with vaccinations received by users and may perform natural language processing on the document data to generate processed document data. The device may process the processed document data, with a machine learning model, to extract vaccination data from the processed document data and may transcribe the vaccination data into corresponding fields of a data structure. The device may receive, from a user device, a request for vaccination data associated with a user and may retrieve the vaccination data from the corresponding fields of the data structure based on the request. The device may provide the vaccination data, to the user device, to enable verification of the vaccination data.