Medical Document Classification via Token Weighting and Highlighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional manual methods for processing medical documents are time-consuming and prone to errors, which can impact patient care and the accuracy of healthcare metrics like HEDIS, necessitating the development of automated processing systems.
Innovation Solution
A method and system utilizing machine learning models to classify medical documents, determine token contribution weights, and visually modify the documents for improved accuracy and efficiency, including the use of techniques like Latent Dirichlet Allocation and Shapley Additive Explanation for classification and highlighting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual methods are used for processing medical documents, then clinicians can review and classify documents, but the process is time-consuming and resource-consuming
Solution Approach 1:
The patent replaces the manual mechanical process of document review with an automated machine learning system. The ML model automatically classifies medical documents by identifying and weighting relevant tokens, eliminating the need for manual clinician review while maintaining high accuracy through sophisticated natural language processing techniques.
Solution Approach 2:
The system enables documents to be processed and classified automatically without human intervention. The machine learning model independently analyzes document content, determines relevance to quality metrics, and generates classifications, making the system self-sufficient for document processing tasks.
2Productivity
If manual methods are used for processing medical documents, then clinicians can identify medical events, but human error leads to mistakes in identification and classification
Solution Approach 1:
The patent replaces human clinicians' manual document analysis with an automated machine learning system. This substitution eliminates human error in identifying and classifying medical events, as the ML model consistently applies the same classification criteria without fatigue or distraction, thereby improving both efficiency and reliability.
3Productivity
If automated processing systems are implemented, then processing speed and accuracy improve, but system complexity increases
Solution Approach 1:
The patent segments the document processing task into distinct components: token extraction, token weighting, quality metric determination, and document classification. This modular approach allows each component to be optimized independently and simplifies the overall system architecture by breaking down the complex processing pipeline into manageable, reusable modules.
4Measurement precision
If comprehensive document analysis is performed to ensure accurate classification, then classification accuracy improves, but processing time increases
Solution Approach 1:
The patent extracts only the most relevant tokens from medical documents using a weighting mechanism that identifies key terms and phrases associated with quality metrics. By focusing analysis on these extracted tokens rather than processing the entire document, the system maintains high classification accuracy while significantly reducing processing time.
Solution Approach 2:
The system performs partial analysis by concentrating computational resources on the most informative portions of documents. The token weighting mechanism identifies and analyzes only the critical elements necessary for accurate classification, avoiding unnecessary processing of redundant information, thus achieving high accuracy with reduced time investment.
Data Source
AI summary
Methods and systems for automatically processing a document may include classifying a document, such as a medical document, as one or more document types based at least in part on one or more machine learning models and one or more tokens extracted from the medical document, determining a token contribution weight of each token towards the classification, modifying the medical document based on the token contribution weights of the one or more tokens, and displaying the modified medical document on a display to a user.


