Clinical Language Understanding for Rich XHTML Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current medical documentation processes, particularly in electronic health records, face challenges in efficiently extracting and annotating medical facts from free-form clinician notes, which hinders the automation of structured data entry and coding, leading to increased manual effort and potential errors.

Innovation Solution

A clinical language understanding (CLU) system utilizing natural language understanding (NLU) engines to convert richly formatted medical documents into plain text, generate annotations, and apply these annotations back to a tokenized XHTML document, enabling the extraction of medical codes and facts while maintaining rich formatting for user review and editing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual extraction and annotation of medical facts from free-form clinician notes is performed, then accuracy and review capability are maintained, but time consumption and manual workload increase significantly

Engineering Contradiction:
Improveaccuracy of medical fact extractionVSAvoidtime consumption for data extraction
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by automatically extracting and annotating medical facts from free-form clinician notes without requiring manual intervention. The NLU engine processes the text independently, identifying medical codes, entities, and relationships, thereby eliminating the time-consuming manual extraction process while maintaining accuracy through automated intelligent processing

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual extraction process with an automated NLU engine that uses natural language processing techniques. This substitution transforms the manual labor-intensive task into an automated system that processes text efficiently, reducing time consumption while maintaining the precision of fact extraction through computational analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated extraction of medical facts is implemented, then productivity and automation level increase, but system complexity and processing requirements worsen

Engineering Contradiction:
Improvespeed of data extractionVSAvoidcomplexity of NLU processing system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex NLU processing into distinct functional modules: text preprocessing, entity recognition, relationship extraction, and annotation generation. This segmentation allows each component to handle specific tasks independently, making the overall system more manageable and easier to implement while maintaining high productivity through automated processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between the input free-form text and the output structured annotations. This intermediary NLU engine acts as a mediator that translates complex natural language into structured medical data, simplifying the interface between raw text processing and structured data output, thereby reducing perceived system complexity while maintaining automation capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If rich formatting is maintained during conversion to plain text for processing, then information completeness is preserved, but processing efficiency and NLU performance deteriorate

Engineering Contradiction:
Improvecompleteness of formatting informationVSAvoidefficiency of NLU processing
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system extracts only the essential textual content from richly formatted documents, separating the meaningful semantic information from the formatting markup. This extraction creates a clean plain text representation that NLU engines can process efficiently without the interference of HTML tags, CSS styles, or complex formatting, thereby improving processing efficiency while preserving information completeness through accurate text extraction

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary text cleaning and formatting removal before the NLU processing step. This preliminary action prepares the text data in advance by converting rich formatted documents into plain text representations, ensuring that the subsequent NLU processing operates on optimized data that enhances both efficiency and accuracy of medical fact extraction

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9971848B2Rich formatting of annotated clinical documentation, and related methods and apparatus
Publication Date: 2018.05.15 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9971848B2 patent drawing
  • US9971848B2 patent drawing
  • US9971848B2 patent drawing

AI summary

Systems and methods for producing and presenting annotations of clinical documents in a rich format are described, for instance for use with medical billing procedures. An initial XHTML document documenting a medical patient encounter and having rich formatting is used to generate a plain text document. A clinical language understanding system generates annotations, such as medical codes, which are used to annotate the XHTML document. The annotated XHTML document is then presented to a user, thus displaying for the user the annotations while retaining the rich formatting of the initial XHTML document.