Document Classification for Personalized Treatment Guidelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Content repositories like PubMed often lack adequate classification and accessibility, making it difficult for users to efficiently search and utilize large volumes of documents for personalized medicine, especially when institutional licenses or payments are required for full access.

Innovation Solution

A computer system processes documents in content repositories by classifying them into functional and clinical categories, annotating relevant sections, and ranking documents based on query terms using machine learning models to generate personalized treatment guidelines.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If documents are stored in large quantities in content repositories, then the quantity of information increases, but accessibility and ease of searching deteriorate

Engineering Contradiction:
Improvequantity of informationVSAvoidaccessibility
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments the large corpus of documents by classifying each document into functional categories (e.g., genomic data, clinical data, drug data) and clinical categories (e.g., evidence level, study type). This segmentation enables users to search and access specific types of documents more easily without being overwhelmed by the total volume of information in the repository.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary classification and annotation system that acts as a mediator between the raw document corpus and user queries. By adding metadata, annotations, and structured classifications to documents, the system enables efficient searching and retrieval without requiring users to manually browse through all documents.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If documents are classified and annotated using machine learning models, then measurement precision and reliability improve, but device complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies machine learning models in advance to automatically classify and annotate documents before they are accessed by users. This preliminary processing extracts key features, assigns categories, and generates metadata upfront, so that when users search for information, the results are already organized and filtered, reducing the need for complex real-time processing during user interactions.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If full-length research documents are made accessible, then information completeness improves, but loss of substance increases due to licensing requirements

Engineering Contradiction:
Improveinformation accessibilityVSAvoidaccess cost
Core Design Contradiction:
Loss of informationVSLoss of substance

Solution Approach 1:

The patent extracts essential information from full-length research documents by classifying and annotating key findings, data, and conclusions. This extraction creates summarized versions or metadata that capture the most important information without requiring users to access the complete original documents, thereby reducing licensing costs while maintaining information value.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10977292B2Processing documents in content repositories to generate personalized treatment guidelines
Publication Date: 2021.04.13 MERATIVE US LP
  • US10977292B2 patent drawing
  • US10977292B2 patent drawing
  • US10977292B2 patent drawing

AI summary

A computer system processes documents in a content repository. Each document of a plurality of documents is classified into one of a functional category and a clinical category. Each document is annotated using one or more corpora to generate document annotations. Documents satisfying one or more query terms are identified by comparing each query term to the document annotations. The identified documents are ranked based on a determined relevance. Guidelines are produced based on the ranking of the identified documents. Embodiments of the present invention further include a method and program product for processing documents in a content repository in substantially the same manner described above.