Medical Fact Verification Using Discrimination Model Pre-training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high requirement for professional knowledge and labeling cost in obtaining large-scale labeled samples for medical fact verification hinders the application of deep learning models in improving medical information extraction, making it difficult to verify medical facts accurately and efficiently.

Innovation Solution

A method and apparatus that utilize a trained discrimination model pre-trained on medical text paragraph pairs and iteratively adjusted using a medical fact sample set with authenticity labeling information to verify medical facts by selecting relevant paragraphs from medical documents, reducing the need for extensive labeled samples and enhancing the accuracy of medical information extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large-scale labeled data is used for training deep learning models in medical fact verification, then the accuracy of medical information extraction is improved, but the labeling cost and professional knowledge requirements increase significantly

Engineering Contradiction:
Improveaccuracy of medical information extractionVSAvoidlabeling cost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-training the discrimination model on large-scale unlabeled medical text paragraph pairs before fine-tuning with small-scale labeled medical fact samples. This preliminary pre-training phase enables the model to learn general medical language patterns and representations without requiring expensive labeled data, thereby reducing labeling costs while maintaining extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses partial action by employing a two-stage training approach where only a small portion of the training process requires expensive labeled data (the fine-tuning stage), while the majority (pre-training stage) uses free unlabeled data. This partial use of labeled data significantly reduces labeling costs compared to traditional approaches that require extensive labeled datasets throughout the entire training process.

Inventive Principle:
Principle #16Partial or excessive action

2Productivity

If deep learning models are applied to medical fact verification, then the efficiency of medical information extraction is improved, but the requirement for large-scale labeled samples increases the complexity of data preparation

Engineering Contradiction:
Improveefficiency of medical information extractionVSAvoidcomplexity of data preparation
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-training the model on unlabeled medical text data before deployment. This preliminary training phase prepares the model with general medical knowledge and language patterns, enabling it to achieve high extraction efficiency with minimal labeled data required for fine-tuning, thereby simplifying the overall data preparation process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by leveraging the pre-trained model's learned representations from unlabeled medical text to handle labeled medical fact verification tasks. The model copies general medical language understanding from the pre-training phase and applies it to the specific verification task, reducing the need for task-specific labeled data and simplifying data preparation complexity.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If traditional information extraction methods are used without deep learning, then the labeling cost is reduced, but the accuracy and efficiency of medical fact verification deteriorates

Engineering Contradiction:
Improvelabeling costVSAvoidaccuracy of medical fact verification
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by pre-training the discrimination model on large-scale unlabeled medical text paragraph pairs, enabling the model to achieve high accuracy with minimal labeled data. This preliminary pre-training compensates for the small amount of labeled data used in fine-tuning, allowing the system to maintain high verification accuracy while keeping labeling costs low compared to traditional methods that require extensive labeled datasets.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11636936B2Method and apparatus for verifying medical fact
Publication Date: 2023.04.25 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11636936B2 patent drawing
  • US11636936B2 patent drawing
  • US11636936B2 patent drawing

AI summary

The present disclosure relates to the field of medical data processing based on natural language processing. Embodiments of the present disclosure disclose a method and apparatus for verifying a medical fact. The method may include: acquiring a description text of the medical fact; selecting a relevant paragraph related to the description text of the medical fact from a medical document; and inputting the description text of the medical fact and the corresponding relevant paragraph into a trained discrimination model for authenticity judgment, to obtain a verification result of the medical fact, the discrimination model being pre-trained based on a medical text paragraph pair extracted from the medical document, and being iteratively adjusted using a medical fact sample set including authenticity labeling information after the pre-training.