Text Analysis Apparatus Using Classification Segmentation for Medical Term Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing techniques struggle to accurately extract term expressions and establish correct relationships between terms in medical texts, which are often unstructured and contain descriptions of various organs and diseases.

Innovation Solution

An information processing apparatus that acquires medical texts, classifies attributes of information into fixed units such as sentence, phrase, word, or character units, and performs analysis within each classification using prediction models subjected to machine learning, enabling accurate term extraction and relationship acquisition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text analysis is performed collectively without classification, then processing efficiency is maintained, but analysis accuracy deteriorates due to mixing different classifications

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the text analysis process into multiple segments by classifying text into different categories (e.g., findings, diagnosis, past comparison) before analysis. Each category is analyzed separately using appropriate processing methods, which improves analysis accuracy while managing complexity through structured segmentation.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If term extraction is performed on unstructured medical text without classification, then processing speed is maintained, but term extraction accuracy deteriorates

Engineering Contradiction:
Improveterm extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary classification of medical text into structured categories before conducting term extraction. This preliminary action organizes the unstructured text into manageable segments, enabling more accurate term extraction while reducing the overall processing time through efficient structured handling.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If relationships between terms are acquired without classification, then processing simplicity is maintained, but relationship accuracy deteriorates due to medically incorrect relationships

Engineering Contradiction:
Improverelationship accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies different processing qualities and methods to different parts of the text based on its classification. For example, findings sections are processed differently from diagnosis sections, ensuring that relationship acquisition is contextually appropriate for each medical domain, thereby improving relationship accuracy while managing complexity through localized processing strategies.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12216695B2Information processing apparatus, information processing method, and program for analyzing text
Publication Date: 2025.02.04 FUJIFILM CORP
  • US12216695B2 patent drawing
  • US12216695B2 patent drawing
  • US12216695B2 patent drawing

AI summary

Provided are an information processing apparatus, an information processing method, and a program capable of performing analysis of a text with high accuracy.An information processing apparatus includes one or more processors and one or more memories that store a command executed by the one or more processors. The one or more processors are configured to acquire a text, classify attributes of information described in the text into a fixed unit of the text, analyze the text for each of the same classifications based on a result of the classification, and output a result of the analysis.