Training Data Generation for Meta-Analysis Literature Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for narrowing down meta-analysis target literatures in the medical field are complex and time-consuming, often including irrelevant literature in training data for machine learning models, which can reduce accuracy.

Innovation Solution

A computer-readable recording medium with a training data generation program that identifies meta-analysis literatures by calculating literature vectors, determining similarity based on feature information, and generating training data only for literatures with a high degree of similarity, thereby distinguishing meta-analysis target literatures from non-targets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual narrowing down of literatures is performed using multiple conditions and steps, then the accuracy of identifying meta-analysis target literatures is improved, but the time and complexity of the process increases significantly

Engineering Contradiction:
Improveaccuracy of identifying meta_analysis target literaturesVSAvoidtime for narrowing down literatures
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of literature review and filtering with an automated machine learning system. The system uses natural language processing to analyze literature titles, abstracts, and content, automatically identifying meta-analysis target literatures without human intervention in the filtering steps, thus dramatically reducing time while maintaining accuracy through algorithmic consistency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables the literature selection process to serve itself by automatically gathering cited literatures from meta-analysis papers, extracting their features, and classifying them without requiring manual curation. The machine learning model autonomously performs the entire workflow from data collection to training data generation, eliminating the need for researchers to manually execute each filtering step.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If all cited literatures in meta-analysis papers are collected for training data, then the quantity of training data increases, but the quality decreases due to inclusion of non-target literatures

Engineering Contradiction:
Improvequantity of training dataVSAvoidquality of training data
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent extracts only the relevant meta-analysis target literatures from the pool of all cited literatures by using a machine learning classification system. The system analyzes features of each cited literature and selectively extracts those that meet the criteria for meta-analysis targets, separating them from non-target literatures that would degrade training data quality, thus achieving both sufficient quantity and high quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the parameters used to evaluate literatures by employing machine learning models that assess multiple features simultaneously (citation context, paper metadata, content characteristics) rather than simple keyword matching. This parameter transformation enables accurate differentiation between target and non-target literatures, ensuring high-quality training data selection.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If researchers manually review and narrow down 60 literatures to identify five target literatures, then the precision of training data selection is improved, but the productivity of the research process decreases

Engineering Contradiction:
Improveprecision of training data selectionVSAvoidproductivity of literature screening
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical review process with automated computational analysis. Machine learning models process and evaluate literature features at scale, performing the same precision-selection function that researchers manually achieved but at much higher speed and throughput, thereby restoring productivity without sacrificing precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary automated filtering and classification of literatures before final selection, pre-processing the data in a way that maintains precision while reducing the manual workload. By preparing and organizing literature data in advance with automated feature extraction and initial filtering, the system enables rapid identification of target literatures, significantly improving overall research productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11769339B2Computer-readable recording medium storing training data generation program, training data generation method, and training data generation apparatus
Publication Date: 2023.09.26 FUJITSU LTD
  • US11769339B2 patent drawing
  • US11769339B2 patent drawing
  • US11769339B2 patent drawing

AI summary

A non-transitory computer-readable recording medium storing a training data generation program for causing a computer to execute processing including: identifying, from among meta-analysis literatures stored in a memory, a plurality of meta-analysis literatures in which a first literature is cited; determining a degree of similarity between the plurality of identified meta-analysis literatures based on feature information of the plurality of identified meta-analysis literatures; and in response to the degree of similarity being equal to or higher than a threshold, generating training data for machine learning including the first literature.