Training Data Generation for Meta-Analysis Literature Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for narrowing down meta-analysis target literatures in the medical field are complex and time-consuming, often including irrelevant literature in training data for machine learning models, which can reduce accuracy.
Innovation Solution
A computer-readable recording medium with a training data generation program that identifies meta-analysis literatures by calculating literature vectors, determining similarity based on feature information, and generating training data only for literatures with a high degree of similarity, thereby distinguishing meta-analysis target literatures from non-targets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual narrowing down of literatures is performed using multiple conditions and steps, then the accuracy of identifying meta-analysis target literatures is improved, but the time and complexity of the process increases significantly
Solution Approach 1:
The patent replaces the manual mechanical process of literature review and filtering with an automated machine learning system. The system uses natural language processing to analyze literature titles, abstracts, and content, automatically identifying meta-analysis target literatures without human intervention in the filtering steps, thus dramatically reducing time while maintaining accuracy through algorithmic consistency.
Solution Approach 2:
The system enables the literature selection process to serve itself by automatically gathering cited literatures from meta-analysis papers, extracting their features, and classifying them without requiring manual curation. The machine learning model autonomously performs the entire workflow from data collection to training data generation, eliminating the need for researchers to manually execute each filtering step.
2Quantity of substance
If all cited literatures in meta-analysis papers are collected for training data, then the quantity of training data increases, but the quality decreases due to inclusion of non-target literatures
Solution Approach 1:
The patent extracts only the relevant meta-analysis target literatures from the pool of all cited literatures by using a machine learning classification system. The system analyzes features of each cited literature and selectively extracts those that meet the criteria for meta-analysis targets, separating them from non-target literatures that would degrade training data quality, thus achieving both sufficient quantity and high quality.
Solution Approach 2:
The system changes the parameters used to evaluate literatures by employing machine learning models that assess multiple features simultaneously (citation context, paper metadata, content characteristics) rather than simple keyword matching. This parameter transformation enables accurate differentiation between target and non-target literatures, ensuring high-quality training data selection.
3Measurement precision
If researchers manually review and narrow down 60 literatures to identify five target literatures, then the precision of training data selection is improved, but the productivity of the research process decreases
Solution Approach 1:
The patent replaces the manual mechanical review process with automated computational analysis. Machine learning models process and evaluate literature features at scale, performing the same precision-selection function that researchers manually achieved but at much higher speed and throughput, thereby restoring productivity without sacrificing precision.
Solution Approach 2:
The system performs preliminary automated filtering and classification of literatures before final selection, pre-processing the data in a way that maintains precision while reducing the manual workload. By preparing and organizing literature data in advance with automated feature extraction and initial filtering, the system enables rapid identification of target literatures, significantly improving overall research productivity.
Data Source
AI summary
A non-transitory computer-readable recording medium storing a training data generation program for causing a computer to execute processing including: identifying, from among meta-analysis literatures stored in a memory, a plurality of meta-analysis literatures in which a first literature is cited; determining a degree of similarity between the plurality of identified meta-analysis literatures based on feature information of the plurality of identified meta-analysis literatures; and in response to the degree of similarity being equal to or higher than a threshold, generating training data for machine learning including the first literature.


