Synthetic Training Data Generation for Scientific Paper Summarization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid increase in scientific paper publications makes it difficult for researchers to keep up with relevant research, and existing automatic summarization methods for scientific papers lack scalable large-scale training data, often relying on human annotations and generating summaries that are too high-level.

Innovation Solution

The method involves using video or audio recordings of presentations to generate extractive content-based summaries for scientific papers by transcribing the recordings, aligning them with the paper text using unsupervised algorithms, and training machine learning models with the generated summaries, employing a hidden Markov model to select important sentences based on semantic similarity and speaker behavior.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If human annotations are used to create training data for scientific paper summarization, then the quality and accuracy of summaries improve, but the scalability and productivity deteriorate due to the substantial manual effort required

Engineering Contradiction:
Improvesummary qualityVSAvoiddata generation scalability
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent creates synthetic training data by copying and adapting existing paper-summary pairs through transformations such as adding noise, paraphrasing, and modifying summary lengths. This allows generation of large-scale training data without requiring proportional human annotation effort, resolving the contradiction between data quality and scalability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary processing of existing high-quality annotated data to create a foundation for generating additional training examples. By pre-processing available data with automated techniques and then using it to generate more examples, the system maintains quality while improving scalability

Inventive Principle:
Principle #10Preliminary action

2Productivity

If automated text summarization is applied to scientific papers, then the ability to process large volumes of research improves, but the lack of large-scale training data deteriorates model performance

Engineering Contradiction:
Improveresearch processing capacityVSAvoidtraining data volume
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent generates synthetic training data by copying and transforming existing paper-summary pairs, creating artificial examples that expand the training dataset. This approach enables accumulation of large-scale training data without requiring equivalent amounts of manual annotation, thus resolving the contradiction between productivity and data quantity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent varies parameters such as summary length, noise levels, and transformation types when generating synthetic training data. By changing these parameters across multiple generated examples, the system creates diverse training data that improves model robustness while maintaining scalability

Inventive Principle:
Principle #35Parameter changes

3Productivity

If summaries are generated based on paper abstracts, then the generation process is simple and fast, but the summaries become too high-level and lose detailed content

Engineering Contradiction:
Improvesummary generation speedVSAvoidcontent detail
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent introduces synthetic training data as an intermediary between the simple abstract-based approach and the desired detailed summaries. By training models on this intermediate synthetic data that bridges abstract and full-paper content, the system achieves both efficiency and information retention

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11270061B2Automatic generation of training data for scientific paper summarization using videos
Publication Date: 2022.03.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11270061B2 patent drawing
  • US11270061B2 patent drawing
  • US11270061B2 patent drawing

AI summary

Embodiments may provide techniques to generate training data for summarization of complex documents, such as scientific papers, articles, etc., that are scalable to provide large scale training data. For example, in an embodiment, a method may be implemented in a computer system and may comprise collecting a plurality of video and audio recordings of presentations of documents, collecting a plurality of documents corresponding to the video and audio recordings, converting the plurality of video and audio recordings of presentations of documents into transcripts of the plurality of presentations, generating a summary of each document by selecting a plurality of sentences from each document using the transcript of the that document, generating a dataset comprising a plurality of the generated summaries, and training a machine learning model using the generated dataset.