Text Generation Model for Speech Style Rewriting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Creating a custom language model for speech recognition requires a sufficient matched corpus in terms of domain and speaking style, but manually transcribed texts are expensive and time-consuming, while existing written texts on target domains are abundant but not formatted for natural speech.

Innovation Solution

A computer-implemented method for generating text using a text generation model trained with a topic corpus for domain matching and a rewriting corpus for style matching, incorporating operation tokens for rewriting, to produce texts that match the target domain and speaking style for language model training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manually transcribed text is used for training corpus, then the corpus matches target domain and speaking style, but the preparation cost and time increase significantly

Engineering Contradiction:
Improvecorpus matching qualityVSAvoidpreparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses a text generation model to automatically generate training corpus texts that copy the style and domain characteristics of manually transcribed texts. The model learns from example texts and generates new texts that replicate the desired speaking style and domain knowledge, providing a time-efficient alternative to manual transcription while maintaining corpus quality.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical manual transcription process with an automated text generation model. Instead of requiring human transcribers to manually convert spoken language to text, the system uses a neural network model that automatically generates texts matching the target domain and speaking style, significantly reducing preparation time and cost.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If existing written texts on target domain are used, then the quantity of available text data increases, but the texts are not formatted for natural speech

Engineering Contradiction:
Improvetext data quantityVSAvoidspeech style formatting
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent changes the parameters of existing written texts by processing them through a text generation model that transforms the text style. The model adjusts linguistic parameters such as sentence structure, vocabulary, and discourse patterns to convert formal written texts into natural speech-style texts while maintaining the target domain content.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The text generation model acts as an intermediary between existing written texts and the required speech-style corpus. It processes the written texts through its training and generates intermediate representations that match natural speech patterns, effectively bridging the gap between formal written language and informal spoken language characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11276391B2Generation of matched corpus for language model training
Publication Date: 2022.03.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11276391B2 patent drawing
  • US11276391B2 patent drawing
  • US11276391B2 patent drawing

AI summary

A computer-implemented method for generating a text is disclosed. The method includes obtaining a first text collection matched with a target domain and a second text collection including a plurality of samples, each of which describes rewriting between a first text and a second text that has a style different from the first text. The method also includes training a text generation model with the first text collection and the second text collection, in which the text generation model has, in a vocabulary, one or more operation tokens indicating rewriting. The method further includes outputting a plurality of texts obtained from the text generation model.