LLM Audio Generation Using Role and Emotion Reference Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio generation technologies require large amounts of labeled data for training, leading to high training costs and low efficiency, which reduces the accuracy and efficiency of audio synthesis.

Innovation Solution

An audio generation method using a large language model that parses text to extract role and emotional information, generates target reference texts and audios, and combines them to produce high-quality audio, leveraging pre-set datasets and a pre-trained audio generation model to enhance authenticity and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional emotion classification model and deep learning-based voice synthesis model are used, then audio synthesis can be achieved, but training costs are high and training efficiency is low due to requirement of large amount of labeled data

Engineering Contradiction:
Improveaudio synthesis accuracyVSAvoidtraining efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system pre-processes text content to extract role information and emotional information before audio synthesis. By performing text parsing and information extraction in advance, the system prepares structured data that can be directly used during the synthesis phase, eliminating the need for extensive labeled training data and reducing training time while maintaining synthesis accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediate processing layer between text input and audio output that extracts role information and emotional information. This intermediary step transforms unstructured text into structured parameters that guide the synthesis process, replacing the traditional approach that relies heavily on labeled training data with a method that uses explicit information extraction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If large amount of labeled data is used for training, then model performance can be improved, but training costs increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining costs
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system extracts role information and emotional information directly from the text content itself through automated parsing, rather than relying on externally provided labeled data. The text serves as its own source of training information, eliminating the need for separate labeling processes and reducing the dependency on large annotated datasets while maintaining model performance.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If traditional voice synthesis approach is used, then audio generation can be performed, but the efficiency of audio generation is reduced

Engineering Contradiction:
Improveaudio generation capabilityVSAvoidaudio generation efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent segments the text content into distinct components: role information and emotional information. By dividing the text processing into these specific segments, the system can efficiently extract and utilize each type of information separately during synthesis, improving the overall efficiency of the audio generation process while maintaining operational simplicity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260065891A1Audio generation method and apparatus based on large language model, electronic device, and storage medium
Publication Date: 2026.03.05 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20260065891A1 patent drawing
  • US20260065891A1 patent drawing
  • US20260065891A1 patent drawing

AI summary

A method of audio generation based on a large language model is disclosed, which involves the fields of artificial intelligence such as large language models, natural language processing, deep learning, and audio generation. The method of audio generation based on a large language model comprises: acquiring a text to be processed; parsing the text to be processed using the large language model to obtain role information and emotional information corresponding to the text to be processed; obtaining a target reference text and a target reference audio according to the role information and the emotional information; and generating a target audio corresponding to the text to be processed according to the text to be processed, the target reference text, and the target reference audio.