Automatic audiobook generation method based on multimodal large language model
Automatically generate audio books through a multimodal large language model, which solves the problem of insufficient character and emotional expression in the existing technology, realizes the consistency of character voices and dynamic adjustment of scene emotions, improves the authenticity and production efficiency of audio books, and reduces production costs.
Patent Information
- Application Number
- CN202310894064.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-19
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-07-19
AI Technical Summary
Existing audio book generation technology cannot effectively reflect different roles, scenes and emotions, resulting in poor experience, low production efficiency, high cost, inconsistent quality, and limited expression.
A multimodal large language model is adopted to generate unique voices and speaking styles based on character attributes, adjust the tone, speech speed and volume, combine background sound generation, and optimize the model through training data and user feedback.
It realizes the consistency of character voices and dynamic adjustment of scene emotions, improves the authenticity and quality of audio books, improves production efficiency, and reduces production costs.
Smart Images

Figure CN116821410B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and digital media, and in particular to automatically generating audiobooks using a multimodal large language model. Background Art
[0002] Advantages of Audiobooks: Audiobooks are popular due to many advantages, such as the ability to listen anywhere and at any time, no need for eye strain caused by reading, and the potential to improve foreign language listening skills.
[0003] Limitations of existing audiobooks: Current audiobooks have several limitations.
[0004] 1. Most of them are generated by speech synthesis engines. They use a single voice to read all texts, which cannot reflect different characters, scenes and emotions, resulting in a poor user experience.
[0005] 2. Some popular audiobooks are read aloud by a single host, which can more effectively express the differences in characters, scenes, and emotions. However, these audiobooks are inefficient and costly to produce because the host needs to do a lot of preparation, planning, performance, and recording work.
[0006] 3. The quality varies greatly, which depends largely on the host's talent and understanding of the original work, resulting in great randomness in quality;
[0007] 4. Limited expressiveness because they usually use a single host to voice all the characters, which leads to monotony, and even male and female voices are done by the same host. Summary of the Invention
[0008] This paper proposes a method for automatically generating audiobooks based on a multimodal large language model. This method leverages the multimodal large language model to generate unique voices and styles based on the character's gender, age, personality, and other characteristics, maintaining voice consistency throughout the entire book. It also adjusts the tone, speed, and volume based on the scene and the character's emotions, and realistically generates background sounds.
[0009] 7. The present invention provides the following technical solution: a method for automatically generating audiobooks based on a large-scale language model, comprising the following steps: Step 1: Preparation of training data: First, multimodal training data is obtained from a variety of sources. These sources may include existing movie scripts and dubbing, audiobooks, and manually annotated datasets. For manually annotated datasets, human annotators evaluate the degree of match between the voices generated by the model and their corresponding characters, rate the voices on a predetermined scale, and generate supervised learning labels. This dataset is used to improve the model's ability to generate diverse voice styles and appropriately match these voices to characters.
[0010] Step 2: Model Training: After acquiring training data, the pre-trained large language model is used to train a multimodal large language model. The language model generates unique voices and speaking styles based on character attributes such as gender, age, and personality. The model associates input text with corresponding voice tags and further learns to link character attributes with specific voice styles. The model generates corresponding unique voices by understanding the text context, identifying the character, its attributes, and emotions. Feedback from human annotators is used to iteratively improve the model, strengthening its ability to generate voices that match the character's nature and adjust to the scene and emotion.
[0011] Step 3: Audiobook Generation: After model training is complete, audiobook generation begins from the given text. The model processes the text to identify characters, attributes, and context, and then generates a unique voice for each character based on previously learned knowledge. The model maintains track of context and adjusts the tone, speed, and volume based on the scene and the character's emotions. In addition, the model generates realistic background sounds based on the scene description.
[0012] Step 4: User Feedback and Continuous Optimization: User feedback is an important part of the continuous optimization of the generation process. User feedback on the generated voices, voice consistency, and match with the character can be incorporated into the training data to further improve the model. Therefore, the process forms an iterative loop of generation, feedback, and improvement, thereby improving the overall quality and realism of the automatically generated audiobooks.
[0013] Preferably, the tone of each character remains consistent throughout the audiobook.
[0014] Preferably, the language model adjusts the character's tone, speaking speed, and volume according to different scenes and the character's emotions.
[0015] Preferably, the language model realistically generates background sounds based on the scene description in the text.
[0016] Preferably, the language model is trained on a dataset comprising existing movie scripts, dubbings, and audiobooks.
[0017] Preferably, the language model is further trained on a specially annotated dataset that includes manual evaluations of the match between the generated voices and their corresponding characters, helping the model learn diverse vocal styles and improve the matching of voices to characters.
[0018] The present invention proposes a method for automatically generating audiobooks based on a multimodal large language model, which overcomes the limitations of existing solutions and has the following technical effects:
[0019] This approach leverages the power of large language models to process and understand text data, creating a more nuanced and authentic audiobook experience. The language model can distinguish different characters in the text, along with their gender, age, personality, and other attributes, based on context. It then generates a unique voice and speaking style for each character.
[0020] 2. In addition, the language model is designed to keep the pitch of each character consistent throughout the audiobook. It achieves this by using internal state and parameters that persist during the generation process, ensuring that the same character is always presented with the same voice.
[0021] 3. The language model can also adjust the character's tone, speed and volume according to different scenes and the character's emotions.
[0022] This dynamic adjustment is achieved through the model’s understanding of the context and generation of appropriate emotional responses;
[0023] 4. In addition, the language model can realistically generate background sounds based on the scene description in the text. It achieves this by learning and generating a variety of possible background sounds, creating the most appropriate sound for each scene based on the context. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flowchart of the method for automatically generating audio books based on a multimodal large language model. DETAILED DESCRIPTION
[0025] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0026] A method for automatically generating audiobooks based on a large-scale language model includes the following steps: Step 1: Preparation of training data: This embodiment first obtains multimodal training data from various sources. These sources may include existing movie scripts and dubbing, audiobooks, and manually annotated datasets. For manually annotated datasets, human annotators evaluate the degree of match between the voices generated by the model and their corresponding characters, rate the voices on a predetermined scale, and generate labels for supervised learning. This dataset is used to improve the model's ability to generate diverse voice styles and appropriately match these voices to characters.
[0027] Step 2: Model training: After obtaining the training data, the pre-trained large language model is used to train a multimodal large language model. The language model generates a unique voice and speaking style based on the character's attributes such as gender, age, and personality. The model associates the input text with the corresponding voice label and further learns to associate the character attributes with a specific voice style. The model generates a corresponding unique voice by understanding the text context, identifying the character and its attributes and emotions. Feedback from human annotators is used for iterative improvement of the model to enhance the model's ability to generate voices that are consistent with the character's nature and adjust according to the scene and emotion.
[0028] Audiobook Generation: After model training is complete, it begins generating audiobooks from a given text. The model processes the text to identify characters, attributes, and context. It then generates a unique voice for each character based on this learned knowledge. The model maintains contextual tracking and adjusts the tone, speed, and volume based on the scene and character's emotions. Furthermore, the model generates realistic background sounds based on the scene description.
[0029] User feedback and continuous optimization: User feedback is an important part of the continuous optimization process. User feedback on the generated voices, voice consistency, and character fit can be incorporated into the training data to further improve the model. This process thus forms an iterative cycle of generation, feedback, and improvement, improving the overall quality and realism of the automatically generated audiobooks.
[0030] Each character's voice maintains consistent intonation throughout the audiobook. The language model adjusts the character's tone, speed, and volume based on the scene and the character's emotions. The language model realistically generates background sounds based on scene descriptions in the text. The language model is trained on a dataset containing existing film scripts, dubbing, and audiobooks. The language model is further trained on a specially annotated dataset that includes manual evaluations of the match between the generated voices and their corresponding characters, helping the model learn diverse vocal styles and improve voice-to-character matching.
[0031] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A method for automatically generating audio books based on a large language model, characterized in that: The following steps are involved: Step 1: Training Data Preparation: First, obtain multimodal training data from various sources. These sources can include existing film scripts and dubbing, audiobooks, and manually annotated datasets. For manually annotated datasets, human annotators evaluate the degree to which the voices generated by the model match their corresponding characters, rating the voices on a predetermined scale and generating supervised learning labels. This dataset is used to improve the model's ability to generate diverse voice styles and appropriately match these voices to characters. Step 2: Model Training: After acquiring training data, the pre-trained large language model is used to train a multimodal large language model. The language model generates unique voices and speaking styles based on character attributes such as gender, age, and personality. The model associates input text with corresponding voice tags and further learns to link character attributes with specific voice styles. The model generates corresponding unique voices by understanding the text context, identifying the character, its attributes, and emotions. Feedback from human annotators is used to iteratively improve the model, strengthening its ability to generate voices that match the character's nature and adjust to the scene and emotion. Step 3: Audiobook Generation: After model training is complete, audiobook generation begins from the given text. The model processes the text to identify characters, attributes, and context, and then generates a unique voice for each character based on previously learned knowledge. The model maintains track of context and adjusts the tone, speed, and volume based on the scene and the character's emotions. In addition, the model generates realistic background sounds based on the scene description. Step 4: User Feedback and Continuous Optimization: User feedback is an important part of the continuous optimization of the generation process. User feedback on the generated voices, voice consistency, and match with the character can be incorporated into the training data to further improve the model. Therefore, the process forms an iterative loop of generation, feedback, and improvement, thereby improving the overall quality and realism of the automatically generated audiobooks.
2. The method for automatically generating audio books based on a large language model according to claim 1, characterized in that: The tone of each character remains consistent throughout the audiobook.
3. The method for automatically generating audio books based on a large language model according to claim 1 or 2, characterized in that: The language model adjusts the character's tone, speed, and volume based on different scenarios and the character's emotions.
4. The method for automatically generating audio books based on a large language model according to claim 3, characterized in that: The language model realistically generates background sounds based on the scene description in the text.
5. The method for automatically generating audio books based on a large language model according to claim 4, characterized in that: The language model is trained on a dataset containing existing movie scripts, dubbings, and audiobooks.
6. The method for automatically generating audio books based on a large language model according to claim 5, characterized in that: The language model is further trained on a specially annotated dataset that includes manual evaluations of the match between the generated voices and their corresponding characters, helping the model learn diverse vocal styles and improve voice-to-character matching.
Citation Information
Patent Citations
Automatic audio content generation
CN113628609A
Speech generation method and device based on pre-training language model, equipment and medium
CN116364055A