A text-to-audio method, device, equipment, storage medium and program product

CN122551764APending Publication Date: 2026-08-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,上述文本转音频的方式,不同角色的音色上较为单一,无法实现角色个性化的语音

Benefits of technology

[0025]通过从目标文本中提取角色信息列表,实现了将散落在目标文本中的角色信息进行结构化存储,避免了人工逐句标注的低效,通过基于角色信息列表识别出每个文本片段对应的角色,实现了精确的角色分配,提高了文本分析的针对性和有效性,通过基于每个角色的角色信息确定每个角色的音色信息,实现了角色的个性化的音色表现,从而增强后续语音合成的真实感和沉浸感,提升具有多角色的目标文本的音频转换效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551764A_ABST
    Figure CN122551764A_ABST
Patent Text Reader

Abstract

The application provides a text-to-audio method, device, equipment, storage medium and program product; the method comprises the following steps: dividing a target text into multiple text segments; extracting a character information list from the target text, wherein the character information list comprises character information of each character in the target text; identifying the character in each text segment based on the character information list; determining the timbre information of each character based on the character information of each character; performing speech synthesis based on each text segment and the timbre information of the character in each text segment to obtain an audio segment corresponding to each text segment; and combining the audio segment corresponding to each text segment to obtain an audio corresponding to the target text. Through the application, text-to-audio of multi-character text can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence technology, and more particularly to a text-to-audio method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the development of internet technology, the demand for text-to-audio conversion is growing, as evidenced by the large number of users listening to novels as audiobooks. Among related technologies, artificial neural networks can be used to convert text into audio. However, the aforementioned text-to-audio methods offer a relatively uniform timbre for different characters, failing to achieve personalized voices. Summary of the Invention

[0003] This application provides a text-to-audio method, apparatus, device, storage medium, and program product that can realize text-to-audio conversion for multi-role text.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a text-to-audio conversion method, the method comprising:

[0006] Segment the target text into multiple text fragments;

[0007] Extract a list of character information from the target text, wherein the list of character information includes character information for each character in the target text;

[0008] Based on the list of character information, identify the character in each of the text fragments;

[0009] The timbre information of each character is determined based on the character information of each character.

[0010] Speech synthesis is performed based on each text segment and the timbre information of the character in each text segment to obtain an audio segment corresponding to each text segment;

[0011] The audio segments corresponding to each text segment are combined to obtain the audio corresponding to the target text.

[0012] This application provides a text-to-audio conversion device, including:

[0013] The text segmentation module is used to divide target text into multiple text fragments;

[0014] A character information extraction module is used to extract a character information list from the target text, wherein the character information list includes the character information of each character in the target text;

[0015] The character recognition module is used to identify the character in each text fragment based on the character information list;

[0016] A timbre allocation module is used to determine the timbre information of each character based on the character information of each character.

[0017] The speech synthesis module is used to perform speech synthesis based on each text segment and the timbre information of the character in each text segment to obtain an audio segment corresponding to each text segment;

[0018] The speech synthesis module is further configured to combine the audio segments corresponding to each text segment to obtain the audio corresponding to the target text.

[0019] This application provides an electronic device, the electronic device comprising:

[0020] Memory is used to store executable instructions or computer programs.

[0021] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the text-to-audio method provided in the embodiments of this application.

[0022] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the text-to-audio method provided in this application when executed by a processor.

[0023] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the text-to-audio method provided in this application.

[0024] The embodiments of this application have the following beneficial effects:

[0025] By extracting a list of character information from the target text, the scattered character information in the target text is stored in a structured manner, avoiding the inefficiency of manual sentence-by-sentence annotation. By identifying the character corresponding to each text segment based on the character information list, accurate character assignment is achieved, improving the targeting and effectiveness of text analysis. By determining the timbre information of each character based on the character information, personalized timbre performance of each character is achieved, thereby enhancing the realism and immersion of subsequent speech synthesis and improving the audio conversion effect of target texts with multiple characters. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the architecture of the text-to-audio system provided in the embodiments of this application;

[0027] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0028] Figure 3 This is a schematic diagram of the text-to-audio method provided in the embodiments of this application;

[0029] Figure 4A This is a first flowchart illustrating the text-to-audio method provided in this application embodiment;

[0030] Figure 4B This is a second flowchart illustrating the text-to-audio method provided in this application embodiment;

[0031] Figure 4C This is a schematic diagram of the third process of the text-to-audio method provided in the embodiments of this application;

[0032] Figure 4D This is a schematic diagram of the fourth process of the text-to-audio method provided in the embodiments of this application;

[0033] Figure 4E This is a schematic diagram of the fifth process of the text-to-audio method provided in the embodiments of this application;

[0034] Figure 4F This is a schematic diagram of the sixth process of the text-to-audio method provided in the embodiments of this application;

[0035] Figure 4G This is a schematic diagram of the seventh process of the text-to-audio method provided in the embodiments of this application;

[0036] Figure 4H This is the eighth flowchart of the text-to-audio method provided in the embodiments of this application;

[0037] Figure 4I This is a ninth flowchart illustrating the text-to-audio method provided in this application embodiment;

[0038] Figure 4J This is a schematic diagram of the tenth process of the text-to-audio method provided in the embodiments of this application;

[0039] Figure 4K This is a schematic diagram of the eleventh step of the text-to-audio method provided in the embodiments of this application;

[0040] Figure 4L This is a schematic diagram of the twelfth step of the text-to-audio method provided in the embodiments of this application;

[0041] Figure 4M This is a schematic diagram of the thirteenth step of the text-to-audio method provided in the embodiments of this application;

[0042] Figure 5This is a schematic diagram of an optional network structure for the speech synthesis model provided in this application embodiment;

[0043] Figure 6 This is a flowchart illustrating the text-to-audio method provided in this application embodiment in an audiobook production scenario;

[0044] Figure 7 This is a schematic diagram of the text-to-audio system interface provided in the embodiments of this application.

[0045] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0047] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0048] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0049] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0050] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0051] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0052] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0053] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0054] 2) Large Language Models (LLMs) are large-scale language models designed to understand and generate human language. They are trained on massive amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their enormous scale, containing billions of parameters that help them learn complex patterns in language data. They are typically based on deep learning architectures. LLMs refer to deep learning models trained on massive amounts of text data, containing billions or even more parameters. They can be used to generate natural language text and understand its meaning. Through training, the models learn the statistical regularities and semantic relationships of language to build a vast language knowledge base, thereby simulating human language understanding and generation capabilities. LLMs have the following characteristics:

[0055] Learning ability: Through training with massive amounts of text data, large language models can learn rich language knowledge and expressions, including grammar, semantics and common expression habits.

[0056] Pattern recognition: Large language models can identify common text patterns and semantic relationships, such as co-occurrence relationships between words, logical structure of sentences, and semantic roles.

[0057] Contextual understanding: Large language models can capture contextual information in text, understand the influence of previous text on subsequent text, and generate corresponding responses based on the context.

[0058] Generative capabilities: Large language models can generate relevant natural language text based on input information, including answering questions, generating articles, and engaging in dialogue.

[0059] Resolving ambiguity: Despite the existence of polysemy and ambiguity in language, large language models resolve ambiguity through contextual information and linguistic rules, providing more accurate and appropriate text generation or understanding.

[0060] Large language models have a wide range of applications, including intelligent customer service, intelligent question answering, natural language generation, advertising recommendation, and games. They can improve the efficiency and accuracy of human-computer interaction and enhance the user experience.

[0061] 3) Transformer: A temporal model based on self-attention. In the encoder part, it can effectively encode temporal information, and its processing capability for temporal information is far superior to Long Short-Term Memory (LSTM) networks, while also being faster. It is widely used in natural language processing, computer vision, machine translation, speech recognition, and other fields.

[0062] 4) Workflow is an effective means of automating business process management using computer technology. It consists of a series of interconnected nodes, each corresponding to a specific stage or task in the business process and possessing independent processing logic and functions. The nodes flow in an orderly manner according to preset rules and conditions, enabling the automatic transfer of documents, information, or tasks among multiple entities to achieve predetermined business goals and ensure the continuity, efficiency, and accuracy of the business process.

[0063] 5) Large Language Model (LLM) workflow refers to deeply embedding a large language model into a traditional workflow framework. In this architecture, each workflow node operates closely around and invokes the large language model. Each node first preprocesses the input data, converting it into a format and content suitable for the large language model's input requirements. Then, it inputs specific instructions to the large language model, customized according to the node's function and task objectives within the workflow. Upon receiving the instructions, the large language model, based on its vast parameter system, deep neural network structure, and pre-trained language knowledge and semantic understanding capabilities, performs in-depth processing and analysis of the input data, executing complex language tasks such as text generation, semantic understanding, and logical judgment. After processing, the large language model outputs results, which the nodes then post-process, including format conversion, information extraction, and result integration, to generate output content that meets the needs of the next workflow node or the final business requirements. This output is then passed to the next workflow node or completes the entire workflow process according to preset rules. Throughout the process, each node forms an efficient interactive loop with the large language model, giving full play to the language processing expertise of the large language model in different business process stages, and realizing the intelligent advancement and automated execution of business processes.

[0064] 6) Natural Language Processing (NLP) is an important branch of computer science and artificial intelligence. It focuses on the interaction between computers and human natural language. Its main goal is to enable computers to understand, generate, and process human language. The understanding process includes syntactic, semantic, and pragmatic analysis of text; the generation process involves converting the computer's internal representation into a natural language expression; and processing encompasses various tasks such as text classification, sentiment analysis, and machine translation. It utilizes various methods, including statistics, machine learning, and deep learning, to build language models. These models are used to preprocess text and extract features to uncover potential information within the text, providing core technical support for numerous practical applications such as text summarization, intelligent customer service, and text correction, enabling computers to handle natural language-related matters in a more intelligent way.

[0065] 7) Named Entity Recognition (NER) is a key task in Natural Language Processing (NLP). It primarily involves locating and classifying various entities within text. Its purpose is to accurately identify entities with specific meanings from unstructured text data. These entities include, but are not limited to, names of people, places, organizations, dates, currencies, etc. The recognition system then labels these entities according to predefined categories, thereby extracting key information units from the text. This provides foundational support for subsequent, more advanced NLP tasks, such as information retrieval, question answering systems, and knowledge graph construction, helping computers better understand the semantic content of text.

[0066] 8) Deep Neural Networks (DNNs) are a crucial component of machine learning. They process data by constructing neural networks with multiple hidden layers, aiming to automatically extract effective features from complex data to build models and perform tasks such as classification and prediction. DNNs encompass various architectures and are widely used in fields such as natural language processing, speech recognition, and computer vision, helping to uncover patterns and rules behind data. DNNs include, but are not limited to, Multilayer Perceptrons (MLPs), Convolutional Neural Networks (CNNs), and Recurrent Neural Networks (RNNs). In natural language processing, they can be used for tasks such as machine translation and text generation; in speech recognition, they can help identify the text content corresponding to speech signals; and in computer vision, they are used in many applications such as image classification and object detection, providing powerful technical support for various complex data processing tasks and helping systems better uncover patterns and rules behind data.

[0067] 9) Text-to-Speech (TTS) systems convert text into speech, their core function being to transform input text into natural and fluent speech. This conversion is achieved primarily through text analysis, speech synthesis, and prosody generation. First, the input text is analyzed for syntax and semantics. Then, either unit-based concatenation or parametric synthesis is used to synthesize speech. Unit-based concatenation relies on pre-recorded speech units, while parametric synthesis calculates parameters based on an acoustic model to generate speech. Finally, appropriate prosody, such as intonation, stress, and rhythm, is generated based on the text's semantic information, resulting in more natural speech. TTS systems have wide applications, commonly found in voice assistants, audiobook production, intelligent navigation systems, and language learning aids, providing services such as voice broadcasting, reading, navigation prompts, and pronunciation learning.

[0068] 10) Automatic Speech Recognition (ASR) systems are systems that convert speech into text. They primarily employ deep learning methods for training, utilizing a large amount of labeled speech-text data. The system analyzes speech content and converts it into text by learning the correlation between speech signal features and language patterns. During the conversion process, inferences are made based on acoustic features and linguistic knowledge, continuously optimizing parameters. After training with a large amount of data, it possesses excellent speech recognition capabilities and can be applied to various fields such as voice assistants and intelligent customer service for tasks such as voice command recognition and speech-to-text transcription.

[0069] 11) A timbre library, or timbre database, stores timbre information, which includes timbre features and audio that matches the timbre features. Timbre features may include a unique identifier for the timbre, the age of the timbre, gender, timbre tags, and other data.

[0070] 12) Character information refers to a series of data and information included in a detailed description of a character. Specifically, character information may include the following aspects:

[0071] Character Name: The name of a character, used to identify and distinguish different characters.

[0072] Character Introduction: A brief introduction to the character, which may include the character's backstory, motivations, goals, and role in the story.

[0073] Character traits: include at least one of the following:

[0074] Character gender: The character's biological sex, such as male, female, or other.

[0075] Character age range: The age range of the character, such as youth, middle age, old age, etc., or it can be a specific age value.

[0076] Character traits: The character's personality traits, such as cheerful, introverted, brave, cunning, etc.

[0077] Role identity: A character's social identity or profession, such as student, doctor, king, etc.

[0078] 13) Timbre information refers to the reference audio (audio data of timbre) and the timbre characteristics of the reference audio. Timbre characteristics may include basic information of the reference audio, such as timbre identifier, gender, age, etc.; timbre description tags, i.e. keywords used to describe timbre, such as "clear", "deep", "soft", "hoarse", "magnetic", "lively", etc.; pitch range (Hz), such as fundamental frequency range (e.g., male average 85-180Hz, female average 165-255Hz); emotion adaptation tags, such as emotions suitable for expression, such as "majestic", "gentle", "cheerful", "sad", etc.; applicable scene tags, such as "animation protagonist", "narrator", "villain", "children's book", etc.

[0079] With the development of internet technology, the demand for converting text to audio is growing, as evidenced by the large number of users listening to novels as audiobooks. Among related technologies, artificial neural networks can be used to convert novels into audio, thus achieving text-to-audio conversion. However, the aforementioned text-to-audio methods are relatively limited.

[0080] This application provides a text-to-audio method, apparatus, device, computer-readable storage medium, and computer program product, capable of converting multi-role text to audio. The following describes exemplary applications of the electronic devices provided in this application. These electronic devices can be implemented as various types of terminal devices such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. The following describes exemplary applications when the electronic device is implemented as a server.

[0081] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the text-to-audio system provided in the embodiments of this application. Figure 1 The system involves server 100, terminal device 200, and network 300. Terminal device 200 is connected to server 100 through network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.

[0082] In some embodiments, the present application embodiments can be implemented collaboratively by a server and a terminal device. For example, terminal device 200 sends target text to server 100, server 100 receives the target text, generates audio corresponding to the target text using the text-to-audio method provided in the present application embodiments, and sends the audio to terminal device 200.

[0083] The text-to-audio system provided in this application can be applied to various scenarios that require text-to-audio conversion, such as virtual assistants and chatbots, podcasts and audiobooks, accessibility assistance, etc. Examples are given below.

[0084] 1) Virtual Assistants and Chatbots: For example, a terminal device (such as a smart speaker) sends a request to a server to synthesize a voice response containing specific information. Upon receiving the request, the server uses a text-to-audio system to generate the voice and sends the voice data back to the terminal device. The terminal device plays the received voice, and the user hears the voice response of a virtual assistant or chatbot.

[0085] 2) Accessible Information Services: For example, a terminal device (such as a mobile phone or computer) recognizes a user's need to read electronic documents, web pages, or e-books and sends text content to the server. The server processes the text content, generates speech through a text-to-audio system, and sends it back to the terminal device. The terminal device then plays the speech for visually impaired users, enabling accessible reading of information.

[0086] 3) Customer Service and Customer Support: For example, terminal devices (such as mobile phones, computers, etc.) interact with the server through an Interactive Voice Response (IVR) system. Users input requests or select service options. Based on the user input, the server generates a corresponding voice response using a text-to-audio system and plays it through the IVR system. Users receive automated voice services via telephone or online channels.

[0087] 4) Podcasts and Audiobooks: For example, a terminal device uploads text content to a server, requesting the generation of corresponding audio narration. The server processes the text content, generates audio through a text-to-audio system, and then sends the audio data back to the terminal device. The terminal device plays the audio, providing the user with the auditory experience of a podcast or audiobook.

[0088] 5) In-vehicle systems: For example, a terminal device (such as an in-vehicle terminal equipped with an in-vehicle infotainment system) requests a server to read aloud navigation instructions, news, or weather forecasts. The server generates audio using a text-to-audio system and sends the audio data to the terminal device. The in-vehicle system plays the audio, providing the driver with real-time voice information.

[0089] 6) News Broadcasting: For example, a terminal device (such as a news organization's broadcasting system) sends news text to a server, requesting automatic voice broadcasting. The server processes the news text, generates audio using a text-to-audio system, and then sends the audio data back to the terminal device. The terminal device plays the audio, achieving automated news broadcasting.

[0090] 7) Games and Entertainment: For example, a terminal device (such as a mobile phone with a game client installed) sends in-game dialogue text or prompts to the server, requesting the generation of voice. The server generates the corresponding audio using a text-to-audio system and sends the audio data to the terminal device. The game system plays the audio, providing players with an immersive gaming experience.

[0091] In other embodiments, the embodiments of this application can be implemented by a terminal device alone. The terminal device 200 generates audio corresponding to the target text using the text-to-audio method provided in the embodiments of this application.

[0092] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The electronic device 400 shown can be either the server 100 or the terminal device 200 mentioned above. Figure 2The illustrated electronic device 400 includes at least one processor 410, a memory 430, and at least one network interface 420. The various components of the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0093] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0094] The memory 430 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 430 may optionally include one or more storage devices physically located away from the processor 410.

[0095] The memory 430 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 430 described in this application embodiment is intended to include any suitable type of memory.

[0096] In some embodiments, memory 430 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0097] Operating system 431 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0098] The network communication module 432 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0099] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A text-to-audio device 433 stored in memory 430 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a text segmentation module 4331, a character information extraction module 4332, a character recognition module 4333, a timbre allocation module 4334, and a speech synthesis module 4335. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0100] In some embodiments, the terminal device or server can implement the text-to-audio method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run; or they can be applets that can be embedded in any APP (e.g., voice assistant clients, e-book reader clients, etc.), i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0101] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the text-to-audio method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0102] The text-to-audio method provided in this application will be described below with reference to exemplary applications and implementations of the server provided in the embodiments of this application, using the server as the execution subject. See also Figure 3 , Figure 3This is a schematic diagram of the text-to-audio method provided in this application embodiment. After obtaining the target text, text segmentation is performed to divide the target text into text segments and text chapters (here, the segmentation can be performed separately, or the target text can be divided into text segments and then composed of text segments to form text chapters; this application embodiment does not impose any restrictions); based on the text chapters, role information is extracted to obtain a role information list; role recognition is performed using the text segments and the role information list to obtain the role corresponding to each text segment; timbre allocation is performed based on the role to obtain the timbre information corresponding to each role; speech synthesis is performed using each text segment and the corresponding timbre information to obtain the audio segment corresponding to each text segment; quality detection is performed on each audio segment; in response to a failed detection, the process is transferred to the speech synthesis node to re-synthesize the audio segment; in response to a passed detection, the audio segments corresponding to each text segment are combined into the audio corresponding to the target text.

[0103] See Figure 4A , Figure 4A This is a first flowchart illustrating the text-to-speech method provided in this application embodiment, which will be combined with... Figure 4A The steps shown are explained.

[0104] In step 101, the target text is segmented into multiple text fragments.

[0105] Here, the target text can include various types, which are determined according to the application scenario and requirements. For example, the target text can be books, articles, news content, etc.

[0106] In some embodiments, before segmenting the target text into multiple text fragments, the following processes may be performed: cleaning the target text to obtain cleaned target text; dividing the cleaned target text into multiple line texts according to the newline characters in the cleaned target text; storing the target text as structured target text according to each line text and the line number of each line text; and proceeding to the process of segmenting the target text into multiple text fragments based on the structured target text.

[0107] For example, text parsing technology is used to read the target text line by line, and each line is scanned and filtered to remove or replace characters that meet the cleaning criteria (such as spaces, emojis, etc.) or replace lowercase characters with uppercase characters. At the same time, metadata such as line numbers and paragraph information are recorded for backtracking and correlation in subsequent processing.

[0108] In some embodiments, the structured target text is segmented according to a pre-set text segmentation rule to obtain multiple text fragments, wherein each text fragment corresponds to a text fragment information, which includes a text fragment number and the line number of the text included in the text fragment.

[0109] For example, pre-defined text segmentation rules can be based on text length, sentence structure, specific delimiters (such as commas, periods, etc.), keywords, or phrases. These rules are applied to structured target text to determine how to segment the text. The segmentation process may include: reading lines of text and their line numbers from the structured target text; segmenting the text according to the pre-defined text segmentation rules to obtain multiple text fragments; and creating an information record for each text fragment, which may include a text fragment number (i.e., assigning each text fragment a unique identifier) ​​and a list of line numbers (i.e., recording the line numbers of all lines of text contained in the text fragment).

[0110] It should be noted that a text fragment can be a paragraph or a sentence in the target text, and the embodiments of this application do not limit the specific granularity of text fragment division.

[0111] See examples Figure 7 In response to a trigger operation on the target text selection control 001, the specified text is retrieved from the text database as the target text (e.g., the text of a novel with a specified novel name is selected as the target text), and then the process of splitting the target text into multiple text fragments is executed.

[0112] In step 102, a list of character information is extracted from the target text, wherein the list of character information includes the character information of each character in the target text.

[0113] It should be noted that the character information list is a set of structured character data extracted from the target text, which includes information such as character name, character characteristics (such as gender, age, identity, etc.) and character introduction.

[0114] In some embodiments, see Figure 4B , Figure 4A Step 102 shown can be implemented through steps 1021 to 1024, which will be explained in detail below.

[0115] In step 1021, a first prompt word is obtained, wherein the first prompt word indicates the extraction of a list of character information from the target text.

[0116] In some embodiments, the first prompt word is pre-set and used to prompt the language model to extract a list of character information from the target text. For example, when the target text is a novel, the first prompt word may be expressed as "Please extract character information from the currently input text. The character information includes the character's name, character characteristics (such as age range, gender, identity, personality, etc.), and character introduction".

[0117] Here, the first cue word is a pre-defined instruction used to guide the language model to extract specific information from the target text. In the role information extraction task, the first cue word explicitly instructs the model to identify and output role information from the text, such as name, characteristics (age, gender, identity, personality, etc.), and a brief introduction. The role of the first cue word can include: 1) task definition, such as clarifying that the task the language model needs to perform is role information extraction; 2) output specification, such as specifying the output format and content, such as role name, characteristics, etc.; 3) contextual guidance, such as helping the language model understand the role-related information in the target text. In other words, the first cue word is an instruction that guides the language model to perform the role information extraction task, ensuring that the output conforms to the expected format and content.

[0118] For example, the input to a language model can include target text (such as a passage from a novel) and a first prompt: "Please extract character information from the currently input text. Character information includes character name, character characteristics (such as age range, gender, identity, personality, etc.), and character description." The output of the language model is a list of character information, for example: Name: Zhang San; Characteristics: Age 30-35, male, company manager, decisive personality; Description: a key figure in the company, with strong decision-making ability.

[0119] In step 1022, the target text is divided into multiple text sections, wherein the granularity of the text section division is greater than that of the text fragment division.

[0120] In some embodiments, paragraphs of the target text can be combined into combined paragraphs using segmentation rules such as double quote matching and length limitation processing, and then multiple combined paragraphs can form a text chapter.

[0121] For example, suppose the target text contains the following paragraphs: Paragraph 1: (“The weather is so nice today,” Xiaoming said, “Let’s go to the park.”); Paragraph 2: (Xiaohong nodded and smiled, saying, “Okay, I’ll bring my picnic basket.”); Paragraph 3: (They walked together to the park, the sun shining on the grass.); Paragraph 4: (“It’s so beautiful here,” Xiaoming exclaimed, “Let’s find a place to sit down.”); Paragraph 5: (Xiaohong opened her picnic basket and took out sandwiches and juice.). The segmentation rule is to merge paragraphs containing dialogue into a single combined paragraph, and the length of each combined paragraph cannot exceed a certain number of characters (e.g., 200 characters). Since paragraphs 1 and 2 contain dialogue and their total length is within the limit, they are combined into paragraph A: ("The weather is so nice today," said Xiaoming, "Let's go to the park." Xiaohong nodded and smiled, saying, "Okay, I'll bring my picnic basket."). Paragraph 3 is descriptive text (i.e., a paragraph without dialogue) and is grouped separately as paragraph B: (They walked together to the park, the sunlight shining on the grass.). Paragraphs 4 and 5 contain dialogue and their total length is within the limit, so they are combined into paragraph C: ("It's so beautiful here," Xiaoming exclaimed, "Let's find a place to sit down." Xiaohong opened her picnic basket and took out sandwiches and juice.). Finally, paragraphs A, B, and C can be combined in sequence to form a text section.

[0122] In other embodiments, the multiple text segments obtained from the above division (see the description of step 101 above) are combined into text chapters according to rules such as length limitation processing, that is, a text chapter includes at least two text segments.

[0123] For example, a length limit rule could be that the total length of each text section does not exceed a certain number of characters (e.g., 150 characters). Suppose the target text is divided into the following text segments (e.g., segments separated by periods): Segment 1: Xiaoming and Xiaohong decided to go to the park. They brought a picnic basket filled with sandwiches and juice. Segment 2: The park was sunny, and the grass was full of wildflowers. They found a quiet spot to sit down and began to enjoy their picnic. Segment 3: Xiaoming said, "This is so comfortable, let's come again next time." Xiaohong smiled and nodded, "Okay, next time we'll bring even more delicious food." Checking the length of each segment to ensure it meets the length limit, finally, segments 1, 2, and 3 are combined into text section A.

[0124] In step 1023, based on the first prompt word and multiple text chapters, a pre-trained first language model is invoked to extract at least one character information from each text chapter.

[0125] For example, after receiving the first prompt word, the first language model performs task understanding, clarifying that the task is to extract role information from the target text. For example, it parses the first prompt word to clarify the task type and target fields. For example, if the task type is information extraction, the task is decomposed into identifying role names, features, and introductions. The implementation method is to use techniques such as NER to extract role information and generate a structured list of role information.

[0126] In some embodiments, each text section includes at least one text fragment, see [link to relevant documentation]. Figure 4C , Figure 4B Step 1023 shown can be implemented by performing steps 10231 to 10237 for each text chapter, as explained in detail below.

[0127] In step 10231, entity recognition processing is performed on the text chapters to obtain at least one character name.

[0128] In some embodiments, entity recognition (i.e., named entity recognition) for each text section can be achieved as follows: First, the text section is cleaned and preprocessed, such as word segmentation, part-of-speech tagging, and stop word removal. Next, an entity recognition algorithm is selected. The entity recognition algorithm can be a rule-based method, such as a regular expression-based method, or a deep learning-based method, such as a Transformer-based entity recognition model. The selected entity recognition algorithm is used to perform entity recognition processing on the preprocessed text to identify the entities. The entity recognition result can include whether each word in the text section is an entity and the type of the entity. The identified entities are converted into entity-dimensional text features. For example, binary representation can be used to indicate whether each word is an entity, or word vectors of entity words and entity type information can be used to construct a multi-dimensional feature vector. The identified entity feature vectors are concatenated to form an entity-dimensional feature vector, which will serve as the representation of the text section in the entity dimension.

[0129] Taking the training of a deep learning-based entity recognition model as an example, text chapter samples containing entities can be collected, and each entity can be labeled to form a training dataset. The label can include the entity's category (such as person's name, place, organization, etc.) and its position in the text chapter sample. Next, the entity recognition model based on the Transformer architecture is trained using the training dataset with entity labels. During the training process, the entity recognition model will learn how to predict the entity's category and position based on the input text sequence. After training, the entity recognition model can be used to perform entity recognition on new text chapters.

[0130] In step 10232, the text chapters are subjected to referential resolution processing based on at least one character name to obtain referential resolution chapters.

[0131] In some embodiments, a referential resolution model (such as Stanford CoreNLP, Hugging FaceTransformers, etc.) is used to resolve the referential meaning of text chapters. This involves analyzing pronouns (such as "he", "she", "they") or noun phrases in the text chapters to determine their referents (such as the character name "Li Ming"). Then, the pronouns are replaced with the corresponding character names to ensure semantic clarity, resulting in a referential resolution chapter.

[0132] For example, the input text section is: Li Ming is a 30-year-old doctor; he is intelligent and brave. Through entity recognition processing in step 10231, the character name is identified as: Li Ming (i.e., the person entity). The referential resolution model identifies "he" in the text section as referring to "Li Ming." "He" is replaced with "Li Ming." The resulting text section with resolved referential resolution is: Li Ming is a 30-year-old doctor; Li Ming is intelligent and brave.

[0133] Here, coreference resolution aims to identify the specific entities (such as personal names, place names, organization names, etc.) referred to by pronouns (such as "he", "she", "it", "they") or noun phrases in the text, and establish the association between these referential components and their corresponding entities, thereby eliminating ambiguity in the text, clarifying the directionality of language expression, and making the semantics of the text clearer and more coherent.

[0134] In step 10233, at least one text fragment corresponding to each character name in the reference resolution chapter is obtained as the character description text corresponding to each character name.

[0135] In some embodiments, the text fragment containing each character name in the resolution chapter is used as the character description text for that character name.

[0136] For example, suppose the text section is "Zhang Lei and Chen Lin are colleagues. Zhang Lei is responsible for project management and he is experienced. Chen Lin is the technical supervisor and she is good at solving complex problems. Today they completed the task together." After the substitution resolution process, the resulting section is "Zhang Lei and Chen Lin are colleagues. Zhang Lei is responsible for project management and Zhang Lei is experienced. Chen Lin is the technical supervisor and Chen Lin is good at solving complex problems. Today Zhang Lei and Chen Lin completed the task together." Taking the text fragment containing each role name in the substitution resolution section as the role description text for that role name, then the role description text for the role name "Zhang Lei" could be "Zhang Lei and Chen Lin are colleagues. Zhang Lei is responsible for project management and Zhang Lei is experienced. Today Zhang Lei and Chen Lin completed the task together.", and the role description text for the role name "Chen Lin" could be "Zhang Lei and Chen Lin are colleagues. Chen Lin is the technical supervisor and Chen Lin is good at solving complex problems. Today Zhang Lei and Chen Lin completed the task together."

[0137] In step 10234, part-of-speech tagging is performed on each word in the description text of each character to obtain the part-of-speech tagging results for each character description text.

[0138] In some embodiments, each word in each character description text is tagged with part-of-speech tags to determine its grammatical role (such as noun, verb, adjective, etc.) as the part-of-speech tagging result for each character description text.

[0139] For example, each word can be tagged with part-of-speech tags using a parser. A parser is a key component in natural language processing, used to analyze and understand the grammatical structure of text data. Parsers can be rule-based parsers, statistical parsers, deep learning parsers, etc. This application does not limit the type of parser.

[0140] For example, training a deep learning-based parser can be achieved as follows: Data collection: First, a large amount of text data needs to be collected for training the parser. Then, the collected text data is manually annotated, assigning part-of-speech tags and dependency relation tags to each word. Choosing a suitable deep learning architecture, such as the Transformer, as the core of the parser allows for pre-training of the model using a large corpus to learn general representations of the language, and fine-tuning the model for specific parsing tasks to adapt to specific annotation rules and language characteristics.

[0141] Continuing with the above example, the词性标注 result of the role description text corresponding to the role name "Zhang Lei": "Zhang Lei and Chen Lin are colleagues. Zhang Lei is responsible for project management and Zhang Lei has rich experience. Today, Zhang Lei and Chen Lin completed a task together." can be expressed as "[“Zhang Lei” / n, “and” / c, “Chen Lin” / n, “are” / v, “colleagues” / n, “.” / w, “Zhang Lei” / n, “is responsible for” / v, “project” / n, “management” / v, “,” / w, “Zhang Lei” / n, “experience” / n, “rich” / a, “.” / w, “Today” / t, “Zhang Lei” / n, “and” / c, “Chen Lin” / n, “together” / d, “completed” / v, “a” / u, “task” / n, “.” / w]". Here, the annotation rules are as follows: Nouns, such as "Zhang Lei" and "Chen Lin", are annotated as n; Verbs, such as "completed", are annotated as v; Time words, such as "Today", are annotated as t; Adjectives, such as "rich", are annotated as a; Conjunctions, such as "and", are annotated as c; Adverbs, such as "together", are annotated as d; Punctuation marks, such as ".", are annotated as "w"; Auxiliary words, such as "a", are annotated as u, where " / n" represents a line break character.

[0142] In step 10235, based on the词性标注 results of each role description text, obtain the role characteristics corresponding to each role name.

[0143] In some embodiments, the role characteristics include at least one of the following: role gender, role age range, role personality, and role identity. Screen the candidate words related to the role characteristics according to the词性标注 results of each role text, and generate the role characteristics based on the candidate words.

[0144] It should be noted that there is an unclear expression "词性标注" in the original text which should be replaced with the accurate English description according to the actual situation. Here it is temporarily retained as it is for translation.Continuing from the previous example, the part-of-speech tagging result for the role description text "Zhang Lei and Chen Lin are colleagues. Zhang Lei is responsible for project management, and Zhang Lei is experienced. Today, Zhang Lei and Chen Lin completed the task together." corresponding to the role name "Zhang Lei" can be represented as "["Zhang Lei" / n, "and" / c, "Chen Lin" / n, "is" / v, "colleague" / n, "." / w, "Zhang Lei" / n, "responsible for" / v, "project" / n, "management" / v, "," / w, "Zhang Lei" / n, "experience" / n, "rich" / a, "." / w, "today" / t, "Zhang Lei" / n, "and" / c, "Chen Lin" / n, "together" / d, "completed" / v, "finished" / u, "task" / n, "." First, extract vocabulary, such as nouns (n): Zhang Lei, colleague, project, management, experience; adjectives (a): rich; verbs (v): responsible, manage, complete. Next, filter vocabulary related to role characteristics, such as: words related to identity / occupation: "colleague," "project," "management"; words related to personality / ability: "rich"; words related to duties / behaviors: "responsible," "manage," "complete". Finally, generate role characteristics: Role Name: Zhang Lei. Identity: Project Management Leader, Chen Lin's colleague. Personality / Ability: Rich in experience.

[0145] It should be noted that, in the embodiments of this application, the character characteristics of each character can be continuously integrated and updated during the character information extraction process, thereby extracting a complete list of character information.

[0146] In step 10236, the character description text corresponding to each character name is processed to generate a summary text, resulting in a character introduction for each character name.

[0147] In some embodiments, important sentences from the character description text are selected and combined to form the character introduction corresponding to the character name. For example, firstly, sentence importance is scored, such as based on position (e.g., first paragraph / last paragraph), keyword coverage, sentence length, etc. Next, redundancy filtering is performed, such as removing duplicate or similar sentences. Finally, the sentences are sorted, for example, arranged according to the original text order or logical relationship, to obtain the character introduction.

[0148] For example, the input might be a character description text containing 5 paragraphs. The output would extract the character synopsis from paragraph 1 (character background, such as name, occupation, gender, etc.), paragraph 3 (character personality traits), and paragraph 5 (age range).

[0149] In other embodiments, after understanding the original text (i.e., the character description text) through a language model, the core content is restated (i.e., new sentences are generated) to serve as the character introduction. For example, the character description text is encoded into vectors through a deep learning model (such as Transformer), and new sentences are generated based on the encoded vectors to cover the key information of the original text. Finally, constraints (such as length limits and keyword coverage) are used to ensure the relevance of the summary, resulting in the final character introduction.

[0150] For example, the input character description text could be: "Li Ming is a 30-year-old doctor. Li Ming is intelligent and brave. Today, Li Ming saw a patient." The output character description could be: "Li Ming is a promising young doctor who saw a patient today."

[0151] In step 10237, each character name, the corresponding character characteristics, and the character introduction are combined to form the character information for each character.

[0152] In some embodiments, the character information for each character is obtained by combining the character name and the correspondence between the character name and the character characteristics and character description.

[0153] Through steps 10231 to 10237, the system uses text chapters as input units and processes such as entity recognition, referential resolution, and part-of-speech tagging to structurally store the role information scattered in the target text. This avoids the inefficiency of manual sentence-by-sentence tagging, covers the multi-dimensional attributes of the roles, and thus provides accurate basic data of the roles for subsequent steps.

[0154] See also Figure 4B In step 1024, at least one character information from each text chapter is combined into a character information list according to the order of the multiple text chapters in the target text.

[0155] In some embodiments, a sub-role information list is formed based on the order in which the characters appear in the text chapters, and then the sub-role information list is combined into a character information list based on the order in which the text chapters appear in the target text.

[0156] For example, see Figure 7 In response to a trigger operation on the character information extraction control 002, a list of character information is obtained. In response to a trigger operation on the character information editing control 003, the character information can be edited. Based on the edited character information, the process proceeds to identify the character in each text fragment based on the character information list.

[0157] In other embodiments, extracting the character information list from the target text can also be achieved by: dividing the target text into multiple text chapters, wherein each text chapter includes at least one text segment; performing the following processing on each text chapter: extracting character information from the text chapter according to a pre-set character matching template to obtain at least one character information in the text chapter; combining the at least one character information in each text chapter into a character information list according to the order of the multiple text chapters in the target text, wherein the text matching template includes: a character name keyword template, a title pattern template, a gender keyword template, an age keyword template, a personality keyword template, an identity keyword template, and a character introduction template.

[0158] For example, first, segmentation rules are defined to identify chapter boundaries in the target text, such as by recognizing specific delimiters (e.g., blank lines, specific punctuation marks, or keywords). Next, the target text is segmented into multiple text chapters according to the segmentation rules. One or more character matching templates are pre-defined to identify character names, characteristics, and descriptions in the text. Templates can be a series of rules or patterns, such as regular expressions, used to match sentences or phrases containing character information. Further, each segmented text chapter is processed, using the pre-defined character matching templates to identify character information within the chapter. Within each text chapter, the patterns defined by the matching templates are searched to extract character names, characteristics, and descriptions. For each text chapter, at least one character information is extracted based on the matched template. The extracted character information from each text chapter is combined according to the order of the text chapters in the target text. This combination can form a new list or data structure containing all extracted character information. Finally, the combined character information list is output or saved for subsequent querying, analysis, and use.

[0159] For example, pre-set role matching templates may include: "Role Name Keyword Template: [Li Ming, Wang Wu, Xiao Hong]; Title Pattern Template: [".*Doctor", ".*Manager"]; Gender Keyword Template: {Male: [He, Mr.], Female: [She, Ms.]}; Personality Keyword Template: [Smart, Brave, Gentle]; Identity Keyword Template: [Doctor, Teacher, Police Officer]; Age Keyword Template: [Teen, Middle-aged, ** years old, ** to ** years old, ** years old - ** years old]; Role Introduction Template: [(Enter the role name here) is a (Enter the role personality here) of (Enter the role identity here)]".

[0160] See also Figure 4A In step 103, the characters in each text segment are identified based on the character information list.

[0161] In some embodiments, see Figure 4D , Figure 4A Step 103 shown can be achieved by performing steps 1031 to 1032 for each text segment, as explained below.

[0162] In step 1031, a second prompt word is obtained, wherein the second prompt word indicates the role corresponding to the identified text fragment.

[0163] In some embodiments, the second prompt is pre-set and used to prompt the language model (i.e., the second language model below) to identify the role corresponding to each text fragment based on the role information list obtained above. For example, the second prompt may be expressed as "Please select the role corresponding to the current text fragment from the currently input role information list" or "Please output the matching score between each role in the role information list and the current text fragment. The matching score ranges from 0 to 1, where 0 represents no match and 1 represents a complete match", so that the role with the highest matching score is taken as the role corresponding to the text fragment.

[0164] Here, the second cue word is used to guide the language model to associate the text fragment with the characters in the character information list. Specifically, it can be direct matching or rating matching. Its core function is to determine the character described by the text fragment.

[0165] For example, the input to a language model can include a text fragment, a list of character information, and a second prompt word. For instance, if the input is the second prompt word: "Please output the matching score (0-1) between each character in the character information list and the current text fragment," the text fragment is: "He treated a patient.", and the character information list is: [Li Ming, Zhang Lei, Chen Lin] (only the character names in the character information list are shown here for ease of description; the input character information list also includes character features and character descriptions), the output would be: Li Ming: 0.9, Zhang Lei: 0.1, Chen Lin: 0.0, then the character corresponding to the text fragment is "Li Ming".

[0166] In step 1032, a pre-trained second language model is invoked based on the second prompt word, the list of character information, the text fragment, and the context of the text fragment to predict the character corresponding to the text fragment.

[0167] In some embodiments, see Figure 4E , Figure 4D Step 1032 shown can be implemented through steps 10321 to 10324, which will be explained in detail below.

[0168] In step 10321, the list of character information, the text fragment, and the context of the text fragment are concatenated to obtain the character prediction text.

[0169] Here, the role information list is used to list the roles and their descriptions in a structured form, the context of the text fragment is used to provide the background of the text fragment to help the language model (i.e., the second language model) understand the semantics, and the text fragment is used to clarify the content to be analyzed.

[0170] For example, the list of character information can be represented as: [Zhang San; Characteristics: Age 30-35, male, company manager, decisive personality; Introduction: Key figure in the company, strong decision-making ability. Li Ming; Characteristics: 30 years old, doctor, intelligent and brave; Introduction: Years of medical experience, skilled in solving difficult and complicated cases. Chen Lin; Characteristics: 25 years old, technical supervisor; Introduction: Key figure in the company, skilled in problem-solving]. The context of the text fragment: Li Ming is a 30-year-old doctor. He is intelligent and brave. Text fragment: He treated a patient. Therefore, the concatenated character prediction text can be represented as: [Zhang San; Characteristics: Age 30-35, male, company manager, decisive personality; Introduction: Key figure in the company, strong decision-making ability. Li Ming; Characteristics: 30 years old, doctor, intelligent and brave; Introduction: Years of medical experience, skilled in solving difficult and complicated cases. Chen Lin; Characteristics: 25 years old, technical supervisor; Introduction: Key figure in the company, skilled in problem-solving]. Li Ming is a 30-year-old doctor. He is intelligent and brave; he treated a patient.

[0171] By combining a list of character information, context, and text fragments, language models can more accurately predict characters and generate structured output.

[0172] In other embodiments, the second prompt word, the list of character information, the text fragment, and the context of the text fragment are concatenated to obtain the character prediction text. Based on the concatenated character prediction text, feature encoding is performed on the character prediction text to obtain the character prediction features.

[0173] In step 10322, the character prediction text is feature encoded to obtain character prediction features.

[0174] In some embodiments, one-hot encoding can be performed on categorical features (such as occupation, gender, etc. in the role information list) in the role prediction text, and normalization processing can be performed on numerical features (such as age in the role information list, numerical values ​​appearing in the text fragment, etc.) in the role prediction text (or numerical values ​​can be used directly). Word embedding encoding can be performed on other textual features in the role prediction text, and then the results of the above encoding are concatenated according to the order of appearance in the role prediction text to obtain the role prediction features.

[0175] For example, the identities (occupations) in the role information list: doctor, teacher, engineer are encoded as [1, 0, 0], [0, 1, 0], and [0, 0, 1] respectively using one-hot encoding. The genders in the role information list: male and female are encoded as 0 and 1 respectively using one-hot encoding.

[0176] For example, a numerical characteristic such as age, like 30, is standardized to (30 - mean) / standard deviation.

[0177] For example, word embedding encoding of text features to obtain the embedding features of each word in the text features can be achieved as follows: Construct a vocabulary that contains all the words or symbols that appear in the text. For each word or symbol in the vocabulary, convert it into a vector of a fixed size. This process is called word embeddings. The dimension of word embeddings is a hyperparameter that can be chosen with values ​​such as 50, 100, and 300. The higher the embedding dimension, the more information the model can capture, but the higher the computational cost. By training models such as Word2Vec and GloVe, the embedding matrix of words or symbols can be learned, thereby converting each word into the corresponding embedding feature.

[0178] In step 10323, feature mapping is performed on the character prediction features to obtain the predicted probability value of each character in the character information list.

[0179] In some embodiments, feature mapping can be implemented using fully connected layers (FC). The dimension of the fully connected layer is set to the number of characters in the character information list, and the probability distribution of each character is output through the Softmax function, thereby mapping the predicted features of the characters to the predicted probability value of each character in the character information list.

[0180] In step 10324, the character with the highest predicted probability value is selected as the character corresponding to the text segment.

[0181] Steps 10321 to 10324 enable text processing on a per-segment basis. By combining information such as the character list and dialogue context with the second language model, the corresponding character for each text segment is determined, achieving accurate character assignment and improving the targeting and effectiveness of text analysis.

[0182] See also Figure 4A In step 104, the timbre information of each character is determined based on the character information of each character.

[0183] For example, see Figure 7 In response to a trigger operation on the timbre assignment control 004, the process transitions to determining the timbre information of each character based on their role information.

[0184] In some embodiments, see Figure 4F , Figure 4A Step 104 shown can be achieved by performing steps 1041 to 1043 for each role, as explained in detail below.

[0185] In step 1041, according to the preset timbre matching rules and the character's role information, multiple candidate timbre features that match the character's role information are queried from the timbre library.

[0186] In some embodiments, a character's character information includes a character name, character traits, and a character description. Character traits include at least one of the following: character gender, character age range, character personality, and character identity; see also Figure 4G , Figure 4E Step 1041 shown can be implemented through steps 10411 to 10413, which will be explained in detail below.

[0187] In step 10411, multiple timbre feature samples matching the character's age range and gender are queried from the timbre library.

[0188] In some embodiments, multiple timbre feature samples that match the character's age range and gender are queried from the timbre library.

[0189] For example, a timbre library can contain the following metadata fields to support precise queries of timbre feature samples: timbre identifier (i.e., a unique identifier for the timbre); gender, such as male, female, or neutral (optional); age range, such as "child," "youth," "middle-aged," "elderly," or a specific age range (e.g., 25-35 years old); timbre description tags, i.e., keywords used to describe the timbre, such as "clear," "deep," "soft," "hoarse," "magnetic," "lively," etc.; pitch range (Hz), such as fundamental frequency range (e.g., male average 85-180Hz, female average 165-255Hz); emotional adaptation tags, such as the emotions suitable for expression, such as "majestic," "gentle," "cheerful," "sad," etc.; applicable scene tags, such as "animation protagonist," "narrator," "villain," "children's book," etc.

[0190] For example, assuming the character Zhang San is a male aged 30-35, then query the voice feature library for all voice feature samples that are male and aged 30-35.

[0191] In step 10412, the matching weights corresponding to the character personality, character identity, and character introduction are obtained from the timbre matching rules, and the one with the highest matching weight among the character personality, character identity, and character introduction is taken as the information to be matched.

[0192] In some embodiments, the voice matching rules pre-set matching weights for character personality, character identity, and character introduction, respectively. For example, if the matching weights for character personality, character identity, and character introduction are 0.6, 0.3, and 0.1 respectively, then character personality is used as the information to be matched.

[0193] In step 10413, from multiple timbre feature samples, timbre feature samples that match the information to be matched are queried as candidate timbre features.

[0194] Following the example above, when character personality is used as the information to be matched, the system queries multiple timbre feature samples to find timbre feature samples that match the character personality, and uses these as candidate timbre features.

[0195] Through steps 10411 to 10413, multiple timbre feature samples in the timbre library are matched according to the character's age range and gender. Then, candidate timbre features are determined from the multiple timbre feature samples based on the information to be matched. This achieves multi-stage timbre screening, improves the efficiency of timbre screening, and enhances the flexibility and accuracy of timbre screening by setting matching weights.

[0196] See also Figure 4F In step 1042, a third prompt word is obtained, wherein the third prompt word indicates the vocal characteristics corresponding to the character to be filtered out.

[0197] In some embodiments, the third prompt word is pre-set and used to prompt the language model to filter out the timbre features corresponding to the character from the candidate timbre features. For example, the third prompt word can be expressed as "Please filter out the timbre features corresponding to the character from the currently input candidate timbre features".

[0198] Here, the third cue word is used to guide the language model (corresponding to the timbre selection model below) to associate the character with timbre features. Specifically, it can be direct matching or rating matching. Its core function is to determine the character's timbre features.

[0199] In step 1043, based on the third prompt word, the character's role information, and multiple candidate timbre features, the character's timbre features are selected from the multiple candidate timbre features, and the timbre features and the audio that matches the timbre features are used as the character's timbre information.

[0200] In some embodiments, a pre-trained third language model is invoked based on a third cue word, and the character's timbre features are selected from multiple candidate timbre features according to the character information.

[0201] For example, a pre-trained timbre selection model (i.e., a third-language model) can be used to select the timbre features of a character from multiple candidate timbre features based on character information. The timbre selection model can be trained as follows: First, data collection is performed, such as collecting a large number of timbre audio samples. Each sample has a corresponding text description indicating its personality traits, such as extroverted, introverted, friendly, serious, etc., as well as timbre features, such as "clear," "deep," "soft," "hoarse," "magnetic," "lively," etc., and emotional fit, such as the emotions the sample is suitable to express, such as "dignified," "gentle," "cheerful," "sad," etc. Next, the text descriptions are preprocessed, such as word segmentation, stop word removal, and part-of-speech tagging, to improve the efficiency of text analysis. Next, feature extraction is performed on the timbre audio samples, such as Mel Frequency Cepstral Coefficients (MFCC), spectral features, and fundamental frequency. Word embedding encoding is then performed on the text description to obtain text features. Using both audio and text features as input data, an architecture based on a Variational Autoencoder (VAE) or Transformer model is selected as the timbre selection model to be trained. Next, model training is performed, feeding the input data into the model and designing a loss function. The matching degree between audio and text features is used as the optimization objective. Gradient descent or other optimization algorithms are used to train the model, and the model parameters are adjusted through multiple iterations to reduce the value of the loss function. The model's performance is evaluated on independent validation and test sets to ensure that the model can match the timbre audio samples with the text descriptions. Based on the evaluation results, model parameters or training strategies are adjusted to improve the model's accuracy and generalization ability.

[0202] For example, see Figure 7 In response to the trigger operation of the timbre information editing control 005, the timbre characteristics corresponding to the character can be modified. Based on the modified timbre information, the processing is switched to speech synthesis based on each text segment and the timbre information of the character in each text segment, so as to obtain the audio segment corresponding to each text segment.

[0203] See also Figure 4A In step 105, speech synthesis is performed based on each text segment and the timbre information of the characters in each text segment to obtain the audio segment corresponding to each text segment.

[0204] In some embodiments, a pre-trained speech synthesis model is used to synthesize speech based on each text segment and the timbre information of the characters in each text segment, to obtain an audio segment corresponding to each text segment.

[0205] For example, see Figure 5 , Figure 5 This is a schematic diagram of an optional network structure for the speech synthesis model provided in this application embodiment. The speech synthesis model includes a phoneme embedding model, a speech discretization model, a language model, and a vocoder. The text (which may be the narration text or dialogue text below) is input into the phoneme embedding model to obtain the phoneme representation of the text. The phoneme representation of the text is used as the text feature. The reference audio (which may be the narration timbre audio below or audio that matches the timbre feature) is input into the speech discretization model to obtain the timbre feature of the reference audio. The text feature and timbre feature are input into the language model (e.g., based on the Transformer architecture). The language model performs audio feature prediction to obtain the audio feature. The audio feature is converted into synthesized audio by a pre-trained vocoder.

[0206] For example, see Figure 7 In response to a trigger operation on the speech synthesis control 006, the process switches to speech synthesis based on each text segment and the timbre information of the character in each text segment, and obtains the audio segment corresponding to each text segment.

[0207] In some embodiments, timbre information includes audio that matches timbre characteristics, and the audio segment includes at least one of narration audio segments and dialogue audio segments, see [link to documentation]. Figure 4H , Figure 4A Step 105 shown can be achieved by performing steps 1051 to 1053 for each text segment, as explained below.

[0208] In step 1051, the text fragment is subjected to narration text detection processing to obtain narration detection results, and the text fragment is subjected to dialogue text detection processing to obtain dialogue detection results.

[0209] In some embodiments, narration text is descriptive content directly spoken by the narrator or a non-character, typically used to explain background, time, scene, or character's psychological activities, and does not contain direct dialogue between characters. The detection rules for narration text may include: no speaker identifier, such as not containing dialogue verbs such as "speak," "ask," or "shout"; using a third-person perspective, such as using "he," "she," "they," or the character's full name; being descriptive content, such as containing time, place, environment, and psychological descriptions; and not containing dialogue identifiers such as quotation marks ("") or dashes (—).

[0210] For example, the input text fragment is: Night falls, and the city is shrouded in a light rain. Li Ming stands by the window, gazing at the flashing neon lights in the distance, his heart filled with confusion. The narration detection result can be represented as: Narration text content: Night falls, and the city is shrouded in a light rain. Li Ming stands by the window, gazing at the flashing neon lights in the distance, his heart filled with confusion. Features: No dialogue verbs, third-person perspective, environmental and psychological descriptions.

[0211] In some embodiments, dialogue text is direct conversation between characters, including speaker identifiers and dialogue content. Dialogue text detection rules may include: including speaker identifiers, such as verbs like “speak,” “ask,” and “answer”; using first / second person pronouns, such as pronouns like “I,” “you,” and “we”; including dialogue content, such as direct or indirect quotations; and including quotation marks (“”), dashes (—), or colons (:).

[0212] For example, given the input text fragment: Li Ming suddenly turned around and shouted to Chen Lin, "We must leave here immediately!" Chen Lin hesitated for a moment and replied in a low voice, "Give me one more minute.", the dialogue detection results can be represented as: Dialogue 1: Speaker: Li Ming, Content: "We must leave here immediately!", Identifier: "shouted"; Dialogue 2: Speaker: Chen Lin, Content: "Give me one more minute.", Identifier: "replied".

[0213] For example, when a text fragment contains both narration and dialogue, the input text fragment is: The rain is getting heavier and the thunder is booming (narration text 1). Li Ming grabbed his coat and said anxiously to Chen Lin (narration text 2): "Hurry up, it's not safe here!" (dialogue text 1). Chen Lin looked out the window, the rain blurring the glass (narration text 3). Detection results: Narration text 1, content: "The rain is getting heavier and the thunder is booming."; Narration text 2, content: "Li Ming grabbed his coat and said anxiously to Chen Lin"; Dialogue text 1, speaker: Li Ming, content: "Hurry up, it's not safe here!", identifier: "said"; Narration text 3, content: "Chen Lin looked out the window, the rain blurring the glass."

[0214] By performing narration and dialogue detection processing on the text fragments, the distinction between narration and dialogue text is achieved. This is because narration text is usually narrative, such as describing scenes or characters' psychology, while dialogue text is the conversation between characters. Therefore, in subsequent speech synthesis, speech synthesis is performed on narration text and dialogue text separately, which helps to enhance the expressiveness and realism of subsequent speech synthesis.

[0215] In step 1052, in response to the narration text result indicating the presence of narration text in the text segment, narration speech generation processing is performed based on the narration text and the pre-set narration audio timbre to obtain a narration audio segment.

[0216] In some embodiments, see Figure 4I , Figure 4H Step 1052 shown can be implemented through steps 10521 to 10525, which will be explained in detail below.

[0217] In step 10521, the second text features of the narration text are extracted to obtain the second text features.

[0218] In some embodiments, the text features of the narration text can be obtained by: performing word segmentation on the narration text to obtain input units; performing word embedding encoding on the input units to obtain word embedding features; and performing self-attention encoding on the word embedding features to obtain the text features of the narration text.

[0219] For example, punctuation marks such as spaces, periods, and commas can be used as segmentation markers to segment the narration text, that is, to divide the narration text into input units. Third-party segmentation tools can also be used to segment the narration text, such as Jieba, NLTK, and SpaCy. These tools can achieve accurate Chinese and English word segmentation through algorithms and language models.

[0220] For example, word embedding models (such as Word2Vec, GloVe, etc.) can be used to encode the input units, representing each input unit as a fixed-length vector, and then forming word embedding features (sequences) in order.

[0221] For example, word embedding features can be used as query vectors (Q), key vectors (K), and value vectors (V). Attention scores are obtained by calculating the dot product of Q and K. Next, the attention scores are normalized, for example by applying a normalization function (such as the softmax function), to transform the attention scores into a probability distribution, which is then used as attention weights. The attention weights are used to weight the word embedding features (corresponding to V above) to generate new feature representations. These new feature representations can be linearly transformed through a linear layer to obtain a rich representation that includes the correlation between different positions in the word embedding features, i.e., the text features of the narration text.

[0222] In other embodiments, the narration text is phoneme-encoded to obtain a phoneme sequence; the phoneme sequence is then processed by phoneme embedding to obtain text features.

[0223] For example, phoneme-based encoding of narration text can be achieved as follows: The narration text is segmented into words to obtain a word sequence; each word in the word sequence is converted into its corresponding phoneme. Here, the phoneme corresponding to each word (data unit) is obtained by converting each word using a pre-set grapheme-to-phoneme (G2P) model; the phonemes corresponding to each word are combined into a phoneme sequence; each phoneme in the phoneme sequence is processed by phoneme vector mapping to obtain a phoneme vector; each phoneme in the phoneme sequence is processed by position encoding to obtain a position vector; the phoneme vector and position vector of each phoneme are fused (e.g., added) to obtain the phoneme feature of each phoneme; the phoneme features of each phoneme are combined into text features (i.e., the phoneme features of each phoneme are combined according to the order of the phonemes in the phoneme sequence).

[0224] For example, the narration text can be segmented using a tokenizer based on Byte Pair Encoding (BPE) or a tokenizer based on Byte-Left Byte Pair Encoding (BLBPE). This application does not limit the specific tokenizer or segmentation method used to segment the narration text.

[0225] For example, the pre-set character-to-phoneme model is trained as follows: multiple word samples and phoneme annotations for each word sample are obtained; for each word sample, the following processing is performed: word embedding is performed on the word sample using the initialized character-to-phoneme model to obtain a word sample vector; the word sample vector is cyclically encoded to obtain a cyclic encoded vector; the cyclic encoded vector is non-linearly mapped to obtain a predicted phoneme; a first loss value is determined based on the phoneme annotations and predicted phonemes for each word sample; the parameters of the initialized character-to-phoneme model are updated based on the first loss value to obtain the trained character-to-phoneme model.

[0226] For example, the phonemes corresponding to each word or each character can be combined into a phoneme sequence. A separator (such as a space) can be added between the phoneme representations of words or characters to distinguish the phonemes of different words, thus combining the phonemes corresponding to each word into a phoneme sequence.

[0227] For example, through a pre-trained phoneme embedding model (corresponding to Figure 5 The phoneme embedding models shown in the figure, such as the transformer-based bidirectional encoder (Phone me-Level BERT, PL-BERT), the phoneme steering model (Phoneme2Vec), etc., perform phoneme vector mapping processing on each phoneme in the phoneme sequence to obtain the phoneme vector of each phoneme.

[0228] For example, taking PL-BERT as the phoneme embedding model, the training of the phoneme embedding model can be achieved as follows: obtain multiple phoneme sequence samples, each phoneme sequence sample including phoneme samples; perform the following processing on each phoneme sequence sample: mask the phoneme sequence sample to obtain masked phoneme sequence samples; predict the masked phoneme sequence samples using the initialized phoneme embedding model to obtain predicted phoneme sequence samples; determine the prediction loss value based on the predicted phoneme sequence samples and the phoneme sequence samples; update the parameters of the initialized phoneme embedding model according to the prediction loss value to obtain the trained phoneme embedding model.

[0229] For example, position embeddings can be generated using a fixed algorithm, such as a combination of sine and cosine functions. For the dimension index i of the position embedding feature of each position (the position of each phoneme in the phoneme sequence) (for example, if the number of dimensions of the position vector of each phoneme is d, then the range of dimension index i is from 0 to (d-1)), the position embedding for even-numbered dimensions uses the sine function, while the position embedding for odd-numbered dimensions uses the cosine function. The embodiments of this application do not limit the specific implementation of position embeddings.

[0230] In other embodiments, after obtaining the second text features, the following processing may be performed: extracting emotion parameters from the narration text to obtain emotion parameters of the narration text; concatenating the emotion parameters with the second text features to obtain new second text features; and performing feature fusion based on the new second text features to obtain second fused features.

[0231] For example, the narration text can be processed by extracting emotion parameters through a pre-configured emotion recognition model to obtain the emotion parameters of the narration text. The emotion recognition results are used to characterize the emotion type and intensity in the narration text.

[0232] For example, emotion types include positive, negative, and neutral, and emotion intensity includes strong and moderate. The emotion recognition results can be quantified to obtain the emotion parameters of the narration text. For example, positive emotion, negative emotion, and neutral emotion can be represented by "+", "-", and a space, respectively. Strong emotion is represented by 1, and moderate emotion is represented by 0. For example, "+1" represents strong positive emotion, and "0" represents moderate neutral emotion.

[0233] It's important to note that emotion type refers to the tendency of an emotion, such as positive, negative, and neutral emotion types. For example, positive emotion types can include happiness, excitement, and exhilaration, while negative emotion types can include sadness, shame, and anger. Emotion intensity is an indicator used to quantify the strength of an emotion; generally, the intensity of a normal emotion is lower than that of a strong emotion.

[0234] Taking the training of a deep learning-based emotion recognition model as an example, texts with strong positive emotions can be manually labeled as positive emotion samples, texts with strong negative emotions as negative samples, and other text samples as neutral emotion samples. The labeled and preprocessed text data can then be used for model training.

[0235] For example, based on the Transformer model framework, the model can learn to identify the emotion category (e.g., positive, negative, and neutral) and emotion intensity (e.g., strong, moderate) of the narration text. The performance of the trained model can be evaluated using evaluation set data, such as by calculating metrics like accuracy and precision. During the inference phase, given new narration text, the emotion recognition model will predict the emotion category and emotion intensity of the narration text.

[0236] In step 10522, the second timbre feature is extracted from the narration audio to obtain the second timbre feature.

[0237] In some embodiments, the narration audio is discretized to obtain multiple discrete symbols; the discrete symbol sequence composed of multiple discrete symbols is used as the second timbre feature.

[0238] For example, the narration audio is first preprocessed, such as removing background noise from the speech signal and identifying and removing silences in the speech. Next, the continuous speech signal is segmented into short audio frames (e.g., 20-30 milliseconds), and acoustic features (e.g., Mel-frequency cepstral coefficients, spectral features, etc.) are extracted from each audio frame. These features can capture the characteristics of the speech signal. Feature mapping is performed on the acoustic features to obtain multiple discrete symbols. The discrete symbol sequence composed of multiple discrete symbols is then used as the second timbre feature. For example, assuming the Mel-frequency cepstral coefficient feature (i.e., acoustic feature) extracted from an audio frame is represented as: ([[0.1, 0.2, 0.3], [0.4, 0.5, 0.6], [0.7, 0.8, 0.9]]), the corresponding discrete symbol sequence composed of multiple discrete symbols can be represented as: ['b', 'a', 'k', 'e', ​​'r'] (assuming these discrete symbols correspond to phonemes).

[0239] For example, this can be achieved using a pre-trained speech discretization model (corresponding to...). Figure 5 The speech discretization model shown in the figure represents the audio of the narration timbre (corresponding to...). Figure 5 The reference audio shown in the figure is discretized and encoded. The speech discretization model can be trained in the following way: acquire a large amount of unlabeled speech data (without text alignment), encode the speech waveform of each speech data into a continuous vector (e.g., through a convolutional neural network), map the continuous vector into discrete symbols through codebook mapping, decode and reconstruct the speech waveform based on the discrete symbols, and minimize the difference between the original speech and the reconstructed speech (e.g., Mel spectrum error) by setting a loss function.

[0240] In step 10523, feature fusion is performed based on the second text feature and the second timbre feature to obtain the second fused feature.

[0241] In some embodiments, the second text feature and the second timbre feature are concatenated to obtain the second fused feature.

[0242] For example, assuming the second text feature is represented as [0.2, 0.5, ..., 0.3] and the second timbre feature is represented as [1, 0, 0, 0, 0], then the concatenated second fusion feature can be represented as [0.2, 0.5, ..., 0.3, 1, 0, 0, 0, 0].

[0243] In step 10524, the second audio feature is predicted based on the second fusion feature to obtain the second audio feature.

[0244] In some embodiments, a second audio feature prediction (autoregressive prediction) is performed based on the second fusion feature to obtain the second audio feature.

[0245] For example, autoregressive predictions can be made using a pre-trained autoregressive model (corresponding to...). Figure 5 The language model shown in the diagram is implemented as follows: The training of the autoregressive model can include: inputting a text-speech pair (second text feature + corresponding second timbre feature (i.e., discrete symbol sequence)); during training, inputting the second text feature; predicting the discrete symbol sequence of speech; and optimizing the model using autoregressive loss (such as cross-entropy). In the inference phase, the narration text is converted into second text features (e.g., extracted via BERT), and discrete symbols from the narration timbre audio are extracted as context (i.e., the second text feature and the second timbre feature are concatenated to obtain the second fused feature), thereby guiding the generation of similar timbres. Autoregression predicts the discrete symbol at each position, thus obtaining the second audio feature (corresponding to...). Figure 5 (The audio features shown in the image).

[0246] In step 10525, the second audio feature is decoded to obtain the narration audio segment.

[0247] In some embodiments, the second audio features can be decoded using a pre-trained vocoder to obtain the narration audio segment (corresponding to...). Figure 5 (Synthesized audio shown in the image).

[0248] Example, vocoder (corresponding) Figure 5 The vocoder shown can be trained as follows: obtain audio samples and corresponding audio feature samples; perform convolution processing on the audio feature samples using the initialized vocoder to obtain convolutional feature samples; perform speech signal mapping processing on the convolutional feature samples to obtain predicted audio; determine the loss value based on the audio samples and predicted audio; update the parameters of the initialized vocoder based on the loss value to obtain the trained vocoder.

[0249] For example, obtain audio samples and corresponding labeled speech text, and perform steps 10521 to 10523 above based on the labeled speech text and audio samples to obtain audio feature samples.

[0250] For example, an initialized vocoder (such as WaveNet) is used to convolve audio feature samples to obtain convolutional feature samples. Taking WaveNet as an example, convolutional processing of audio feature samples to obtain convolutional feature samples can be achieved by sequentially performing convolution processing on the audio feature samples using causal convolutional layers and dilated causal convolutional networks.

[0251] For example, speech signal mapping can be performed on convolutional feature samples using activation functions (such as Sofatmax, ReLU, etc.) to obtain predicted audio. For instance, after performing non-linear mapping on convolutional feature samples using the ReLU activation function, convolution is performed through a convolutional layer (Conv1d) with a 1×1 kernel. The resulting convolution is then mapped using the Sofatmax function to obtain the predicted audio.

[0252] For example, the loss value can be determined based on the audio samples and the predicted audio using a pre-set loss function (such as cross-entropy loss).

[0253] For example, the gradient information of the loss value with respect to each parameter of the vocoder can be obtained through the backpropagation algorithm. Then, the parameters of the initialized vocoder can be updated using the obtained gradient information according to the gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). The above process is repeated until a certain number of iterations are reached or the vocoder converges, thereby obtaining the trained vocoder.

[0254] It should be noted that the embodiments of this application do not limit the specific vocoder used, and other vocoders such as WaveGlow can also be used.

[0255] Through steps 10521 to 10525, a narration audio segment is obtained based on the narration text in the text segment, the emotional parameters of the text segment (such as the category of emotion, such as anger, joy, etc., and the intensity value of the emotion), and the pre-set narration timbre, thus achieving a highly infectious dubbing effect.

[0256] See also Figure 4H In step 1053, in response to the dialogue detection result indicating that there is dialogue text in the text segment, dialogue speech generation processing is performed based on the dialogue text and timbre information to obtain the dialogue audio segment.

[0257] In some embodiments, timbre information includes timbre features and audio matching the timbre features, see [link to documentation]. Figure 4J , Figure 4H Step 1053 shown can be implemented through steps 10531 to 10535, which will be explained in detail below.

[0258] In step 10531, the first text features of the dialogue text are extracted to obtain the first text features.

[0259] For details, please refer to the explanation of step 10521 above; it will not be repeated here.

[0260] In step 10532, the first timbre feature is extracted from the audio that matches the timbre feature to obtain the first timbre feature.

[0261] For details, please refer to the explanation of step 10522 above; it will not be repeated here.

[0262] In step 10533, feature fusion is performed based on the first text feature and the first timbre feature to obtain the first fused feature.

[0263] For details, please refer to the explanation of step 10523 above; it will not be repeated here.

[0264] In step 10534, the first audio feature is predicted based on the first fusion feature to obtain the first audio feature.

[0265] For details, please refer to the explanation of step 10524 above; it will not be repeated here.

[0266] In step 10535, the first audio feature is decoded to obtain the dialogue audio segment.

[0267] For details, please refer to the explanation of step 10525 above; it will not be repeated here.

[0268] Through steps 10531 to 10535, dialogue audio segments are obtained based on the dialogue text in the text segment, the emotional parameters of the text segment (such as the category of emotion, such as anger, joy, etc., and the intensity value of the emotion), and the character's timbre corresponding to the dialogue text, thus achieving a highly infectious dubbing effect.

[0269] For example, see Figure 7 In response to a trigger operation on the audio clip editing control 007, background music can be added to each audio clip. Based on the edited audio clips, the audio clips corresponding to each text clip are combined to obtain the audio corresponding to the target text.

[0270] See also Figure 4A In step 106, the audio segments corresponding to each text segment are combined to obtain the audio corresponding to the target text.

[0271] In some embodiments, the narration audio segment and dialogue audio segment corresponding to each text segment are combined according to the order of their appearance in the text segment to obtain the audio segment corresponding to each text segment. The audio segments corresponding to each text segment are then combined to obtain the audio corresponding to the target text.

[0272] In other embodiments, see Figure 4K After obtaining the audio segment corresponding to each text segment, the following steps 201 to 204 can be performed on the audio segment corresponding to each text segment, which are explained in detail below.

[0273] In step 201, the audio segment is processed by speech recognition to obtain speech-recognized text.

[0274] In some embodiments, an automatic speech recognition system can be used to process audio segments for speech recognition to obtain speech-recognized text.

[0275] Here, Automatic Speech Recognition (ASR) is used to convert human speech into text. It is a speech processing technology that analyzes and decodes the features of speech signals to convert them into corresponding text.

[0276] In step 202, the quality parameters of the audio segment are determined based on the text segment and the speech recognition text.

[0277] In some embodiments, see Figure 4L , Figure 4K Step 202 shown can be achieved through steps 2021 to 2024, as explained in detail below.

[0278] In step 2021, using the text fragment as reference text, the number of missing words in the speech recognition text relative to the reference text and the number of replaced words in the speech recognition text relative to the reference text are obtained.

[0279] In some embodiments, firstly, the reference text and the speech recognition text are segmented into units at the word or character level. This allows for precise comparison of each unit in the two texts. Then, by comparing the segmentation results of the reference text and the speech recognition text, a mapping relationship is established for each word or character. This can be achieved using a dynamic programming algorithm, such as the Levenshtein distance algorithm, to determine the minimum edit distance between the two texts. In the mapping relationship, if a word or character in the reference text has no corresponding match in the speech recognition text, then that word or character is considered missing. The number of missing words or characters is obtained by counting all missing words or characters. In the mapping relationship, if a word or character in the reference text has a match in the speech recognition text, but the content is different, then that word or character is considered replaced. The number of replaced words or characters is obtained by counting all replaced words or characters.

[0280] For example, suppose the reference text is: "The weather is really nice today"; the speech recognition text is: "The air is really nice today". After word segmentation: the reference text is segmented as: ["today", "weather", "really nice", "."] and the speech recognition text is segmented as: ["today", "air", "really nice", "."]. Using a dynamic programming algorithm to establish a mapping, we can obtain the following mapping relationship: [today] corresponds to [today], [weather] corresponds to [air] (a substitution has occurred here), [really nice] corresponds to [really nice], [."] corresponds to [."]. Missing characters: none, number of missing characters: 0; replaced characters: weather corresponds to air, number of replaced characters: 1. That is, based on the comparison between the reference text and the speech recognition text, we can obtain that the number of missing characters is 0 and the number of replaced characters is 1.

[0281] In step 2022, a pre-set audio duration range is obtained based on the text length of the speech recognition text.

[0282] In some embodiments, a word-per-minute (Words per minute) is set at a normal speaking speed, for example, 300 Words per minute, corresponding to an average duration of 60 / 300 = 0.2 seconds per Word. Audio duration ranges can be represented as follows: Short text (1-10 words): 220 to 300 Words per minute, to avoid excessively fast speaking speed. Medium text (11-25 words): Maintaining a baseline speaking speed of 300 Words per minute. Long text (26 words or more): 300 to 350 Words per minute, to prevent excessively long total duration; that is, different word-per-minute ranges correspond to different text lengths for speech recognition.

[0283] In step 2023, the duration parameter of the audio segment is determined based on the duration of the audio segment and the audio duration range.

[0284] In some embodiments, the number of words per minute of an audio segment is obtained based on the duration of the audio segment and the text length of the speech-recognized text. If the number of words per minute of the audio segment is within the range of the corresponding audio duration interval, the duration parameter of the audio segment is 0. If the number of words per minute of the audio segment is not within the range of the corresponding audio duration interval, the duration parameter of the audio segment is 1.

[0285] In step 2024, the missing word count, the number of replaced words, and the duration parameter are weighted and summed to obtain the quality parameters of the audio segment.

[0286] In some embodiments, the missing word count, the number of replaced words, and the duration parameter are weighted and summed according to a pre-set weighting coefficient to obtain the quality parameters of the audio segment.

[0287] For example, the weighting coefficients can be expressed as: number of missing characters (0.4), number of replaced characters (0.4), and duration parameter (0.2).

[0288] Steps 2021 to 2024 are used to score the quality of audio segments (characterized by quality parameters, with larger quality parameters indicating lower audio segment quality), thus ensuring the quality of audio segments.

[0289] See also Figure 4K In step 203, in response to the quality parameter being greater than or equal to a preset quality parameter threshold, the process proceeds to combining the audio segments corresponding to each text segment to obtain the audio corresponding to the target text.

[0290] In some embodiments, when the quality parameter is greater than or equal to a preset quality parameter threshold, the process proceeds to step 106 above.

[0291] In step 204, in response to the quality parameter being less than a preset quality parameter threshold, the process switches to speech synthesis based on each text segment, the role in each text segment, and the timbre information of the role.

[0292] In some embodiments, when the quality parameter is less than a preset quality parameter threshold, the process proceeds to step 105 above.

[0293] In other embodiments, sentiment analysis is performed based on the content of the original text (text fragments) to obtain the sentiment category of the original text; features such as speech rate, tone, and volume of the audio fragments are extracted; pre-set speech rate range, tone range, and volume range (pre-set through empirical data) are obtained according to the sentiment category of the original text (and the gender of the character); if any of the speech rate, tone, or volume of the audio fragment exceeds a reasonable range, the current audio fragment fails the quality detection and proceeds to speech synthesis processing based on each text fragment, the character in each text fragment, and the character's timbre information.

[0294] In some embodiments, the text-to-audio method is implemented by calling processing nodes in the workflow computing framework; segmentation (i.e., segmenting the target text) is implemented by calling segmentation nodes in the workflow framework; the role information list is obtained by calling role information extraction nodes in the workflow framework; roles are implemented by role recognition nodes in the workflow framework; timbre information is obtained by timbre allocation nodes in the workflow computing framework; and audio segments are synthesized by dubbing nodes in the workflow framework. See also Figure 4M Before segmenting the target text into multiple text fragments, the following steps 301 to 304 can be performed, which are explained in detail below.

[0295] In step 301, the workflow settings interface is displayed.

[0296] In some embodiments, the workflow settings interface provides a visual, configurable operating platform, thereby breaking down the text-to-audio task flow into manageable steps and automating the process through flexible parameter adjustments and logical orchestration. For example, multi-step tasks (such as text segmentation, character information list extraction, character recognition, timbre information determination, speech synthesis, data export, etc.) can be integrated into a repeatable pipeline.

[0297] In step 302, in response to the first setting operation in the workflow settings interface, multiple processing nodes are set in the workflow settings interface.

[0298] In some embodiments, in response to a first setting operation in the workflow settings interface, a text segmentation node (for performing step 101 above), a role extraction node (for performing step 102 above), a role affiliation judgment node (for performing step 103 above), a timbre allocation node (for performing step 104 above), a dubbing node (for performing step 105 above), a quality detection node (for performing steps 201 to 204 above), and a data export node (for performing step 106 above) are set in the workflow settings interface.

[0299] In step 303, in response to a second setup operation for each processing node, working logic is applied to each processing node, wherein the working logic is used to perform specific processing on the target text.

[0300] Continuing from the previous example, the workflow logic (i.e., the processing logic for each step at each processing node) can be encapsulated as an independent module, supporting hot-swapping (e.g., a faster speech synthesis engine can be replaced at the dubbing node). Each processing node can access upstream data (e.g., the role attribution judgment node uses the content output by the role extraction node). Through this mechanism, each processing node in the workflow can accurately apply specific logic to process the target text according to the user-configured "secondary setting operation," ultimately achieving end-to-end automated text-to-audio conversion.

[0301] In step 304, in response to a connection operation for multiple processing nodes, the multiple processing nodes are connected into a workflow framework, wherein, for any two processing nodes that are connected, the output of the preceding processing node is the input of the following processing node.

[0302] In some embodiments, the output port of a preceding node (such as Node A) is compatible with the input port data of a subsequent node (such as Node B). During system execution, the system automatically resolves node dependencies, and each processing node executes its work logic sequentially, calling the processing function of each node in turn and passing upstream data.

[0303] Through steps 301 to 304, based on the workflow, after the target text is input into the text-to-audio system, the data between each step is automatically transferred without manual intervention, thus improving the conversion efficiency of text to audio.

[0304] The text-to-audio method provided in this application embodiment can be applied to various scenarios that require text-to-audio conversion, such as virtual assistants and chatbots, podcasts and audiobooks, accessibility assistance, etc. The application of the text-to-audio method provided in this application embodiment in the audiobook production scenario is described below.

[0305] The audiobook production process mainly includes the following key steps. First is text preprocessing, which involves acquiring the text content, checking and correcting text formatting, punctuation, etc., and dividing the text into chapters for subsequent operations. Next is character analysis, similar to dialogue attribution techniques in novels, determining character information in the text, including personality and identity. This helps assign appropriate voices to different characters, enhancing the story's appeal and comprehensibility, especially important in multi-character audiobooks. Then comes the voice synthesis or dubbing stage. If voice synthesis is used, a suitable voice library and synthesis software need to be selected; if human dubbing is used, suitable voice actors need to be selected and their performances tailored to the character's traits. Finally, post-production begins, adjusting audio volume and speaking speed, adding background music and sound effects, such as adding atmospheric music to emotionally charged scenes and transitional sound effects during scene changes, making the audiobook more attractive and professional, ultimately resulting in a high-quality audiobook product. The audiobook production systems in related technologies can be mainly divided into the following four categories:

[0306] 1) Purely human voice-over production method: Professional voice actors dub the audiobook sentence by sentence. However, this production method is expensive and the recording work often takes a long time to process a large amount of text content, making it difficult to quickly launch audiobook products.

[0307] 2) Simple speech synthesis software production method: Using some basic speech synthesis software, the text is input and the speech is generated directly. However, the synthesized speech has a mechanical and stiff tone, lacks emotion and naturalness, and sounds like a robot reading. It is difficult for the audience to have an emotional resonance, which affects the quality of the audiobook.

[0308] 3) Semi-human, semi-technical production method: This method combines human voice-over with audio editing software and simple voice processing techniques. For example, voice actors record the main dialogue first, and then software is used to edit, splice, and add sound effects to improve the overall quality of the audiobook. However, the human voice-over portion is still limited by the cost and time of voice actors. The degree of technological assistance is limited, making it difficult to achieve accurate character extraction and intelligent timbre allocation.

[0309] 4) Traditional template-based speech synthesis production method: Based on a pre-defined speech template library, appropriate templates are selected for speech synthesis according to the type and style of the audiobook text. For example, there are specific speech style templates for different genres such as martial arts novels and romance novels. During production, the text is adapted to these templates. However, due to the significant limitations of the templates, it is difficult to meet the diverse and personalized needs of audiobooks. Once the audiobook content exceeds the preset range of the templates, such as the appearance of innovative plots or unique character settings, the speech synthesis effect will be greatly reduced, resulting in an audiobook that lacks appeal and realism.

[0310] This application provides an audiobook production system based on a large language model and TTS technology (corresponding to the text-to-audio system mentioned above), which can achieve efficient, accurate and low-cost audiobook production. Its key function is to build an automated and intelligent production process by using a large language model and TTS technology to achieve accurate processing of characters, dialogues, timbre and other aspects, thereby completing high-quality audiobook production. The audiobook production system includes data preprocessing nodes, character extraction nodes, dialogue extraction and attribution determination nodes, timbre allocation nodes, speech synthesis nodes, quality detection nodes, and large language model adaptation and invocation nodes. Each node has its own unique technical implementation. The data preprocessing node receives novel text by chapter and combines paragraphs into smaller paragraphs using rules such as double quote matching and length limit processing. Multiple paragraphs form a small chapter storage unit as the basic input unit for the system. It also constructs and maintains basic data such as novel type, atmosphere enumeration (used to match background voice and sound effects, and can also affect the emotional effect during TTS), emotion enumeration, character attribute enumeration, and timbre library, ensuring the system's data foundation is complete and accurate, providing strong data support for the smooth operation of subsequent functional nodes. The character extraction node uses part-of-speech analysis to locate character information, infers attributes through semantic understanding, and outputs a pre-defined format (JavaScript Object) after model training. The data is in Notation (JSON); the dialogue extraction and attribution judgment node uses natural language processing technology to analyze sentence structure and uses machine learning or deep learning models to determine the corresponding role and emotion of the sentence; the timbre allocation node selects the timbre for each role based on pre-set matching rules between roles and timbres and a pre-trained large language model; the speech synthesis node generates corresponding audio content according to pre-set speech synthesis rules and algorithms, which effectively ensures high-quality audiobook production; the quality inspection node is used to perform quality inspection and evaluation of the audio content of the audiobook; the large language model adaptation and invocation node is responsible for deeply integrating the large model into the workflow framework, ensuring that each workflow node can interact efficiently with the large language model.

[0311] It should be noted that the aforementioned processing nodes can also be called modules, which are abstractions of a processing stage in the workflow. Nodes can be deployed on one or more electronic devices.

[0312] Below, in conjunction with Figure 6 The process shown in the diagram will be explained in detail.

[0313] Step 11: Access the novel resource database.

[0314] For example, first, design the database model to determine the data structure to be stored, such as the novel's basic information (title, author, category, etc.), content, chapters, etc. Then, build the database on the server, such as MySQL or PostgreSQL. Enter the novel resources into the database, either manually or by writing a program to automatically import them from other data sources (such as text files, other databases, etc.). Design an Application Programming Interface (API) so that the audiobook production system can interact with the database through these interfaces to perform CRUD operations. During database calls, server-side scripting languages ​​such as Java, Python, and PHP can be used to write backend logic to handle requests from the frontend, interact with the database, and return processing results. In the frontend application (such as a web page or mobile application), JavaScript, jQuery, etc., can be used to write code to interact with the backend API through asynchronous JavaScript and XML (JSON can be used instead of XML as the data exchange format) technologies to retrieve or send novel data. Throughout the entire process, Hypertext Transfer Protocol Secure (HTTPS) can be used to encrypt communication, verify user identity, and prevent attacks such as SQL injection, thereby ensuring data security.

[0315] Step 12: Based on the novel text content extracted from the novel resource database, call the text segmentation node.

[0316] In some embodiments, in response to retrieving specified novel text content from a novel resource database (e.g., calling novel text content with a specified name from a novel resource database), a text segmentation node is invoked to perform text segmentation processing on the novel text content, resulting in multiple text fragments.

[0317] For example, the input to the text segmentation node is the novel text content preprocessed by the preprocessing node, and the output is the text fragments segmented according to specific rules, such as a list of text blocks or a list of sentences in paragraph form.

[0318] For example, text segmentation algorithms, such as punctuation-based segmentation methods (periods, question marks, exclamation marks, etc.) or sentence segmentation functions based on natural language processing libraries (such as NLTK, Spacy, etc.), are used to divide the novel text into independent segments. Simultaneously, a unique identifier (such as a segment number) is added to each segment, and the segment's position information in the original text (start line number, end line number, etc.) is recorded for subsequent nodes to associate and process.

[0319] Step 13: Save the output of the text segmentation nodes to the text fragment database.

[0320] In some embodiments, the text fragments output by the text segmentation node are saved to a text fragment database.

[0321] Step 14: Based on the input from the text fragment database, call the role extraction node.

[0322] In some embodiments, in response to receiving multiple text fragments as input from a text fragment database, a character extraction node is invoked to extract a list of character information from the multiple text fragments, wherein the list of character information includes character information for each character in the novel's text content.

[0323] Here, the function of the character extraction node is to input the language model into a small chapter (which can be a small chapter composed of text fragments), and to deeply analyze and extract character information based on various predefined information, including character name, character characteristics (such as personality traits, gender, identity, and age range), and character introduction, and output it in JSON array format to provide basic character data for subsequent processing. For the specific method of extracting the character information list, please refer to the explanation in step 102 above, which will not be repeated here.

[0324] For example, the input to the character extraction node is: the text fragments output by the text segmentation node (which can be input into the large language model in units of small chapters, and a small chapter can include multiple text fragments); the output is: a list of character information extracted from the text fragments (or small chapters) (including the character name, character characteristics, and character introduction for each character).

[0325] For example, by inputting a text fragment into a language model and specifying the task as character name extraction (corresponding to obtaining the first cue word above, where the first cue word indicates the extraction of a list of character information from the target text), the language model identifies and classifies entities such as character names and titles in the text, returning a list of character information. Furthermore, it can analyze descriptive information about the characters, such as physical appearance and clothing descriptions, providing more clues for subsequent dialogue attribution determination. For the custom rule approach, some common character name keywords and title patterns can be predefined, and a text matching algorithm can be used to match character information within the text fragment.

[0326] Step 15: Save the output of the character extraction node to the character information database.

[0327] In some embodiments, the character information extracted by the character extraction node is saved to the character information database.

[0328] For example, design a character information database table, and write the JSON array output from the character extraction nodes into the character information database according to the character information database table. The character information database table can include the following fields: id (a unique identifier for each character); name (character name); age (age range); gender; traits (character traits); description (character description); created_at (creation time); updated_at (update time).

[0329] Step 16: Perform post-processing on the newly added characters extracted from the character extraction node.

[0330] In some embodiments, when the character extracted by the character extraction node is not in the pre-set character name list (e.g., the character name of the character is not stored in the pre-set character name list), the character information of the newly extracted character is post-processed, that is, the character information of the newly extracted character extracted by the character extraction node is converted from JSON data into the structure of database records according to the character information database table.

[0331] In other embodiments, when the character extraction node extracts new character information for a certain character, the extracted new character information is processed by character post-processing, that is, the new character information extracted by the character extraction node is converted from JSON data into the structure of database records according to the character information database table.

[0332] Step 17: Save the newly added characters after post-processing to the character database.

[0333] In some embodiments, a database connection pool or a single connection is used to ensure that the connection parameters are correct (such as host, port, username, password, database name, etc.) and the post-processed role information of the newly added role is inserted into the role database table.

[0334] Step 18: Based on the text fragments and extracted role information, call the dialogue recognition and attribution judgment node.

[0335] In some embodiments, the dialogue recognition and attribution determination node is invoked to perform the following processing on each text segment: based on the role information list, the text segment, and the context of the text segment, a pre-trained language model (corresponding to the second language model mentioned above) is invoked to predict the role corresponding to the text segment; the text segment is processed for dialogue text detection and narration text detection to obtain the dialogue text detection results and narration text detection results; and the text segment is processed for emotion parameter extraction.

[0336] For example, the specific implementation of predicting the role corresponding to a text fragment using a language model can be found in the description of step 1032 above. The specific implementation of performing dialogue text detection processing and narration text detection processing on the text fragment can be found in the description of step 1051 above, and will not be repeated here.

[0337] For example, the pre-configured emotion recognition model can be invoked through the dialogue recognition and attribution judgment node to extract emotion parameters from the text fragment, thereby obtaining the emotion parameters of the text fragment. The emotion recognition result is used to characterize the emotion type and emotion intensity of the text fragment.

[0338] For example, emotion types include positive, negative, and neutral, and emotion intensity includes strong and moderate. The emotion recognition results can be quantified to obtain the emotion parameters of the text fragment. For example, positive emotion, negative emotion, and neutral emotion can be represented by "+", "-", and a space, respectively. Strong emotion is represented by 1, and moderate emotion is represented by 0. For example, "+1" represents strong positive emotion, and "0" represents moderate neutral emotion.

[0339] It's important to note that emotion type refers to the tendency of an emotion, such as positive, negative, and neutral emotion types. For example, positive emotion types can include happiness, excitement, and exhilaration, while negative emotion types can include sadness, shame, and anger. Emotion intensity is an indicator used to quantify the strength of an emotion; generally, the intensity of a normal emotion is lower than that of a strong emotion.

[0340] Taking the training of a deep learning-based emotion recognition model as an example, texts with strong positive emotions can be manually labeled as positive emotion samples, texts with strong negative emotions as negative samples, and other text samples as neutral emotion samples. The labeled and preprocessed text data can then be used for model training.

[0341] For example, based on the Transformer model framework, the model can learn how to identify the emotion category (e.g., positive, negative, and neutral) and emotion intensity (e.g., strong, moderate) of text fragments. The performance of the trained model can be evaluated using evaluation set data, such as by calculating metrics like accuracy and precision. During the inference phase, given a new text fragment, the emotion recognition model will predict the emotion category and emotion intensity of the text fragment.

[0342] Here, the dialogue extraction and attribution judgment node function is as follows: process the text by sentence (a text segment includes at least one sentence), detect the role and emotion to which the sentence belongs, and return the role name and emotion information (including emotion type and intensity) in JSON format to provide emotional parameters for the dubbing process and enhance the expressiveness of the audiobook.

[0343] For example, the input to the dialogue recognition and attribution judgment node is: the text fragment output by the text segmentation node and the list of character information output by the character extraction node; the output is: the character identifier of each text fragment (including the identifier of whether a certain sentence in the text fragment is narration text) and confidence information.

[0344] For example, text fragments are integrated with contextual information (such as preceding and following paragraphs, adjacent dialogues, etc.) to construct input instructions for the language model (e.g., the second prompt word obtained above). For instance, given dialogue statements (i.e., the dialogue text in the text fragment), contextual text, and a list of character information as input, the large language model is asked to determine the most likely character to which the dialogue belongs. Based on its language understanding and semantic analysis capabilities, the large language model comprehensively considers factors such as the dialogue's language style, vocabulary usage habits, and the degree of matching with character features, providing a character attribution judgment result along with a confidence score indicating the reliability of the judgment. After receiving the output of the large language model, the dialogue attribution judgment node organizes and records the results, associating and storing the dialogue statements with the corresponding character identifiers and confidence information.

[0345] Step 19: Save the output of the dialogue recognition and attribution judgment node to the dialogue / narration database.

[0346] In some embodiments, the role identifier (including an identifier indicating whether a certain statement in the text fragment is narration text) and confidence information of each text fragment are saved to the dialogue / narration database.

[0347] Step 20: Call the timbre database.

[0348] In some embodiments, the timbre database (i.e., timbre library) stores timbre-related information, including timbre unique identifiers, age, gender, timbre tags, etc., to provide a rich timbre for subsequent character timbre assignment.

[0349] Step 21: Based on the timbre database and the role database, call the timbre allocation node.

[0350] In some embodiments, the timbre information corresponding to each character is determined based on the character information of each character in the character database and the timbre information in the timbre database (different timbre information is represented by different timbre identifiers).

[0351] For example, the specific implementation of determining the timbre information of each character can be found in the description of step 104 above, and will not be repeated here.

[0352] For example, the voice assignment node plays a crucial role in audiobook production systems. Its core function is to intelligently select suitable voices from a voice library based on the characteristics of the characters. These character characteristics encompass many aspects, such as the character's identity (protagonist, villain, supporting character, or extra), personality (e.g., calm, lively, introverted, extroverted), and age (child, teenager, adult, elderly). By comprehensively considering these character attributes, the system accurately matches each character with the most fitting voice, making the audiobook's voice presentation more reasonable and vivid. This helps enhance listeners' recognition of the characters and their immersion in the story. For instance, a mature and dignified voice might be assigned to a calm protagonist, while a clear and innocent voice might be matched to a lively and adorable child character, giving each character a distinct personality at the vocal level.

[0353] For example, the input to the timbre allocation node is: the character information and timbre library output by the character extraction node; the output is: the character and the timbre ID corresponding to the character.

[0354] Step 22: Save the output of the timbre assignment node to the character-sound database.

[0355] In some embodiments, the character and the corresponding timbre ID are saved to the character-voice database.

[0356] Step 23: Based on the dialogue / narration database and the character-voice database, call the voice-over node.

[0357] In some embodiments, the voice-over node includes a dialogue voice-over node and a narration voice-over node.

[0358] For example, the input to the dialogue dubbing node is: the dialogue text output by the dialogue attribution judgment node and the corresponding character identifier, and the voice corresponding to the character output by the voice assignment node; the output is: the audio of the dialogue text (i.e., the dialogue audio segment) and the corresponding dialogue text. The dialogue dubbing node utilizes text-to-speech (TTS) technology to combine the dialogue text corresponding to the character identifier with the relevant voice information, and generates corresponding audio content according to pre-set speech synthesis rules and algorithms. This provides audio material for the dialogue sections of the audiobook, ensuring that the dialogue is presented to the listener with appropriate voices and vivid speech, thus enhancing the overall auditory experience and storytelling of the audiobook.

[0359] For example, the specific implementation of the dialogue voice generation of the dialogue dubbing node can be found in the description of steps 10531 to 10535 above, and will not be repeated here.

[0360] For example, the input to the narration / dubbing node is the narration text output by the dialogue attribution judgment node, along with a specified narration voice (i.e., narration voice audio) configured in the system. The output is: the audio of the narration text (i.e., the narration audio clip) and the corresponding narration text. The narration / dubbing node identifies all content other than the dialogue text as narration text, and then uses text-to-speech (TTS) technology and the specified narration voice configured in the system to generate the corresponding audio content, providing audio material for the subsequent narration portion of the audiobook.

[0361] For example, the specific implementation of the narration voice generation of the narration voice node can be found in the description of steps 10521 to 10525 above, and will not be repeated here.

[0362] For example, the speech synthesis node aims to utilize advanced TTS technology, using text content, character information, and assigned timbres as basic elements to synthesize high-quality speech. It can also flexibly adjust the emotional features of the synthesized speech by referencing emotional parameters, ultimately achieving a captivating dubbing effect. These dubbing contents are then systematically integrated to complete the production of an audiobook. This functionality transforms audiobooks from mechanical text readings into vivid presentations based on plot, character emotions, and other factors, allowing listeners to better immerse themselves in the story world created by the audiobook and significantly improving the overall quality and auditory experience. Specifically, the TTS node generates speech based on the text's inherent prosodic features (such as pauses, intonation variations, and stress placement, derived from grammatical structure, semantic understanding, and contextual analysis), combined with previously extracted emotional parameters (such as emotion categories like anger and joy, and emotional intensity), and a selected appropriate timbre, using advanced speech synthesis algorithms.

[0363] Step 24: Save the output of the dubbing node to the dubbing result database to be tested.

[0364] In some embodiments, the dialogue text and audio of the dialogue text output by the dubbing node, as well as the narration text and audio of the narration text, are saved to the dubbing result database to be detected.

[0365] Step 25: Based on the database of dubbing results to be detected, call the quality detection node.

[0366] In some embodiments, the input to the quality inspection node is the audio output from the dubbing node and the corresponding text; the output is the audio score. This node is mainly used to perform quality inspection on the audio output by the TTS system. For example, the audio is input into the ASR system to obtain the recognized text content, and the audio is scored based on the original text and the recognized text content. If the score meets the requirements, the audio is adopted; otherwise, it is sent back to the dubbing node for reprocessing.

[0367] For example, the primary function of the ASR-based quality inspection node is to perform quality inspection and evaluation of the audio content of audiobooks. It can convert speech to text and compare the converted text with the original audiobook script to check for omissions, misreadings, or extra readings, thus determining the accuracy of the speech content. Simultaneously, it can analyze aspects such as speech clarity and fluency, detecting issues like unclear pronunciation, pauses, and disjointedness to ensure the basic quality of the audiobook's audio. Furthermore, this node can assess whether the speech's speed, tone, and volume match the style and emotional expression requirements of the audiobook. For instance, it can determine whether the speed and tone are appropriate during tense scenes and whether the volume is moderate during lyrical passages, thereby improving the overall auditory effect and artistic appeal of the audiobook and ensuring that the audiobook achieves a high level of accuracy in both content and speech performance.

[0368] For example, firstly, the quality inspection node receives the audio output from the dubbing node as input, uses ASR technology to recognize the speech, and converts it into text. Then, it compares the converted text with the pre-stored original text word by word, marking inconsistencies such as mispronounced words. Based on these analysis results, combined with preset quality assessment standards, such as a reasonable range for allowable error rates, the speech quality is quantitatively scored. Finally, a detailed quality inspection report is generated, including the types and number of speech errors to ensure the accuracy and traceability of the quality inspection.

[0369] For example, the specific implementation of quality inspection can be found in the description of steps 201 to 204 above, and will not be repeated here.

[0370] Step 26: Save the dubbing results that pass the test to the final dubbing result database.

[0371] In some embodiments, in response to the quality detection of the dubbing result (i.e., the audio output by the dubbing node) passing, the dubbing result is saved to the final dubbing result database.

[0372] Step 27: Transfer the dubbing results that fail the test to the dubbing node and dub them again.

[0373] In some embodiments, in response to the failure of the quality detection of the dubbing result (i.e., the audio and corresponding text output by the dubbing node), the text corresponding to the dubbing result is transferred to the dubbing node for re-dubbing.

[0374] For example, the narration audio clips and dialogue audio clips corresponding to each text segment are combined according to their order of appearance in the text segment to obtain the audio clips corresponding to each text segment. The audio clips corresponding to each text segment are then combined to obtain the audio corresponding to the novel's text content, thus completing the production of the audiobook.

[0375] Through steps 11 to 27, this application provides an audiobook production system based on a large language model and TTS technology, which can achieve efficient, accurate, and low-cost audiobook production. Its key function lies in building an automated and intelligent production process by leveraging a large language model and TTS technology, and has the following features:

[0376] Beneficial effects:

[0377] Sufficient Data Foundation: Audiobook production requires a complete and accurate data foundation to support the operation of each stage. In the data preprocessing stage, this application receives novel texts by chapter, constructs small chapter storage units, and maintains a series of basic data such as novel type, atmosphere enumeration, emotion enumeration, character attribute enumeration, and voice library. This provides a solid data guarantee for the smooth operation of the entire audiobook production system and avoids problems such as missing or inaccurate data affecting the production effect.

[0378] Cost Reduction: Traditional audiobook production relies heavily on human voice-over, resulting in high costs. This application's embodiment reduces reliance on professional voice-over personnel, optimizes resource allocation through a technology-driven model, reduces labor costs, and effectively solves the problem of high audiobook production costs. This enhances the industry's economic efficiency and competitiveness, making audiobook production more cost-effective. The cost is less than 1% of the cost of manual production processes.

[0379] Improved Production Efficiency: Previously, audiobook production suffered from low efficiency. For example, when human voice-overs were faced with large amounts of text, the recording cycle was long, making it difficult to quickly respond to market demands. This application's embodiments utilize an automated and intelligent production process, with efficient collaboration across all stages, significantly reducing manual operations and substantially shortening the audiobook production cycle, thus enabling timely fulfillment of the market's demand for rapid audiobook output.

[0380] Precise character extraction: Related technologies have shortcomings in precise character extraction. However, the embodiments of this application input text into a large language model in units of small chapters. Based on the original text of the chapters and predefined enumerated information such as personality, identity, and age, the system deeply analyzes and extracts character information, covering multi-dimensional attributes of the character. The system outputs the information as a JSON array, providing accurate basic data of the character for subsequent steps. This achieves precise character extraction and overcomes the previous difficulty in accurately obtaining relevant character information.

[0381] Accurate Dialogue Recognition and Attribution: In terms of dialogue recognition and attribution, existing technologies struggle to accurately determine the character to whom a dialogue belongs. This application's embodiment processes text on a sentence-by-sentence basis, combining a character list and dialogue context information with a large language model to determine attribution. Simultaneously, it analyzes the dialogue's emotional state, accurately outputting character names and emotion-related parameters in JSON format. This provides a basis for dubbing, effectively solving the problems of inaccurate dialogue attribution and emotional grasp, and optimizing the content logic of audiobooks.

[0382] Intelligent timbre allocation: In terms of timbre allocation, previous methods required manual selection of suitable timbres for characters based on book content, which was costly and inefficient. This application's embodiment accumulates sufficient character profile information during character extraction and, combined with a timbre library, can match suitable timbres for each character, significantly improving efficiency while ensuring the diversity of character timbres.

[0383] Dubbing quality is guaranteed: A scoring rule based on the ASR model is designed to verify the audio quality generated by the TTS system. First, the audio is input into the ASR model to obtain the text content of the audio. By comparing it with the original text, points are deducted at different weights for missing words, replaced words, and whether the audio duration and text length are within the corresponding range (e.g., in a text of 5 words and a text of 20 words, the audio of 5 words is shorter and the audio of 20 words is longer). Then, the score of this audio segment is obtained. If the score meets the passing score, it passes the verification and the audio is adopted. If it does not meet the passing score, it is sent back to the generation stage for regeneration.

[0384] Fully automated process: Based on the workflow, after the novel content is input into the system, the data between each step flows automatically without human intervention. Finally, a structured audiobook dataset is generated and stored in the computer storage medium. Developers only need to interact according to the standardized protocol to develop an audiobook system for users.

[0385] The following description continues to illustrate the exemplary structure of the text-to-audio device 433 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the text-to-audio device 433 in the memory 430 may include:

[0386] The text segmentation module 4331 is used to segment target text into multiple text fragments.

[0387] The character information extraction module 4332 is used to extract a character information list from the target text, wherein the character information list includes the character information of each character in the target text.

[0388] The role recognition module 4333 is used to identify the role in each of the text fragments based on the role information list.

[0389] The timbre allocation module 4334 is used to determine the timbre information of each character based on the character information of each character.

[0390] The speech synthesis module 4335 is used to perform speech synthesis based on each text segment and the timbre information of the character in each text segment to obtain an audio segment corresponding to each text segment.

[0391] In some embodiments, the speech synthesis module 4335 is further configured to combine the audio segments corresponding to each text segment to obtain the audio corresponding to the target text.

[0392] In some embodiments, the character information extraction module 4332 is further configured to: obtain a first prompt word, wherein the first prompt word indicates that the character information list is extracted from the target text; divide the target text into multiple text chapters, wherein the granularity of the text chapter division is greater than the granularity of the text segment division; based on the first prompt word and the multiple text chapters, call a pre-trained first language model to extract at least one character information in each text chapter; and combine the at least one character information in each text chapter into a character information list according to the order of the multiple text chapters in the target text.

[0393] In some embodiments, each text chapter includes at least one text fragment. The role information extraction module 4332 is further configured to perform entity recognition processing on the text chapter to obtain at least one role name; perform referential resolution processing on the text chapter based on at least one role name to obtain a referential resolution chapter; obtain at least one text fragment corresponding to each role name in the referential resolution chapter as the role description text corresponding to each role name; perform part-of-speech tagging on each word in each role description text to obtain the part-of-speech tagging result of each role description text; obtain the role features corresponding to each role name based on the part-of-speech tagging result of each role description text; perform summary text generation processing on the role description text corresponding to each role name to obtain the role introduction corresponding to each role name; and combine each role name, the role features corresponding to each role name, and the role introduction to form the role information of each role.

[0394] In some embodiments, the role information extraction module 4332 is further configured to divide the target text into multiple text chapters, wherein each text chapter includes at least one text segment; and perform the following processing on each text chapter: extract role information from the text chapter according to a pre-set role matching template to obtain at least one role information in the text chapter; and combine the at least one role information in each text chapter into a role information list according to the order of the multiple text chapters in the target text.

[0395] In some embodiments, the role recognition module 4333 is further configured to perform the following processing on each text segment: obtain a second prompt word, wherein the second prompt word indicates that the role corresponding to the text segment is identified; and predict the role corresponding to the text segment by invoking a pre-trained second language model based on the second prompt word, the role information list, the text segment, and the context of the text segment.

[0396] In some embodiments, the role recognition module 4333 is further configured to concatenate the role information list, the text fragment, and the context of the text fragment to obtain role prediction text; perform feature encoding on the role prediction text to obtain role prediction features; perform feature mapping on the role prediction features to obtain the prediction probability value of each role in the role information list; and select the role with the highest prediction probability value as the role corresponding to the text fragment.

[0397] In some embodiments, the timbre allocation module 4334 is further configured to: query a plurality of candidate timbre features that match the role information of the role from a timbre library according to a preset timbre matching rule and the role information of the role; obtain a third prompt word, wherein the third prompt word indicates the timbre features corresponding to the role to be filtered out; based on the third prompt word, the role information of the role and the plurality of candidate timbre features, filter out the timbre features of the role from the plurality of candidate timbre features, and use the timbre features and the audio that matches the timbre features as the timbre information of the role.

[0398] In some embodiments, the character information includes character name, character characteristics, and character introduction. The character characteristics include at least one of the following: character gender, character age range, character personality, and character identity. The timbre allocation module 4334 is further configured to: query multiple timbre feature samples that match the character age range and character gender from the timbre library; obtain the matching weights corresponding to the character personality, character identity, and character introduction from the timbre matching rules; and take the one with the highest matching weight among the character personality, character identity, and character introduction as the information to be matched; and query timbre feature samples that match the information to be matched from the multiple timbre feature samples as the candidate timbre features.

[0399] In some embodiments, the timbre allocation module 4334 is further configured to invoke a pre-trained third language model based on the third prompt word, and to select the timbre features of the character from the plurality of candidate timbre features according to the character information.

[0400] In some embodiments, the audio segment includes at least one of a narration audio segment and a dialogue audio segment. The speech synthesis module 4335 is further configured to perform the following processing for each text segment: perform narration text detection processing on the text segment to obtain a narration detection result, and perform dialogue text detection processing on the text segment to obtain a dialogue detection result; in response to the narration text result indicating that there is narration text in the text segment, perform narration speech generation processing based on the narration text and a pre-set narration timbre audio to obtain the narration audio segment; in response to the dialogue detection result indicating that there is dialogue text in the text segment, perform dialogue speech generation processing based on the dialogue text and the timbre information to obtain the dialogue audio segment.

[0401] In some embodiments, the timbre information includes timbre features and audio matching the timbre features. The speech synthesis module 4335 is further configured to extract first text features from the dialogue text to obtain first text features; extract first timbre features from the audio matching the timbre features to obtain first timbre features; perform feature fusion based on the first text features and the first timbre features to obtain first fused features; perform first audio feature prediction based on the first fused features to obtain first audio features; and perform first audio feature decoding on the first audio features to obtain the dialogue audio segment.

[0402] In some embodiments, the speech synthesis module 4335 is further configured to extract second text features from the narration text to obtain second text features; extract second timbre features from the narration audio to obtain second timbre features; perform feature fusion based on the second text features and the second timbre features to obtain second fused features; predict second audio features based on the second fused features to obtain second audio features; and decode the second audio features to obtain the narration audio segment.

[0403] In some embodiments, the speech synthesis module 4335 is further configured to perform the following processing on the audio segment corresponding to each text segment: perform speech recognition processing on the audio segment to obtain speech-recognized text; determine the quality parameters of the audio segment based on the text segment and the speech-recognized text; in response to the quality parameters being greater than or equal to a preset quality parameter threshold, proceed to the processing of combining the audio segments corresponding to each text segment to obtain the audio corresponding to the target text; in response to the quality parameters being less than the preset quality parameter threshold, proceed to the processing of speech synthesis based on each text segment, the role in each text segment, and the timbre information of the role.

[0404] In some embodiments, the speech synthesis module 4335 is further configured to: use the text segment as a reference text to obtain the number of missing characters in the speech recognition text relative to the reference text, and to obtain the number of replaced characters in the speech recognition text relative to the reference text; obtain a pre-set audio duration interval based on the text length of the speech recognition text; determine the duration parameter of the audio segment based on the duration of the audio segment and the audio duration interval; and perform a weighted summation of the number of missing characters, the number of replaced characters, and the duration parameter to obtain the quality parameter of the audio segment.

[0405] In some embodiments, the text segmentation module 4331 is further configured to perform text cleaning processing on the target text to obtain cleaned target text; divide the cleaned target text into multiple line texts according to the newline characters in the cleaned target text; store the target text as structured target text according to each line text and the line number of each line text; and transfer the target text into the processing of segmenting the target text into multiple text fragments based on the structured target text.

[0406] In some embodiments, the text segmentation module 4331 is further configured to segment the structured target text according to a preset text segmentation rule to obtain the plurality of text segments, wherein each text segment corresponds to a text segment information, and the text segment information includes a text segment number and a line number of the text included in the text segment.

[0407] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the text-to-audio method described above in this application.

[0408] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the text-to-audio method provided in this application. For example, ... Figure 4A The text-to-audio conversion method is shown.

[0409] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0410] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0411] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0412] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0413] In summary, through the embodiments of this application, a list of character information is extracted from the target text, realizing the structured storage of character information scattered throughout the target text. This avoids the inefficiency of manual sentence-by-sentence annotation. By identifying the character corresponding to each text segment based on the character information list, accurate character allocation is achieved, improving the targeting and effectiveness of text analysis. By determining the timbre information of each character based on their character information, the timbre performance of the characters is optimized, thereby enhancing the realism and immersion of subsequent speech synthesis and improving the audio conversion effect of target texts with multiple characters.

[0414] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for converting text to audio, the method comprising: The method includes: Segment the target text into multiple text fragments; Extract a list of character information from the target text, wherein the list of character information includes character information for each character in the target text; Based on the list of character information, identify the character in each of the text fragments; The timbre information of each character is determined based on the character information of each character. Speech synthesis is performed based on each text segment and the timbre information of the character in each text segment to obtain an audio segment corresponding to each text segment; The audio segments corresponding to each text segment are combined to obtain the audio corresponding to the target text.

2. The method of claim 1, wherein, The step of extracting the list of character information from the target text includes: Obtain a first prompt word, wherein the first prompt word indicates that the list of character information be extracted from the target text; The target text is divided into multiple text sections, wherein the granularity of the text section division is greater than the granularity of the text segment division; Based on the first prompt word and the multiple text chapters, a pre-trained first language model is invoked to extract at least one character information in each of the text chapters; The at least one character information in each of the text chapters is combined into a character information list according to the order of the multiple text chapters in the target text.

3. The method of claim 2, wherein, Each of the text chapters includes at least one of the text segments, and the step of extracting at least one character information from each of the text chapters by calling a pre-trained first language model based on the first prompt word and the plurality of text chapters includes: Entity recognition processing is performed on the text chapters to obtain at least one character name; Based on at least one of the character names, the text chapter is subjected to referential resolution processing to obtain a referential resolution chapter; Obtain at least one text fragment corresponding to each character name in the chapter on the resolution of references, and use it as the character description text corresponding to each character name; Part-of-speech tagging is performed on each word in each of the character description texts to obtain the part-of-speech tagging results for each of the character description texts; Based on the part-of-speech tagging results of each character description text, obtain the character features corresponding to each character name; Summarize the character description text corresponding to each character name to obtain a character introduction for each character name; The character information for each character is composed of each character name, the corresponding character characteristics, and the character description.

4. The method of claim 1, wherein, The step of extracting the list of character information from the target text includes: The target text is divided into multiple text sections, wherein each text section includes at least one text fragment; Perform the following processing on each of the aforementioned text sections: The text chapter is processed to extract character information according to a pre-set character matching template, so as to obtain at least one character information in the text chapter. The at least one character information in each of the text chapters is combined into a character information list according to the order of the multiple text chapters in the target text.

5. The method of claim 1, wherein, The process of identifying the character in each text fragment based on the character information list includes: Perform the following processing on each of the text fragments: Obtain a second prompt word, wherein the second prompt word indicates the role corresponding to the text fragment; Based on the second prompt word, the character information list, the text fragment, and the context of the text fragment, a pre-trained second language model is invoked to predict the character corresponding to the text fragment.

6. The method of claim 5, wherein, The second language model, pre-trained based on the second prompt word, the character information list, the text fragment, and the context of the text fragment, predicts the character corresponding to the text fragment, including: The character information list, the text fragment, and the context of the text fragment are concatenated to obtain the character prediction text; The character prediction text is feature-encoded to obtain character prediction features; The predicted features of the characters are mapped to obtain the predicted probability value of each character in the character information list; The character with the highest predicted probability value is selected as the character corresponding to the text segment.

7. The method of claim 1, wherein, The process of determining the timbre information of each character based on the character information of each character includes: According to the pre-set timbre matching rules and the role information of the character, query multiple candidate timbre features that match the role information of the character from the timbre library; Obtain a third prompt word, wherein the third prompt word indicates the vocal timbre feature corresponding to the character to be filtered out; Based on the third prompt word, the character's role information, and the multiple candidate timbre features, the character's timbre features are selected from the multiple candidate timbre features, and the timbre features and the audio that matches the timbre features are used as the character's timbre information.

8. The method of claim 7, wherein, The character information includes character name, character traits, and character description. The character traits include at least one of the following: character gender, character age range, character personality, and character identity. The step of querying a timbre library for multiple candidate timbre traits that match the character information, according to pre-set timbre matching rules and the character information, includes: From the timbre library, query multiple timbre feature samples that match the character's age range and gender; The matching weights corresponding to the character personality, the character identity, and the character introduction are obtained from the timbre matching rules, and the one with the highest matching weight among the character personality, the character identity, and the character introduction is taken as the information to be matched; From the plurality of timbre feature samples, query the timbre feature samples that match the information to be matched, and use them as the candidate timbre features; The step of filtering out the character's vocal features from the multiple candidate vocal features based on the third prompt word, the character's role information, and the multiple candidate vocal features includes: Based on the third prompt word, a pre-trained third language model is invoked, and the character's timbre features are selected from the multiple candidate timbre features according to the character information.

9. The method according to any one of claims 1 to 8, characterized in that, The audio segments include at least one of narration audio segments and dialogue audio segments. The process of speech synthesis based on each text segment and the timbre information of the character within each text segment to obtain an audio segment corresponding to each text segment includes: For each of the text fragments, the following processing is performed: The text fragment is subjected to narration text detection processing to obtain narration detection results, and the text fragment is subjected to dialogue text detection processing to obtain dialogue detection results; In response to the narration text result indicating that there is narration text in the text segment, narration speech generation processing is performed based on the narration text and the pre-set narration audio timbre to obtain the narration audio segment; In response to the dialogue detection result indicating the presence of dialogue text in the text segment, dialogue speech generation processing is performed based on the dialogue text and the timbre information to obtain the dialogue audio segment.

10. The method of claim 9, wherein, The timbre information includes timbre features and audio matching the timbre features. The dialogue speech generation process based on the dialogue text and the timbre information to obtain the dialogue audio segment includes: The first text features are extracted from the dialogue text to obtain the first text features; The first timbre feature is extracted from the audio that matches the timbre feature to obtain the first timbre feature; Based on the first text feature and the first timbre feature, feature fusion is performed to obtain the first fused feature; Based on the first fused feature, a first audio feature is predicted to obtain the first audio feature; The first audio feature is decoded to obtain the dialogue audio segment.

11. The method of claim 9, wherein, The process of generating narration speech based on the narration text and pre-set narration audio timbre to obtain the narration audio segment includes: The second text features are extracted from the narration text to obtain the second text features; The second timbre feature is extracted from the narration audio to obtain the second timbre feature; Based on the second text feature and the second timbre feature, feature fusion is performed to obtain the second fused feature; Based on the second fusion feature, the second audio feature is predicted to obtain the second audio feature; The second audio feature is decoded to obtain the narration audio segment.

12. The method according to any one of claims 1 to 8, characterized in that, After obtaining the audio segment corresponding to each text segment, the method further includes: Perform the following processing on the audio segment corresponding to each of the text segments: The audio segment is subjected to speech recognition processing to obtain speech-recognized text; Based on the text segment and the speech-recognized text, determine the quality parameters of the audio segment; In response to the quality parameter being greater than or equal to a preset quality parameter threshold, the process proceeds to combining the audio segments corresponding to each text segment to obtain the audio corresponding to the target text. In response to the quality parameter being less than a preset quality parameter threshold, the process proceeds to speech synthesis based on each text segment, the character in each text segment, and the timbre information of the character.

13. The method of claim 12, wherein, The step of determining the quality parameters of the audio segment based on the text segment and the speech-recognized text includes: Using the text segment as a reference text, the number of missing characters in the speech recognition text relative to the reference text is obtained, as well as the number of replaced characters in the speech recognition text relative to the reference text is obtained; Based on the text length of the speech-recognized text, obtain the pre-set audio duration range; The duration parameter of the audio segment is determined based on the duration of the audio segment and the audio duration range; The quality parameters of the audio segment are obtained by performing a weighted summation of the number of missing characters, the number of replaced characters, and the duration parameter.

14. The method according to any one of claims 1 to 8, characterized in that, The text-to-audio method is implemented by calling the processing node in the workflow framework. The segmentation is implemented by calling the segmentation node in the workflow framework. The role information list is obtained by calling the role information extraction node in the workflow calculation framework. The role is implemented by the role recognition node in the workflow framework. The timbre information is obtained by the timbre allocation node in the workflow framework. The audio segment is synthesized by the dubbing node in the workflow calculation framework. Before segmenting the target text into multiple text fragments, the method further includes: Display the workflow settings interface; In response to a first setting operation in the workflow settings interface, multiple processing nodes are set in the workflow settings interface; In response to a second setting operation for each of the processing nodes, working logic is applied to each of the processing nodes, wherein the working logic is used to perform specific processing on the target text; In response to a connection operation for multiple processing nodes, the multiple processing nodes are connected to form the workflow framework, wherein, for any two processing nodes that are connected, the output of the preceding processing node is the input of the following processing node.

15. The method of claim 1, wherein, Before segmenting the target text into multiple text fragments, the method further includes: The target text is cleaned to obtain the cleaned target text. Based on the newline characters in the cleaned target text, the cleaned target text is divided into multiple lines of text; The target text is stored as structured target text according to each line of text and the line number of each line of text, and then the process of segmenting the target text into multiple text fragments is carried out based on the structured target text.

16. The method of claim 15, wherein, The process of segmenting the target text into multiple text fragments includes: The structured target text is segmented according to a pre-set text segmentation rule to obtain multiple text segments. Each text segment corresponds to a text segment information, which includes a text segment number and the line number of the text included in the text segment.

17. A text-to-audio converter, characterized in that, The device includes: The text segmentation module is used to divide target text into multiple text fragments; A character information extraction module is used to extract a character information list from the target text, wherein the character information list includes the character information of each character in the target text; The character recognition module is used to identify the character in each text fragment based on the character information list; A timbre allocation module is used to determine the timbre information of each character based on the character information of each character. The speech synthesis module is used to perform speech synthesis based on each text segment and the timbre information of the character in each text segment to obtain an audio segment corresponding to each text segment; The speech synthesis module is further configured to combine the audio segments corresponding to each text segment to obtain the audio corresponding to the target text.

18. An electronic device, comprising: The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the text-to-audio method according to any one of claims 1 to 16.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program comprise the steps of claim 18. When the computer-executable instructions or computer program are executed by a processor, the text-to-audio method according to any one of claims 1 to 16 is implemented.

20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the text-to-audio method according to any one of claims 1 to 16 is implemented.