Audio playing method and device, computer equipment and computer readable storage medium
By analyzing the character map and dialogue labeling information in the audio text, different character tones are automatically synthesized, which solves the problem of high human resources consumption in the existing technology and realizes efficient audio synthesis.
Patent Information
- Application Number
- CN202410189324.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-20
- Publication Date
- 2025-08-26
AI Technical Summary
Different tones in existing audio works require a lot of human resources and are less efficient.
By obtaining the target text, analyzing the character map and dialogue annotation information, determining the target tone parameters corresponding to different characters, and synthesizing the target audio based on these parameters to achieve automated tone synthesis.
Efficiently synthesize audio with multiple character tones without manual dubbing, improving synthesis efficiency.
Smart Images

Figure CN120544534A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and specifically to an audio playback method, apparatus, computer equipment, and computer-readable storage medium. Background Art
[0002] With the advancement of technology and the changing needs of people, new forms of audio works such as audiobooks and radio dramas are becoming more and more popular. Such audio works have plots and can be played directly. Some of them will interpret different characters through different timbres, and the voices are rich, varied and interesting. However, in related audio works, different timbres are used to interpret different characters, and most of them are dubbed by real people. This process requires more human resources and is less efficient. Summary of the Invention
[0003] Embodiments of the present application provide an audio playback method, apparatus, computer device, and computer-readable storage medium that can efficiently synthesize target audio having at least two character timbres.
[0004] The present invention provides an audio playback method, including:
[0005] Obtaining a target text, where the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles;
[0006] Performing content analysis on the target text to obtain a role map and at least two dialogue annotation information, wherein the role map includes role identity information of at least two characters, and the dialogue annotation information includes the identifiers of the dialogue characters to which the dialogue text belongs;
[0007] Searching for target character identity information corresponding to the dialogue character identifier from the character map, and determining target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters for different dialogue characters are different;
[0008] generating a dialogue audio according to the target timbre parameter and the dialogue text, and synthesizing a target audio of the target text based on at least two dialogue audios;
[0009] In response to a play operation on the target text, the target audio is played.
[0010] Accordingly, the present application provides an audio playback device, comprising:
[0011] An acquisition module is used to acquire a target text, where the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles;
[0012] An analysis module is configured to perform content analysis on the target text to obtain a role map and at least two dialogue annotation information, wherein the role map includes role identity information of at least two characters, and the dialogue annotation information includes the identifiers of the dialogue characters to which the dialogue text belongs;
[0013] A timbre module is used to search for target character identity information corresponding to the dialogue character identifier in the character map, and determine target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters of different dialogue characters are different;
[0014] an audio module, configured to generate a dialogue audio according to target timbre parameters and a dialogue text, and synthesize a target audio of the target text based on at least two dialogue audios;
[0015] The playing module is used to play the target audio in response to the playing operation on the target text.
[0016] In some embodiments of the present application, the target text further includes non-dialogue text, the dialogue annotation information further includes position information of the dialogue text in the target text, and the audio playback device further includes:
[0017] A non-dialogue module, configured to obtain preset timbre parameters for the non-dialogue text and generate non-dialogue audio based on the preset timbre parameters and the non-dialogue text;
[0018] At this time, the audio module is specifically used for:
[0019] According to the position information of each dialogue text, each dialogue audio is inserted into the non-dialogue audio to obtain the target audio of the target text.
[0020] In some embodiments of the present application, the dialogue annotation information also includes the emotional information of the dialogue characters in the dialogue text, and the timbre module includes an identity submodule and a timbre submodule, wherein:
[0021] The identity submodule is used to find the target role identity information corresponding to the dialogue role identifier from the role graph;
[0022] A timbre submodule is used to determine target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters of different dialogue characters are different;
[0023] In some embodiments of the present application, the timbre submodule includes an initial unit and a target unit, wherein:
[0024] An initialization unit, used to determine initial timbre parameters corresponding to the dialogue text based on the target character identity information;
[0025] The target unit is used to adjust the initial timbre parameters according to the emotional information to obtain the target timbre parameters corresponding to the dialogue text.
[0026] In some embodiments of the present application, the character map also includes the character's personality information, and the initial unit includes a basic sub-unit and an initial sub-unit, wherein:
[0027] A basic subunit, used to determine basic timbre parameters corresponding to the dialogue text based on the target character identity information;
[0028] The initial subunit is used to adjust the basic timbre parameters according to the character information of the dialogue character to obtain the initial timbre parameters corresponding to the dialogue text.
[0029] In some embodiments of the present application, the target role identity information includes target age information and target gender information, and the basic subunit is specifically used to:
[0030] Obtaining a preset timbre data table, the timbre data table including a plurality of preset timbre parameters, each preset timbre parameter corresponding to preset gender information and preset age information;
[0031] From a plurality of preset timbre parameters, basic timbre parameters corresponding to both target gender information and target age information are searched.
[0032] In some embodiments of the present application, the audio playback device further includes a parameter display module and a parameter adjustment module, wherein:
[0033] A parameter display module, configured to display target timbre parameters of at least two characters of the target text on the client;
[0034] a parameter adjustment module for obtaining updated timbre parameters corresponding to the selected character in response to an adjustment operation on the target timbre parameters of the selected character, and replacing the target timbre parameters of the selected character with the updated timbre parameters;
[0035] At this time, the audio module is specifically used to generate dialogue audio according to the updated timbre parameters and the dialogue text of the selected character.
[0036] In some embodiments of the present application, the audio playback device further includes a storage module, wherein:
[0037] A saving module, used for saving the target audio in a target storage system;
[0038] At this time, the playback module is specifically used for:
[0039] In response to a play operation on the target text, the target audio is acquired from the target storage system, and the player is controlled to play the target audio.
[0040] In some embodiments of the present application, the audio playback device further includes:
[0041] The playback progress module is used to obtain the playback progress information of the target audio in the player and query the paragraph information of the target text corresponding to the playback progress information;
[0042] The text progress module is used to control the reader to display the text corresponding to the paragraph information in the target text, so that the audio playback progress of the player matches the text display progress of the reader.
[0043] In some embodiments of the present application, the audio playback device further includes:
[0044] A subtitle generation module is used to generate a subtitle file corresponding to the target audio. The subtitle file includes the time information of each speech in the target audio, the corresponding text of the speech, and the paragraph information of the paragraph to which the text belongs in the target text;
[0045] The subtitle display module is used to display real-time subtitles on the client through the subtitle file while the player is playing the target audio;
[0046] At this time, the playback progress module is specifically used to: determine the playback voice corresponding to the playback progress information, and query the paragraph information corresponding to the playback voice from the subtitle file.
[0047] In some embodiments of the present application, the audio playback device further includes:
[0048] A text generation module is used to generate analysis prompt text based on the target text, and the analysis prompt text is used to prompt content analysis of the target text;
[0049] The text input module is used to input analysis prompt text into the large language model so as to perform content analysis on the target text through the large language model to obtain a role map and at least two dialogue annotation information.
[0050] Accordingly, an embodiment of the present application also provides a computer device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the audio playback method provided in the embodiment of the present application.
[0051] Accordingly, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to implement the steps in the audio playback method provided in the embodiment of the present application.
[0052] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which are executed by a processor to implement the steps in the audio playback method provided in the embodiment of the present application.
[0053] The embodiment of the present application can process the target text to obtain a character map and dialogue annotation information. The character map may include the role identity information of all characters in the target text, and the dialogue annotation information may include the dialogue role identifier to which the dialogue text belongs. The present application can search for the target role identity information corresponding to the dialogue role identifier from the character map, and then obtain the target timbre parameters that match the role identity information. The role identity information of each character is different, so the target timbre parameters are also different. Finally, the present application can synthesize the dialogue audio of each character based on different target timbre parameters, and finally obtain the target audio with at least two character timbres. This process does not require manual dubbing and has high synthesis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0055] Figure 1 Schematic diagram of a scenario of the audio playback method provided in an embodiment of the present application;
[0056] Figure 2 Schematic diagram of the flow of the audio playback method provided in the embodiment of the present application;
[0057] Figure 3 This is another flowchart of the audio playback method provided in an embodiment of the present application;
[0058] Figure 4 Schematic diagram of the architecture of the audio playback method provided in an embodiment of the present application;
[0059] Figure 5 is a structural diagram of an audio playback device provided in an embodiment of the present application;
[0060] Figure 6 It is a structural diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0062] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the numbers used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0063] It is understandable that in several embodiments of the present application, related data such as user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0064] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0065] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0066] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. Pre-trained models such as WavLM and UniSpeech, which leverage the Transformer architecture, possess strong generalization and versatility, enabling them to effectively handle a wide range of speech processing tasks.
[0067] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, the language people use daily, and is closely related to linguistics. It also involves computer science and mathematics. It is a discipline that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, the language people use daily, and is closely connected to linguistics. Pre-trained models, a key technology for model training in artificial intelligence, are derived from large language models (LLMs) in the NLP field. After fine-tuning, LLMs can be widely applied to downstream tasks. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, knowledge graphs, and other technologies.
[0068] The embodiments of the present application relate to fields such as natural language processing and speech technology of artificial intelligence. For example, content analysis of a target text is performed through relevant technologies of natural language processing. For another example, speech technology is used to generate dialogue audio according to the target timbre parameters and the dialogue text, and so on.
[0069] The embodiments of the present application provide an audio playback method, apparatus, computer device, and computer-readable storage medium. The audio playback apparatus can be integrated into an audio playback system, and the audio playback system can be integrated into at least one computer device, which can include at least one of a terminal and a server.
[0070] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0071] In an embodiment of the present application, the audio playback system can be integrated into at least one computer device. For example, the audio playback system can be integrated into a terminal. The terminal can obtain a target text, which includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles. The terminal can perform content analysis on the target text to obtain a role map and at least two dialogue annotation information. The role map includes role identity information of at least two roles, and the dialogue annotation information includes the dialogue role identifier to which the dialogue text belongs; the target role identity information corresponding to the dialogue role identifier is searched from the role map, and the target timbre parameters corresponding to the dialogue text are determined based on the target role identity information, wherein the target timbre parameters of different dialogue roles are different. The terminal can also generate dialogue audio based on the target timbre parameters and the dialogue text, and synthesize the target audio of the target text based on the at least two dialogue audios. The terminal can also play the target audio in response to a playback operation for the target text.
[0072] For example, see Figure 1 The audio playback system can be integrated into at least one server and at least one terminal. The server can obtain a target text, which includes at least two dialogue texts, each belonging to a different dialogue role. The server can perform content analysis on the target text to obtain a role map and at least two dialogue annotation information. The role map includes role identity information of at least two roles, and the dialogue annotation information includes the dialogue role identifier to which the dialogue text belongs. The server can search the role map for target role identity information corresponding to the dialogue role identifier and determine target timbre parameters corresponding to the dialogue text based on the target role identity information, wherein the target timbre parameters for different dialogue roles are different. The server can generate dialogue audio based on the target timbre parameters and the dialogue text, and synthesize the target audio of the target text based on the at least two dialogue audios. The terminal can respond to a play operation for the target text by sending a target audio acquisition request to the server. The server can send the target audio to the terminal based on the target audio acquisition request, so that the terminal plays the target audio.
[0073] In some embodiments of the present application, the terminal responds to a timbre adjustment operation on a target text by sending a request for obtaining timbre information to a server. The server may send the target timbre parameters of each character in the target text and the preset timbre parameters of the narration (corresponding to non-dialogue text) to the terminal based on the timbre information acquisition request, and present the target timbre parameters of each character at the terminal. The terminal may respond to a timbre adjustment operation on a selected character or narration, obtain updated timbre parameters of the selected character or narration, and replace the target timbre parameters of the selected character with the updated timbre parameters, or replace the preset timbre parameters of the narration with the updated timbre parameters, and then synthesize the target text into target audio according to the new target timbre parameters, and play the target audio in response to a play operation on the target text.
[0074] In some embodiments of the present application, the server may also send a character map and dialogue annotation information of the target text to the terminal based on a terminal request. The terminal may search for the target role identity information corresponding to the dialogue role identifier from the character map, and determine the target timbre parameters corresponding to the dialogue text based on the target role identity information, wherein the target timbre parameters of different dialogue roles are different. The terminal may then generate dialogue audio based on the target timbre parameters and the dialogue text, and synthesize the target audio of the target text based on at least two dialogue audios. Finally, the terminal may play the target audio in response to a play operation for the target text.
[0075] Figure 1 This is an example of an application scenario of the audio playback system of the present application, which is mainly used to introduce but not limit the audio playback system of the present application. In the process of actually applying the technical solution described in the embodiment of the present application, the computer devices included in the audio playback system and the steps performed by each computer device can be flexibly adjusted, and are not limited to Figure 1 The content described in .
[0076] The audio playback method of the present application will be further described below in conjunction with embodiments. It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments.
[0077] Figure 2 A flow chart of the audio playback method of the present application is shown as follows: Figure 2 , the audio playback method may include:
[0078] 110. Obtain a target text, where the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles.
[0079] Among them, the target text may include literary works of various genres, for example, the target text may be a novel text, a non-fiction work text, an essay text, a script text, a fable text, a news release text, etc. The target text includes at least one character, and the character may include a dialogue character or a non-dialogue character. The dialogue character may include a character that dialogues with itself or with other characters. The target text includes dialogue text that depicts the dialogue character's dialogue with itself (such as inner monologue or soliloquy) or other characters. The target text may also include non-dialogue text. The non-dialogue text can be used to advance the plot, depict scenes, depict character actions, explain the background, etc.
[0080] Specifically, the target text can be pre-stored locally on a computer device or other computer device, and can be obtained through passive reception or active request. The specific implementation can be flexibly carried out according to actual conditions, and this application does not impose any restrictions on this.
[0081] 120. Perform content analysis on the target text to obtain a role map and at least two dialogue annotation information. The role map includes role identity information of at least two characters, and the dialogue annotation information includes the dialogue role identifiers to which the dialogue text belongs.
[0082] Among them, the role map may include a map that records the relevant information of each role in the target text and the relationship information between the roles. The roles can be distinguished by role identification. The relevant information of the role may include role name information, role identity information, role personality information, etc. In the embodiment of the present application, the role may include various types of roles that appear in the text work. The role name information may include information used to identify and distinguish a certain role in the target text, such as name, nickname, etc. The role identity information may include information that reflects the role identity, such as gender information, age (age) information, occupation information, etc. The role personality information may include information that reflects the characterization of the role. The relationship information between the roles may include information that reflects the relationship between the roles.
[0083] Among them, the dialogue annotation information can record relevant information involved in the dialogue text. For example, the dialogue annotation information can include the dialogue role identifier of the dialogue role to which the dialogue text belongs, the emotional information, tone information, position information, etc. of the dialogue role in the dialogue text. The emotional information can reflect the emotions of the dialogue role in the context of the dialogue text, the tone information can reflect the tone of the dialogue role when speaking the dialogue text, and the position information can represent the position of the dialogue text in the target text.
[0084] The various information obtained from the content analysis of the target text here can be the original text extracted from the target text, or the information obtained after the original text is converted. For example, the character name information can be the original text, such as "Lao Wang", "Xiao Bao", "Zhang San", etc. For another example, the age information can be the original text, such as "thirty years old", or the information obtained after the original text is converted, such as converting "jiguan" to "20", converting "jigui" to "15", etc. For another example, different codes can be set in advance to identify different occupations. The original text "master", "teacher", "teacher", "teacher", "taifu", "teacher", etc. can all be converted into digital codes 023. Similarly, different identifiers can also be set in advance for relationship information. For example, the original text "father and son" can be converted into the character identifier "fs1", and the original text "a pair of ill-fated mandarin ducks", "lovers", "biren", etc. can be converted into "cl".
[0085] For example, emotional information may include original emotional text such as "happy", "angry", "nervous", "disgusted", "contemptuous", "excited", "excited", etc., and may also include converting the original emotional text into a limited number of emotional identifiers that represent basic emotions. For example, 6 pre-set emotional identifiers can correspond to 6 basic emotions. The original emotional text can be converted into specific emotional identifiers according to pre-set classification principles.
[0086] Similarly, the tone information may also include original tone text such as "hesitant", "contemptuous", "shouting", "slow", "harsh", "sharp", etc., and the original tone text may also be converted into a limited number of tone identifiers representing basic tones. For example, the tone identifiers pre-set by 8 people can correspond to 8 basic tones. The original tone text may be converted into specific tone identifiers according to pre-set classification principles.
[0087] In the embodiments of the present application, the position information can be represented by at least one of the following: a line identifier (e.g., a line number), a paragraph identifier (e.g., a paragraph number), a page identifier (e.g., a page number), a chapter identifier (e.g., a chapter number, a chapter name), etc., indicating the line where the dialogue text is located in the target text. For example, the position information of dialogue text 1 may include: chapter number A, paragraph number 2.
[0088] Specifically, there are many ways to perform content analysis on the target text. For example, you can build and train a content analysis model based on role recognition technology, you can obtain a content analysis model built based on role recognition technology and train it, or you can call a content analysis model that can be directly used based on role recognition technology, and then use the content analysis model to perform content analysis on the novel text to obtain a role map and dialogue annotation information.
[0089] In some embodiments of the present application, a large language model can be used to perform content analysis on a target text. Specifically, analysis prompt text can be first generated based on the target text, and then this analysis prompt text can be input into the large language model. The large language model can then perform content analysis on the target text based on the analysis prompt text to obtain a character map and at least two dialogue annotation information. The large language model herein can include a large language model trained based on a number of sample texts.
[0090] For example, the target text can be combined with the preset prompt text module to obtain analysis prompt text 1: Please construct a role map for all characters in the [target text]. The role map includes xxx1 (a text description of the content that the role map must include). Please identify the dialogue annotation information of each dialogue in the [target text]. The dialogue annotation information includes xxx2 (a text description of the content that the dialogue annotation information must include). Input the analysis prompt text 1 into the large language model to obtain the role map 1 and N (positive integer) dialogue annotation information output by the large language model.
[0091] In some embodiments of the present application, roles can be extracted from the target text, and a role identifier can be set for each identified role. Then, the relevant information (age information, gender information, personality information, name information, etc.) recorded in the target text for each role, as well as the relationship information between the roles, can be specifically identified. Then, the role identifier can be used as an entity, and the relevant information of the role identifier can be used as different attribute values of the entity. The entities with relationship information can be connected, and the relationship information can be used as the attribute value of the edge between the entities, and finally a role map of the target text can be obtained.
[0092] In some embodiments of the present application, the dialogue text can be determined by identifying key characters and the dialogue annotation information of the dialogue text. For example, the dialogue text can be determined by identifying key characters such as "say:", "think:", "say:", and "shout:". The name information, emotional information, tone information, etc. of the dialogue role to which the dialogue text belongs can be extracted from several characters adjacent to the dialogue text. The target text can be pre-annotated with line numbers, paragraph numbers, page numbers, etc., and its position information can be determined after the dialogue text is determined.
[0093] 130. Search the target role identity information corresponding to the dialogue role identifier in the role map, and determine the target timbre parameters corresponding to the dialogue text based on the target role identity information, wherein the target timbre parameters of different dialogue roles are different.
[0094] Specifically, the role map includes the role identity information of at least two characters. For each character pair, after determining the dialog role identifier to which the dialogue text belongs, the target role identity information corresponding to the dialog role identifier can be searched from the role map. For example, if the role map includes role identity information 1 corresponding to role identifier psx and role identity information 2 corresponding to role identifier cd, the target role identity information corresponding to the dialog role identifier cd can be searched from the role map: role identity information 2.
[0095] Among them, timbre parameters may include parameters that reflect the sound characteristics of the timbre. The specific meanings of timbre parameters in different application scenarios may be different. For example, a certain timbre may correspond to a set of timbre parameters, including pitch, frequency, fundamental wave and harmonics, etc., and each timbre parameter has a different parameter value; for another example, a certain timbre may correspond to a parameter value of a timbre parameter, and different values of the timbre parameter may represent different timbres, such as the timbre of a girl, the timbre of a child, the timbre of an elderly person, etc.
[0096] There are many ways to determine the target timbre parameters corresponding to the dialogue text (i.e., the dialogue character) based on the target character identity information. For example, the target character identity information can be digitally quantized based on a preset quantization algorithm to obtain a corresponding target identity value. Then, a preset timbre set can be obtained, which includes multiple preset identity values and preset timbre parameters corresponding to each preset identity value. Then, the target timbre parameters corresponding to the target identity value can be searched from the preset timbre set.
[0097] In some embodiments of the present application, the basic timbre parameters corresponding to the dialogue character can be first determined based on the target character identity information, and then the basic timbre parameters can be adjusted according to the personality information of the dialogue character. The initial timbre parameters can then be further adjusted according to the emotional information in the dialogue text to obtain the target timbre parameters corresponding to the dialogue text.
[0098] The process of determining basic timbre parameters may include: obtaining a preset timbre data table, the timbre data table including multiple preset timbre parameters, each corresponding to preset gender information and preset age information; and searching, from the multiple preset timbre parameters, for basic timbre parameters corresponding to both the target gender information and the target age information. In some credit scenarios, the preset timbre parameters may correspond to preset gender information, preset age information, and preset occupation information. Accordingly, basic timbre information corresponding to the target gender information, target age information, and target occupation information of the dialogue character may be searched from the timbre dataset.
[0099] The process of adjusting the basic timbre parameters through personality information and adjusting the initial timbre parameters through emotional information can be carried out based on the same logic. For example, a preset adjustment amplitude data set can be obtained, and the adjustment amplitude data set can include parameter adjustment amplitude values corresponding to different personality information / emotional information. The parameter adjustment amplitude value corresponding to the personality information / emotional information of the dialogue character can be found from the adjustment amplitude data set, and the basic timbre parameters / initial timbre parameters can be corrected based on the parameter adjustment amplitude value to obtain the initial timbre parameters / target timbre parameters.
[0100] In some embodiments of the present application, a parameter adjustment model and a parameter generation model can also be constructed and trained. Basic timbre parameters / initial timbre parameters and personality information / emotional information can be input into the parameter adjustment model, and the parameter adjustment model can output the adjusted initial timbre parameters / target timbre parameters accordingly; the target role identity information, personality information and emotional information of the dialogue role can be input into the parameter generation model, and the target role audio of the dialogue role can be directly output through the parameter generation model.
[0101] For example, the preset timbre parameter p corresponding to the age information 1 and the gender information 2 is searched from the preset timbre data table 1. The preset timbre parameter p is the basic timbre parameter of the dialogue character A. Then, the first adjustment amplitude value corresponding to the personality information of the dialogue character A and the second adjustment amplitude value corresponding to the emotional information are determined. Therefore, the basic timbre parameter is adjusted by the first adjustment amplitude value and the second adjustment amplitude value, and finally the target timbre parameter is obtained.
[0102] 140. Generate dialogue audio according to the target timbre parameter and the dialogue text, and synthesize the target audio of the target text based on at least two dialogue audios.
[0103] Among them, the conversation audio may include the audio of the conversation text synthesized by the target timbre parameters. The generation process of the conversation audio can be realized by the relevant synthesis model in the field of speech synthesis. The synthesis model can be set on various computer devices with different computing and storage resources. The more computing and storage resources the synthesis model uses, the richer the synthesized audio, the more details, and the better the user's listening experience.
[0104] The target audio can include all dialogue text and audio corresponding to non-dialogue text within the target document. Synthesis of non-dialogue audio can also be achieved using relevant synthesis models in the field of speech synthesis. The overall process is similar to that of synthesizing dialogue audio. Preset timbre parameters for non-dialogue text can be pre-set by relevant personnel.
[0105] During the implementation of this application, an online synthesis model can be set up on the server and an offline synthesis model can be set up on the terminal. The online synthesis model can provide higher quality audio, and the offline synthesis model does not need to interact with the server and can directly and quickly synthesize audio.
[0106] In some embodiments of the present application, the user can adjust the initial audio parameters / target timbre parameters corresponding to the dialogue text / dialogue role, the preset timbre parameters corresponding to the non-dialogue text, etc. Specifically, the timbre parameters corresponding to the role and the preset timbre parameters corresponding to the narration (non-dialogue text) can be transmitted to the terminal, and the terminal can display the timbre parameters of the role and the narration on the client, as well as the timbre playback controls and adjustment controls for each audio parameter. The user can trigger the timbre playback control for the selected role. If not satisfied, the timbre parameters can be modified through the adjustment control to obtain the updated timbre parameters of the selected role. This can improve the flexibility of timbre adjustment, allowing users to adjust the timbre of the role according to their personal preferences, thereby improving the user experience.
[0107] After speech synthesis is performed on all dialogue texts and non-dialogue texts, at least two dialogue audios and multiple non-dialogue audios can be obtained. The dialogue identification information also includes the position information of the dialogue text in the target text. The position information can be used to determine the texts (dialogue text or non-dialogue text) adjacent to the left and right ends of the dialogue text. Therefore, the dialogue audio can be inserted into the non-dialogue audio according to the position information to obtain the target audio.
[0108] 150. In response to a play operation on the target text, play the target audio.
[0109] The playback operation for the target text may include the user triggering the audio playback control, the client automatically generating the playback instruction through a timer, etc.
[0110] It should be noted that in actual application scenarios, the response process to the target text playback operation can be before or after the target audio synthesis. For example, in response to a trigger operation for the target text, dialogue audio can be generated based on the target timbre parameters and the dialogue text, and the target audio of the target text can be synthesized based on at least two dialogue audios before the target audio is played. The timing of the response can be flexibly adjusted according to actual circumstances and is within the scope of protection of the technical solution of this application.
[0111] In some embodiments of the present application, the generated target audio can be saved in a storage system, which can be a distributed storage system, a cloud storage system, etc. If the user triggers the audio playback control for the target text on the terminal, the terminal can quickly obtain the target audio from the storage system and play the target audio on the terminal player.
[0112] In some embodiments of the present application, the terminal can play audio through the player while automatically highlighting the text corresponding to the current audio through the reader. As the playback progress of the target audio changes, the highlighting progress of the reader also changes accordingly, so that the playback progress of the target audio is consistent with the text display progress of the reader, thereby improving the user's experience of listening and reading at the same time and reducing the complexity of the user's operation.
[0113] Specifically, the playback progress information of the target audio in the player can be obtained, wherein the playback progress information can reflect the playback progress, such as the current playback time information, the text corresponding to the currently playing audio, the position information of the text in the target text, etc. If the playback progress information is related to the text, such as the text corresponding to the currently playing audio and the position information of the text in the target text, the playback progress information can be sent to the reader, and the text corresponding to the playback progress information can be displayed in real time by the reader, so that the text display progress of the reader is consistent with the audio playback progress of the player; if the playback progress information is not related to the text, such as the current playback time information, the paragraph information or line information of the target text corresponding to the playback progress information can be queried, and then the reader can be controlled to display the text corresponding to the paragraph information or line information in the target text, so that the audio playback progress of the player and the text display progress of the reader match.
[0114] In some embodiments of the present application, the terminal can display the target text in the form of subtitles while playing audio through a player. Specifically, a subtitle file corresponding to the target audio can be generated, and the subtitle file includes the time information of each sentence in the target audio, the corresponding text of the speech, and the paragraph information of the paragraph to which the text belongs in the target text; in the process of the player playing the target audio, real-time subtitles are displayed on the client through the subtitle file, thereby enabling the text to be displayed synchronously in real time in the form of subtitles, thereby improving the convenience of users listening and reading at the same time and improving the user experience.
[0115] The time information in the subtitle file matches the time information of the target audio. When the playback progress signal represents the current playback time, the paragraph information of the playback voice corresponding to the current playback time can be found in the subtitle file.
[0116] For example, the playback progress information is playback time 1, the subtitle file includes multiple voice time periods, the text corresponding to the voice in each time period, and the paragraph information of the paragraph to which the text belongs in the target text. Determine the time period x where playback time 1 is located, and determine the paragraph information 1 of the text p corresponding to the time period x.
[0117] The embodiment of the present application can process the target text to obtain a character map and dialogue annotation information. The character map may include the role identity information of all characters in the target text, and the dialogue annotation information may include the dialogue role identifier to which the dialogue text belongs. The present application can search for the target role identity information corresponding to the dialogue role identifier from the character map, and then obtain the target timbre parameters that match the role identity information. The role identity information of each character is different, so the target timbre parameters are also different. Finally, the present application can synthesize the dialogue audio of each character based on different target timbre parameters, and finally obtain the target audio with at least two character timbres. This process does not require manual dubbing and has high synthesis efficiency.
[0118] The audio playback method of the present application will be further introduced below in conjunction with the embodiments. The audio playback method can be integrated with an audio playback system, and the audio playback system can be integrated into a computer device, such as a server and a terminal.
[0119] Specifically, Figure 3 A flowchart of the audio playback method of the present application is shown as follows: Figure 3 As shown, the audio playback method may include:
[0120] 210. The server obtains a target text, where the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles.
[0121] The audio playback method of this application can be combined with Figure 4 Further, the target text may include a novel text. The audio playback method may include a first stage and a second stage. The first stage may be implemented by a computer device in a specific local area network. The first stage may include a character analysis module, a speech synthesis module, a storage system, and other main bodies. Different main bodies may be set up in the same computer device in the specific local area network, or in multiple computer devices in the specific local area network.
[0122] In the first stage, the server may obtain the novel text, specifically, the server may export the novel text and add it to a queue to be processed.
[0123] A novel text may include multiple dialogue texts and non-dialogue texts. Different dialogue texts may belong to the same dialogue role or to different dialogue roles.
[0124] 220. The server performs content analysis on the target text to obtain a role map and at least two dialogue annotation information. The role map includes the role identity information and personality information of at least two characters. The dialogue annotation information includes the dialogue role identifier to which the dialogue text belongs and the location information of the dialogue text in the target text.
[0125] The character analysis module can be integrated with a content analysis model, thereby performing content analysis (including character analysis and content annotation) on the target text through the content analysis module to obtain a character map and text annotation information, which can include multiple dialogue annotation information. In some cases, embodiments of the present application can also build an annotation information management platform to display the text annotation information of the novel text on the page, facilitating manual review of the dialogue annotation information.
[0126] The character map also includes the character's name information, etc. The location information in the dialogue annotation information can be the chapter location number of the dialogue text.
[0127] 230. The server searches the role map for target role identity information and target personality information corresponding to the dialogue role identifier.
[0128] 240. The server determines the target timbre parameters of the dialogue character based on the target character identity information and target personality information.
[0129] 250. The server transmits target timbre parameters of at least two dialogue characters and preset timbre parameters of the narration sound to the terminal based on the timbre acquisition request transmitted by the terminal.
[0130] 260. The terminal displays the target timbre parameters of each dialogue character and the preset timbre parameters of the narration sound.
[0131] 270. In response to the timbre parameter adjustment operation for the selected character or narration sound, the terminal obtains updated timbre parameters of the selected character or narration sound, and transmits the updated timbre parameters to the server, so that the server replaces the target timbre parameters of the selected character with the updated timbre parameters, or replaces the preset timbre parameters of the narration sound with the updated timbre parameters.
[0132] In this way, users can flexibly adjust the timbre parameters of each character based on personal preferences, personal understanding of the character, etc. The resulting audiobook audio is more in line with user needs and improves the user experience.
[0133] 280. The server generates dialogue audio corresponding to the dialogue text according to the target timbre parameters of the dialogue character, generates non-dialogue audio corresponding to the non-dialogue text according to the preset timbre parameters, and synthesizes the target audio of the target text based on the non-dialogue audio, the respective position information of at least two dialogue texts, and the dialogue audio.
[0134] Therefore, this application can obtain dialogue audio that is consistent with the character's identity, personality and emotions when speaking, and then obtain audio book audio of the novel text. This audio book audio is similar to the real-person dubbing effect, but does not require a lot of manpower and is more efficient and quick.
[0135] 290. The server sends the target audio to the terminal based on the audio acquisition request transmitted by the terminal, so that the terminal responds to the audio playback operation of the target text and plays the target audio.
[0136] The speech synthesis module integrates a speech synthesis model, which allows it to synthesize audiobook audio corresponding to the novel text based on multiple dialogue annotation information and character maps. The audiobook audio can be a single audio file or a combination of multiple chapter audio files. The server can also generate subtitle files corresponding to the audiobook audio. Correspondingly, the subtitle file can be a single file or multiple chapter subtitle files. The subtitle file can include the start time of each sentence, the chapter location label, and the text.
[0137] The server can store the audio of the audio book in the distributed storage system cos. When receiving the audio acquisition request sent by the terminal, the server can obtain the audio of the audio book from the distributed storage system cos and send the audio book audio to the terminal.
[0138] The audio playback operation of the target text can be, for example, a user triggering a multi-character audiobook playback control of a novel.
[0139] While the terminal is playing the audiobook audio through the player, the novel text can be displayed in real time through the reader. The playback progress of the audiobook audio is consistent with the display progress of the novel text. For example, when the audiobook audio plays to the audio corresponding to a certain sentence, the reader will brightly display the text.
[0140] While the terminal is playing the audio book audio through the player, it can display real-time subtitles through the subtitle file. The real-time subtitles include the novel text that matches the playback progress of the audio book audio.
[0141] See also Figure 4 In the second stage of the audio playback method, the novel text (main text content), target audio (audio address of the target audio), character map and main text annotation information (multi-character information) can be obtained through the computer device (background). For the terminal audio book client, it can include offline single-tone speech synthesis (offline conventional TTS), offline multi-character speech synthesis, online multi-character speech synthesis and other speech synthesis models. The audio book client can also provide users with a page for adjusting the narration sound, the timbre parameters of each character, and a page for selecting the speech synthesis mode. The terminal can also obtain paragraph information, chapter information, line number information, etc. of the novel text from the background, and can also obtain the target audio.
[0142] When the audiobook audio is played on the player, the reader of the audiobook client can determine the chapter information, paragraph information, etc. corresponding to the audio at the current playback progress, and automatically display the corresponding text, etc. as the playback progresses.
[0143] The embodiment of the present application can process the target text to obtain a character map and dialogue annotation information. The character map may include the role identity information of all characters in the target text, and the dialogue annotation information may include the dialogue role identifier to which the dialogue text belongs. The present application can search for the target role identity information corresponding to the dialogue role identifier from the character map, and then obtain the target timbre parameters that match the role identity information. The role identity information of each character is different, so the target timbre parameters are also different. Finally, the present application can synthesize the dialogue audio of each character based on different target timbre parameters, and finally obtain the target audio with at least two character timbres. This process does not require manual dubbing and has high synthesis efficiency.
[0144] In order to better implement the above method, the embodiment of the present application also provides an audio playback device, such as Figure 5 As shown, the audio playback device may include an acquisition module 310, an analysis module 320, a timbre module 330, an audio module 340, and a playback module 350, wherein:
[0145] An acquisition module 310 is configured to acquire a target text, wherein the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles;
[0146] An analysis module 320 is configured to perform content analysis on the target text to obtain a character map and at least two dialogue annotation information, wherein the character map includes role identity information of at least two characters, and the dialogue annotation information includes the identifiers of the dialogue characters to which the dialogue text belongs;
[0147] The timbre module 330 is used to search the target character identity information corresponding to the dialogue character identifier from the character map, and determine the target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters are different for different dialogue characters;
[0148] An audio module 340 is configured to generate a dialogue audio according to a target timbre parameter and a dialogue text, and synthesize a target audio of the target text based on at least two dialogue audios;
[0149] The playing module 350 is configured to play the target audio in response to a play operation on the target text.
[0150] In some embodiments of the present application, the target text further includes non-dialogue text, the dialogue annotation information further includes position information of the dialogue text in the target text, and the audio playback device further includes:
[0151] A non-dialogue module, configured to obtain preset timbre parameters for the non-dialogue text and generate non-dialogue audio based on the preset timbre parameters and the non-dialogue text;
[0152] At this time, the audio module is specifically used for:
[0153] According to the position information of each dialogue text, each dialogue audio is inserted into the non-dialogue audio to obtain the target audio of the target text.
[0154] In some embodiments of the present application, the dialogue annotation information also includes the emotional information of the dialogue characters in the dialogue text, and the timbre module includes an identity submodule and a timbre submodule, wherein:
[0155] The identity submodule is used to find the target role identity information corresponding to the dialogue role identifier from the role graph;
[0156] A timbre submodule is used to determine target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters of different dialogue characters are different;
[0157] In some embodiments of the present application, the timbre submodule includes an initial unit and a target unit, wherein:
[0158] An initialization unit, used to determine initial timbre parameters corresponding to the dialogue text based on the target character identity information;
[0159] The target unit is used to adjust the initial timbre parameters according to the emotional information to obtain the target timbre parameters corresponding to the dialogue text.
[0160] In some embodiments of the present application, the character map also includes the character's personality information, and the initial unit includes a basic sub-unit and an initial sub-unit, wherein:
[0161] A basic subunit, used to determine basic timbre parameters corresponding to the dialogue text based on the target character identity information;
[0162] The initial subunit is used to adjust the basic timbre parameters according to the character information of the dialogue character to obtain the initial timbre parameters corresponding to the dialogue text.
[0163] In some embodiments of the present application, the target role identity information includes target age information and target gender information, and the basic subunit is specifically used to:
[0164] Obtaining a preset timbre data table, the timbre data table including a plurality of preset timbre parameters, each preset timbre parameter corresponding to preset gender information and preset age information;
[0165] From a plurality of preset timbre parameters, basic timbre parameters corresponding to both target gender information and target age information are searched.
[0166] In some embodiments of the present application, the audio playback device further includes a parameter display module and a parameter adjustment module, wherein:
[0167] A parameter display module, configured to display target timbre parameters of at least two characters of the target text on the client;
[0168] a parameter adjustment module for obtaining updated timbre parameters corresponding to the selected character in response to an adjustment operation on the target timbre parameters of the selected character, and replacing the target timbre parameters of the selected character with the updated timbre parameters;
[0169] At this time, the audio module is specifically used to generate dialogue audio according to the updated timbre parameters and the dialogue text of the selected character.
[0170] In some embodiments of the present application, the audio playback device further includes a storage module, wherein:
[0171] A saving module, used for saving the target audio in a target storage system;
[0172] At this time, the playback module is specifically used for:
[0173] In response to a play operation on the target text, the target audio is acquired from the target storage system, and the player is controlled to play the target audio.
[0174] In some embodiments of the present application, the audio playback device further includes:
[0175] The playback progress module is used to obtain the playback progress information of the target audio in the player and query the paragraph information of the target text corresponding to the playback progress information;
[0176] The text progress module is used to control the reader to display the text corresponding to the paragraph information in the target text, so that the audio playback progress of the player matches the text display progress of the reader.
[0177] In some embodiments of the present application, the audio playback device further includes:
[0178] A subtitle generation module is used to generate a subtitle file corresponding to the target audio. The subtitle file includes the time information of each speech in the target audio, the corresponding text of the speech, and the paragraph information of the paragraph to which the text belongs in the target text;
[0179] The subtitle display module is used to display real-time subtitles on the client through the subtitle file while the player is playing the target audio;
[0180] At this time, the playback progress module is specifically used to: determine the playback voice corresponding to the playback progress information, and query the paragraph information corresponding to the playback voice from the subtitle file.
[0181] In some embodiments of the present application, the audio playback device further includes:
[0182] A text generation module is used to generate analysis prompt text based on the target text, and the analysis prompt text is used to prompt content analysis of the target text;
[0183] The text input module is used to input analysis prompt text into the large language model so as to perform content analysis on the target text through the large language model to obtain a role map and at least two dialogue annotation information.
[0184] The embodiment of the present application can process the target text to obtain a character map and dialogue annotation information. The character map may include the role identity information of all characters in the target text, and the dialogue annotation information may include the dialogue role identifier to which the dialogue text belongs. The present application can search for the target role identity information corresponding to the dialogue role identifier from the character map, and then obtain the target timbre parameters that match the role identity information. The role identity information of each character is different, so the target timbre parameters are also different. Finally, the present application can synthesize the dialogue audio of each character based on different target timbre parameters, and finally obtain the target audio with at least two character timbres. This process does not require manual dubbing and has high synthesis efficiency.
[0185] The present application also provides a computer device, such as Figure 6 , which shows a schematic diagram of the structure of a computer device involved in an embodiment of the present application. The computer device may be a terminal or a server, etc. Specifically:
[0186] The computer device may include one or more processing core processors 401, one or more computer readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will understand that Figure 6 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0187] Processor 401 is the control center of the computer device. It connects all components of the computer device using various interfaces and circuits. It executes computer programs and / or modules stored in memory 402 and accesses data stored in memory 402 to perform various computer functions and process data. Optionally, processor 401 may include one or more processing cores. Preferably, processor 401 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and computer programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.
[0188] The memory 402 can be used to store computer programs and modules. The processor 401 executes various functional applications and data processing by running the computer programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, a computer program required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0189] The computer device also includes a power supply 403 for supplying power to various components. Preferably, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0190] The computer device may further include an input unit 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0191] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to one or more computer program processes into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:
[0192] Acquire a target text, the target text including at least two dialogue texts, the at least two dialogue texts belonging to different dialogue roles; perform content analysis on the target text to obtain a role map and at least two dialogue annotation information, the role map including role identity information of at least two roles, and the dialogue annotation information including dialogue role identifiers to which the dialogue text belongs; search for target role identity information corresponding to the dialogue role identifiers in the role map, and determine target timbre parameters corresponding to the dialogue text based on the target role identity information, wherein target timbre parameters for different dialogue roles are different; generate dialogue audio according to the target timbre parameters and the dialogue text, and synthesize target audio of the target text based on the at least two dialogue audios; and play the target audio in response to a play operation on the target text.
[0193] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0194] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by a computer program, or by controlling related hardware through a computer program. The computer program may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0195] To this end, an embodiment of the present application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the audio playback methods provided in the embodiments of the present application. For example, the computer program can execute the following steps:
[0196] Acquire a target text, the target text including at least two dialogue texts, the at least two dialogue texts belonging to different dialogue roles; perform content analysis on the target text to obtain a role map and at least two dialogue annotation information, the role map including role identity information of at least two roles, and the dialogue annotation information including dialogue role identifiers to which the dialogue text belongs; search for target role identity information corresponding to the dialogue role identifiers in the role map, and determine target timbre parameters corresponding to the dialogue text based on the target role identity information, wherein target timbre parameters for different dialogue roles are different; generate dialogue audio according to the target timbre parameters and the dialogue text, and synthesize target audio of the target text based on the at least two dialogue audios; and play the target audio in response to a play operation on the target text.
[0197] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0198] Since the computer program stored in the computer-readable storage medium can execute the steps of any audio playback method provided in the embodiments of the present application, the beneficial effects that can be achieved by any audio playback method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0199] The present application also provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the aforementioned audio playback method.
[0200] The above is a detailed introduction to an audio playback method, device, computer equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An audio playback method, characterized in that: include: Acquire a target text, wherein the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles; Performing content analysis on the target text to obtain a role map and at least two dialogue annotation information, wherein the role map includes role identity information of at least two characters, and the dialogue annotation information includes an identifier of a dialogue character to which the dialogue text belongs; Searching for target character identity information corresponding to the dialogue character identifier from the character map, and determining target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters for different dialogue characters are different; generating a dialogue audio according to the target timbre parameter and the dialogue text, and synthesizing a target audio of the target text based on at least two of the dialogue audios; In response to a play operation on the target text, the target audio is played.
2. The method according to claim 1, characterized in that The target text also includes non-dialogue text, the dialogue annotation information also includes position information of the dialogue text in the target text, and the method further includes: Obtaining preset timbre parameters for the non-dialogue text, and generating non-dialogue audio according to the preset timbre parameters and the non-dialogue text; The step of synthesizing the target audio of the target text based on at least two of the conversation audios comprises: According to the position information of each of the dialogue texts, each of the dialogue audios is inserted into the non-dialogue audio to obtain the target audio of the target text.
3. The method according to claim 2, characterized in that The dialogue annotation information further includes the emotion information of the dialogue character in the dialogue text, and determining the target timbre parameter corresponding to the dialogue text based on the target character identity information includes: Determining initial timbre parameters corresponding to the dialogue text based on the target character identity information; The initial timbre parameters are adjusted according to the emotional information to obtain target timbre parameters corresponding to the dialogue text.
4. The method according to claim 3, characterized in that The character map also includes the character's personality information. The determining of the initial timbre parameters corresponding to the dialogue text based on the target character's identity information includes: Determining basic timbre parameters corresponding to the dialogue text based on the target character identity information; The basic timbre parameters are adjusted according to the character information of the dialogue character to obtain initial timbre parameters corresponding to the dialogue text.
5. The method according to claim 4, characterized in that The target character identity information includes target age information and target gender information. The determining of basic timbre parameters corresponding to the dialogue text based on the target character identity information includes: Obtaining a preset timbre data table, the timbre data table including a plurality of preset timbre parameters, each of the preset timbre parameters corresponding to preset gender information and preset age information; From the plurality of preset timbre parameters, basic timbre parameters corresponding to both the target gender information and the target age information are searched.
6. The method according to claim 1, characterized in that The method further comprises: Displaying target timbre parameters of at least two characters of the target text on the client; In response to an adjustment operation on a target timbre parameter of a selected character, obtaining updated timbre parameters corresponding to the selected character, and replacing the target timbre parameters of the selected character with the updated timbre parameters; Generating the dialogue audio according to the target timbre parameter and the dialogue text includes: Generate dialogue audio based on the updated timbre parameters and the dialogue text of the selected character.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Saving the target audio in a target storage system; Playing the target audio in response to a play operation on the target text includes: In response to a play operation on the target text, the target audio is acquired from the target storage system, and a player is controlled to play the target audio.
8. The method according to claim 7, characterized in that The method further comprises: Obtaining playback progress information of the target audio on the player, and querying paragraph information of the target text corresponding to the playback progress information; The reader is controlled to display the text corresponding to the paragraph information in the target text, so that the audio playback progress of the player matches the text display progress of the reader.
9. The method according to claim 8, characterized in that The method further comprises: Generate a subtitle file corresponding to the target audio, the subtitle file including time information of each speech in the target audio, text corresponding to the speech, and paragraph information of the paragraph to which the text belongs in the target text; During the process of the player playing the target audio, displaying real-time subtitles on the client through the subtitle file; The querying of the paragraph information of the target text corresponding to the playback progress information includes: The playback voice corresponding to the playback progress information is determined, and paragraph information corresponding to the playback voice is searched from the subtitle file.
10. The method according to claim 1, characterized in that The performing of content analysis on the target text to obtain a character map and at least two dialogue annotation information includes: generating an analysis prompt text based on the target text, wherein the analysis prompt text is used to prompt a content analysis of the target text; The analysis prompt text is input into the large language model so as to perform content analysis on the target text through the large language model to obtain a role map and at least two dialogue annotation information.
11. An audio playback method and device, characterized in that: include: An acquisition module is used to acquire a target text, wherein the target text includes at least two dialogue texts, and the at least two dialogue texts belong to different dialogue roles; An analysis module, configured to perform content analysis on the target text to obtain a role map and at least two dialogue annotation information, wherein the role map includes role identity information of at least two characters, and the dialogue annotation information includes an identifier of a dialogue character to which the dialogue text belongs; a timbre module, configured to search the character map for target character identity information corresponding to the dialogue character identifier, and determine target timbre parameters corresponding to the dialogue text based on the target character identity information, wherein the target timbre parameters are different for different dialogue characters; an audio module, configured to generate a dialogue audio according to the target timbre parameter and the dialogue text, and synthesize a target audio of the target text based on at least two of the dialogue audios; A playing module is used to play the target audio in response to a play operation on the target text.
12. A computer device, characterized in that: It comprises a memory and a processor; the memory stores an application, and the processor is used to run the application in the memory to execute the steps in the audio playback method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the audio playback method according to any one of claims 1 to 10.