Audio generation method, apparatus, and related product

CN121214910BActive Publication Date: 2026-08-21BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511398375.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-08-21
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

[0002]相关技术中,当根据用户提供的文本生成音频时,通常根据默认的配置信息生成音频,如根据男生音色搭配女生音色的方式生成音频,导致音频效果有限,降低了音频效果的多样性

Benefits of technology

[0007] Fourthly, embodiments of this disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214910B_ABST
    Figure CN121214910B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an audio generation method and device, and related products, wherein the method comprises: in response to an instruction of generating an audio according to a text, determining a first text; the first text is used to generate a first audio; displaying at least one audio configuration information corresponding to the first audio; the audio configuration information is used to represent a number of virtual roles in the first audio and expression style information of the virtual roles; the audio configuration information is generated according to text content of the first text; in response to a selection instruction for a first configuration information in each of the audio configuration information, outputting the first audio; the first audio is generated according to the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an audio generation method, apparatus and related products. Background Technology

[0002] In related technologies, when generating audio based on user-provided text, the audio is usually generated according to default configuration information, such as generating audio by matching male voices with female voices, which results in limited audio effects and reduces the diversity of audio effects. Summary of the Invention

[0003] This disclosure provides an audio generation method, apparatus, and related products that can improve the diversity of generated audio effects in scenarios where audio is generated from text.

[0004] In a first aspect, embodiments of this disclosure provide an audio generation method, including: In response to an instruction to generate audio from text, a first text is determined; the first text is used to generate the first audio. Display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; In response to a selection instruction for the first configuration information in each of the audio configuration information, the first audio is output; the first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

[0005] Secondly, embodiments of this disclosure provide an audio generation apparatus, comprising: A text determination unit is configured to determine a first text in response to an instruction to generate audio based on the text; the first text is used to generate the first audio. An information display unit is used to display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; An audio output unit is configured to output the first audio in response to a selection instruction for the first configuration information in each of the audio configuration information; the first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

[0006] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in the first aspect above.

[0007] Fourthly, embodiments of this disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in the first aspect.

[0008] Fifthly, embodiments of this disclosure provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect above.

[0009] In one or more embodiments of this disclosure, firstly, in response to an instruction to generate audio from text, a first text is determined. The first text is used to generate a first audio. Then, at least one audio configuration information corresponding to the first audio is displayed. The audio configuration information is used to represent the number of virtual characters and the expression style information of the virtual characters in the first audio. The audio configuration information is generated based on the text content of the first text. In response to an instruction to select first configuration information from each audio configuration information, the first audio is output. The first audio is generated based on the text content of the first text and the first configuration information. The text content of the first text is expressed in the first audio through the first configuration information. Therefore, through this embodiment, at least one audio configuration information corresponding to the first audio can be determined based on the text content of the first text, and the first audio can be generated based on the first configuration information selected from each audio configuration information. Thus, the first audio can express the text content of the first text according to the number of virtual characters and the expression style information of the virtual characters represented by the first configuration information. This alleviates the problem of limited audio effects when generating audio based on default configuration information by generating one or more selectable audio configuration information based on the text content, and improves the diversity of generated audio effects. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in one or more embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic flowchart of an audio generation method provided in an embodiment of this disclosure; Figure 2aA schematic diagram illustrating the determination of a first text according to an embodiment of this disclosure; Figure 2b A schematic diagram illustrating the determination of a first text as provided in another embodiment of this disclosure; Figure 3a This is a schematic diagram illustrating the display of audio configuration information according to an embodiment of the present disclosure; Figure 3b This is a schematic diagram of the output of a first audio signal provided in an embodiment of the present disclosure; Figure 4 This is a schematic diagram illustrating the display of audio configuration information according to yet another embodiment of the present disclosure; Figure 5a This is a schematic diagram of the output of a first audio signal provided in another embodiment of the present disclosure; Figure 5b This is a schematic diagram of the output of a first audio signal provided in yet another embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an audio generation apparatus provided in an embodiment of the present disclosure; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0011] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this disclosure, the technical solutions in one or more embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of the embodiments. Based on one or more embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this disclosure.

[0012] It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, relevant parties should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and authorization from the relevant parties should be obtained.

[0013] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0014] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0015] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0016] This disclosure provides an audio generation method, apparatus, and related products, which can improve the diversity of generated audio effects in scenarios where audio is generated from text. The audio generation method can be applied to and executed by a terminal device, including but not limited to laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smartphones, smart speakers, smartwatches, smart TVs, in-vehicle terminals, and other types of user terminals.

[0017] Figure 1 This is a flowchart illustrating an audio generation method provided in an embodiment of the present disclosure, as shown below. Figure 1 As shown, the process includes: Step S102: In response to the instruction to generate audio from text, determine the first text; the first text is used to generate the first audio. Step S104: Display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to represent the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; Step S106: In response to the selection instruction for the first configuration information in each audio configuration information, output the first audio; the first audio is generated based on the text content of the first text and the first configuration information; the text content of the first text is expressed in the first audio through the first configuration information.

[0018] In this embodiment, firstly, in response to an instruction to generate audio from text, a first text is determined. This first text is used to generate the first audio. Then, at least one audio configuration information corresponding to the first audio is displayed. This audio configuration information represents the number of virtual characters and their expressive style information in the first audio. The audio configuration information is generated based on the text content of the first text. In response to an instruction to select the first configuration information from each audio configuration information, the first audio is output. The first audio is generated based on the text content of the first text and the first configuration information. The first audio expresses the text content of the first text through the first configuration information. Therefore, through this embodiment, at least one audio configuration information corresponding to the first audio can be determined based on the text content of the first text, and the first audio can be generated based on the first configuration information selected from each audio configuration information. Thus, the first audio can express the text content of the first text according to the number of virtual characters and their expressive style information represented by the first configuration information. By generating one or more selectable audio configuration information based on the text content, the problem of limited audio effects when generating audio based on default configuration information is alleviated, and the diversity of generated audio effects is improved.

[0019] In some embodiments, a smart assistant runs on the terminal device. The terminal device can run the smart assistant by opening the corresponding webpage in a browser, or by installing and running the application containing the smart assistant. The smart assistant can be various products generated based on AI (Artificial Intelligence) models, such as AI dialogue assistants or intelligent agents. AI dialogue assistants or intelligent agents can help users perform various tasks through dialogue, such as one or more of information retrieval, image generation, text generation, and audio generation. When performing generative tasks such as image generation, text generation, and audio generation, the smart assistant can call the generative model in the AI ​​model to generate the data required by the user. The audio generation methods in the various embodiments of this disclosure can be executed through the smart assistant on the terminal device. For ease of explanation, the following figures illustrate the smart assistant as an AI dialogue assistant.

[0020] In step S102 above, the terminal device, in response to the instruction to generate audio based on text, determines the first text. The first text is the text used to generate the first audio, and the first text can be text of any format and content, without limitation here. The first audio refers to the audio generated based on the first text. The audio content of the first audio is determined based on the text content of the first text, and the audio content of the first audio is consistent with the text content of the first text. Here, "consistent" can mean that the audio content of the first audio is consistent with the text content of the first text sentence by sentence, or that the audio content of the first audio is not completely consistent with the text content of the first text sentence by sentence, but the audio content of the first audio can express the core meaning of the text content of the first text. The first audio can be a podcast generated based on the first text. A podcast is an audio program published via the Internet, which users can listen to online or download to their terminal devices for offline consumption. Podcasts typically consist of one or more hosts narrating the content to be conveyed.

[0021] In some embodiments, the terminal device responds to an instruction to generate audio from text, and if the instruction is associated with uploaded text, the uploaded text is used as the first text. In this embodiment, a user can upload text to a smart assistant and perform an operation requesting the generation of audio from the text. The smart assistant generates an instruction to generate audio from the text based on the user's operation of uploading text and requesting the generation of audio from the text, and responds to the instruction by using the user-uploaded text as the first text.

[0022] In some embodiments, in response to an instruction to generate audio from text, if the instruction is associated with an uploaded text link, the text corresponding to the uploaded text link is used as the first text. In this embodiment, a user can upload a text link to a smart assistant and perform an operation requesting the generation of audio from the text. The smart assistant generates an instruction to generate audio from the text based on the user's operation of uploading the text link and requesting the generation of audio from the text, and in response to the instruction, uses the text corresponding to the user's uploaded text link as the first text. This text link can be a storage link of the text on the user's terminal device, or a webpage link of the text on the network.

[0023] Figure 2a This is a schematic diagram illustrating the determination of a first text according to an embodiment of the present disclosure, such as... Figure 2aAs shown, users can trigger the "AI Podcast" control in the interface where they interact with the smart assistant. The smart assistant responds to this trigger by displaying two options: "Upload Document" and "Paste Article Link." Users can upload the first text by triggering the "Upload Document" option, or upload a text link of the first text by triggering the "Paste Article Link" option. After the user uploads the first text, the smart assistant generates an instruction to generate audio from the text based on the user's actions of triggering the "AI Podcast" control and uploading the first text, and responds to this instruction to determine the first text. The first text is either the text uploaded by the user or the text corresponding to the text link uploaded by the user.

[0024] In some embodiments, determining the first text in response to the instruction to generate audio based on the text includes: In response to an instruction to generate audio from text, if the instruction is associated with a selected dialogue message, the selected dialogue message is used as the first text; the dialogue message includes dialogue messages sent to the smart assistant or dialogue messages sent by the smart assistant; the smart assistant is used to generate the aforementioned audio configuration information and the first audio.

[0025] In this embodiment, the user engages in a dialogue with the intelligent assistant, selecting either a dialogue message sent by the user to the intelligent assistant or a dialogue message sent by the intelligent assistant to the user. Based on the selected dialogue message, the user requests the generation of audio from the text. The intelligent assistant generates an instruction to generate audio from the text based on the user's selected dialogue message and the request to generate audio from the text. In response to the instruction, the user-selected dialogue message associated with the instruction is used as the first text. The selected dialogue message can be a single dialogue message or multiple dialogue messages. When it is a single dialogue message, it can be all or part of the content of that single dialogue message.

[0026] Figure 2b A schematic diagram illustrating the determination of the first text provided in another embodiment of this disclosure, such as... Figure 2b As shown, users can send messages to the smart assistant, such as "Generate an article about family life," within the interface where they interact with the smart assistant. The smart assistant responds to this message and displays the generated article. Users can also select from the dialogue messages sent by the smart assistant, such as selecting a portion or all of the generated article (the image illustrates the entire article as an example). The smart assistant displays the operation panel corresponding to the user's selected dialogue message, and the user can trigger the "Generate Podcast" control in the operation panel. Based on the user's selection of the dialogue message and the triggering of the "Generate Podcast" control, the smart assistant generates an instruction to generate audio from the text and, in response to this instruction, identifies the user's selected dialogue message as the first text.

[0027] Through this embodiment, users can also select the desired dialogue message as the first text during a conversation with the intelligent assistant, and generate the first audio based on the first text, thereby improving the diversity and convenience of the way users provide the first text.

[0028] After determining the first audio, in step S104 above, the intelligent assistant displays at least one audio configuration information corresponding to the first audio. Specifically, the intelligent assistant can invoke a generative model to generate at least one audio configuration information corresponding to the first audio based on the text content of the first text, and then display each audio configuration information. The audio configuration information is used to represent the number of virtual characters in the first audio and the expressive style information of the virtual characters.

[0029] A virtual character is a role used to express the audio content in the first audio recording; it can be considered the anchor in the AI-generated first audio recording. The expressive style information of a virtual character includes at least one of the following: timbre information, speaking speed information, and emotional expression information. Specifically, timbre information represents the virtual character's voice, speaking speed information represents the speed of speech, and emotional expression information represents the virtual character's emotions. Examples of timbre information include: male timbre, female timbre, child timbre, student timbre, elderly timbre, lyrical timbre, bright timbre, and deep timbre. Examples of speaking speed information include: slow speaking speed, fast speaking speed, and moderate speaking speed. Examples of emotional expression information include: cheerful, sad, calm, relaxed, gentle, and rational.

[0030] Each audio configuration information represents the number of virtual characters and their expressive style. Each audio configuration is generated based on and matches the text content of the first text. Users can select the desired audio configuration as the first configuration and generate the first audio based on it.

[0031] Figure 3a This is a schematic diagram illustrating the display of audio configuration information according to an embodiment of this disclosure, such as... Figure 3a As shown, assuming the user is based on Figure 2b The scenario request generates the first audio for the selected dialogue message, such as generating a podcast. After the user triggers the "Generate Podcast" control, the smart assistant in the terminal device generates and displays the various audio configuration information. Figure 3aThe system displays three audio configuration information entries. The first entry indicates one virtual character in the first audio, with a cheerful and joyful middle-aged female voice. The second entry indicates two virtual characters in the first audio, with a cheerful and joyful middle-aged female voice and a gentle and kind middle-aged male voice, respectively. The third entry indicates three virtual characters in the first audio, with a cheerful and joyful middle-aged female voice, a gentle and kind middle-aged male voice, and a childlike male voice, respectively.

[0032] In some embodiments, the audio configuration information is generated based on the text content in the following way: Determine the content type of the text content, and determine the number of narrative perspectives included in the text content; The number of virtual characters is determined based on the number of narrative perspectives, and the expressive style information of the virtual characters matched in the first text is determined based on the content type. Based on the determined number of virtual characters and the expression style information of the virtual characters matched with the first text, audio configuration information is generated.

[0033] Taking the example of an intelligent assistant calling a generative model to generate at least one audio configuration information corresponding to a first audio based on the text content of a first text, the generative model first determines the content type of the first text. Content types can include, for example, stories, popular science, history, etc. Next, the generative model determines the number of narrative perspectives included in the text content. A narrative perspective refers to the viewpoint used to describe an event within the text content. For example, a narrative perspective could be the various characters in the event described in the text content, the various viewpoints in the event described in the text content, or the various time points in the event described in the text content. Finally, the number of narrative perspectives is determined, which includes, but is not limited to, one or more of the following: the number of characters, the number of viewpoints, and the number of time points in the text content. Time points can be exemplified by various natural days, various eras, and various seasons.

[0034] Then, using a generative model, the number of virtual characters is determined based on the number of narrative perspectives, resulting in one or more possible numbers. Next, using the generative model, the expressive style information of the virtual characters matched in the first text is determined based on the content type, resulting in one or more expressive style pieces of information. Finally, using the generative model, audio configuration information is generated based on the determined number of virtual characters and the expressive style information of the virtual characters matched in the first text. For example, combining the number of virtual characters with the expressive style information yields the audio configuration information.

[0035] Through this embodiment, the audio configuration information matching the first text can be accurately determined from two aspects: the content type of the first text and the number of narrative perspectives included in the first text.

[0036] In some embodiments, determining the number of virtual characters based on the number of narrative perspectives includes: The number of virtual characters is determined to be less than or equal to the number of narrative perspectives.

[0037] In this embodiment, when determining the number of virtual characters based on the number of narrative perspectives using a generative model, it can be determined that the number of virtual characters is less than or equal to the number of narrative perspectives. For example, a prompt message can be pre-input into the generative model, indicating that the number of generated virtual characters is less than or equal to the number of narrative perspectives. For instance, if the number of narrative perspectives is 3, then the number of virtual characters generated by the generative model is less than or equal to 3, meaning that the first audio recording has a maximum of 3 virtual characters used to express the audio content.

[0038] Generative models can intelligently analyze the text content of the first text when the number of virtual characters is less than or equal to the number of narrative perspectives, determining how many virtual characters can be matched with the first text. For example, in Figure 2b and Figure 3a In the example, the first text is an article about family life, describing two scenes: one in spring and one in summer for a family of three. Therefore, the generative model determines that the first text includes three narrative perspectives: the father, the mother, and the child. Based on this, the generative model determines that the number of virtual characters can be one, two, or three. Specifically, when there is one virtual character, the first audio uses that single virtual character to introduce the text content; when there are two virtual characters, the first audio uses two virtual characters to introduce the text content; and when there are three virtual characters, the first audio uses all three virtual characters to introduce the text content.

[0039] In other examples, the generative model can determine that the first text includes two narrative perspectives, namely spring and summer. Based on this, the generative model can determine that the number of virtual characters can be one or two. Specifically, when there is one virtual character, the first audio uses that one virtual character to introduce the text content of the first text; when there are two virtual characters, the first audio uses both virtual characters to introduce the text content of the first text. This situation is not illustrated in the accompanying diagram.

[0040] Therefore, through this embodiment, a generative model can be used to intelligently analyze the text content of the first text when the number of virtual characters is less than or equal to the number of narrative perspectives, and determine the number of virtual characters that match the first text, thereby accurately and intelligently determining the number of one or more virtual characters.

[0041] In some embodiments, determining the expressive style information of the virtual character matching the first text, based on the content type, includes: Among the existing expression style information of various virtual characters, the expression style information of the virtual character that matches the first text is selected according to the content type.

[0042] In this embodiment, a generative model can be used to select one or more expressive style information that matches the first text from the existing expressive style information of various virtual characters, based on the content type. The existing expressive style information of virtual characters refers to the expressive style information that has been created in advance through the generative model. Since one expressive style corresponds to one virtual character, this step can be considered as selecting one or more virtual characters that match the first text from the various virtual characters created by the generative model, based on the content type.

[0043] For example, in Figure 2b and Figure 3a In the example, the first text is an article about family life, describing a scene of a family of three—a father, mother, and child—during spring and another during summer. Therefore, the generative model determines the content type of the first text as "family life." Furthermore, the generative model identifies pre-created virtual characters matching the family life content, including a father, mother, and child. The father's expressive style information—"gentle and kind middle-aged male voice"—the mother's expressive style information—"cheerful and joyful middle-aged female voice"—and the child's expressive style information—"childlike boy's voice"—are identified as the expressive style information of the virtual characters matching the first text. Therefore, the expressive style information of the virtual characters matching the first text includes: a cheerful and joyful middle-aged female voice, a gentle and kind middle-aged male voice, and a childlike boy's voice.

[0044] In other examples, generative models can be used to determine the expressive style information of virtual characters matching family life content. This includes a cheerful and happy, slow-paced middle-aged female voice; a gentle and kind, slow-paced middle-aged male voice; and a childlike, slightly faster-paced male voice. In this example, the expressive style information also includes speaking speed information, which reflects the gentle and calm nature of the parents and the lively personality of the child in the first audio clip. This situation is no longer illustrated with accompanying diagrams.

[0045] Therefore, through this embodiment, a generative model can be used to select the expression style information of the virtual character that matches the first text from the existing expression style information of various virtual characters, based on the content type, thereby obtaining one or more expression style information and accurately and intelligently determining the expression style information.

[0046] In some embodiments, audio configuration information is generated based on the determined number of virtual characters and the expressive style information of the virtual characters matched with the first text, including: Based on the determined number of virtual characters, select one or more expression style information from the expression style information of the virtual characters matched in the first text; Based on the determined number of virtual characters and the selected expression style information, audio configuration information is generated.

[0047] In this step, firstly, using a generative model, one or more expression style information pieces are selected from the expression style information of the virtual characters matched in the first text, according to a determined number of virtual characters. The number of selected expression style information pieces is the same as the number of virtual characters. For example, in Figure 2b and Figure 3a In the example, the generative model determines that the number of virtual characters can be 1, 2, or 3. Then, from the expression style information of the virtual characters matched in the first text, we select 1 expression style information, 2 expression style information, and 3 expression style information.

[0048] Since the maximum number of expression style information selected from the expression style information of the first text matching is equal to the maximum number of virtual characters previously determined, it is required that the number of expression style information of the virtual characters in the first text matching is greater than or equal to the maximum number of virtual characters. Therefore, after selecting the expression style information of the virtual characters in the first text matching from the existing expression style information of each virtual character through the generative model, it can be determined whether the number of selected expression style information is greater than or equal to the maximum number of virtual characters. If so, the process of generating audio configuration information is executed; if not, the generative model is prompted to reselect expression style information from the existing expression style information of each virtual character until the above requirement is met.

[0049] Next, using a generative model, audio configuration information is generated based on the determined number of virtual characters and the selected expression style information. For each type of virtual character, this number can be combined with the expression style information selected based on that number to obtain one audio configuration information, and then multiple audio configuration information can be obtained. The number of audio configuration information is equal to the maximum value of the determined number of virtual characters.

[0050] For example, in Figure 2b and Figure 3a In the example, the generative model determines that the number of virtual characters can be 1, 2, or 3. Then, from the expression style information of the virtual characters matched in the first text, one expression style information is selected, such as a cheerful and happy middle-aged female voice; two expression style information information are selected, such as a cheerful and happy middle-aged female voice and a gentle and kind middle-aged male voice; and three expression style information information are selected, such as a cheerful and happy middle-aged female voice, a gentle and kind middle-aged male voice, and a childlike male voice. Finally, the combination yields three audio configuration information sets. The first audio configuration set indicates that there is one virtual character in the first audio, with the expression style information being a cheerful and joyful middle-aged female voice. The second audio configuration set indicates that there are two virtual characters in the first audio, with the expression style information being a cheerful and joyful middle-aged female voice and a gentle and kind middle-aged male voice, respectively. The third audio configuration set indicates that there are three virtual characters in the first audio, with the expression style information being a cheerful and joyful middle-aged female voice, a gentle and kind middle-aged male voice, and a childlike male voice, respectively.

[0051] Of course, in other examples, the following three audio configuration information can also be obtained. The first audio configuration information indicates that the first audio has one virtual character, with a vocal style of either a young boy's voice or a gentle, middle-aged man's voice. The second audio configuration information indicates that the first audio has two virtual characters, with vocal styles of either a cheerful, middle-aged woman's voice and a young boy's voice, or a gentle, middle-aged man's voice and a young boy's voice. The third audio configuration information indicates that the first audio has three virtual characters, with vocal styles of either a cheerful, middle-aged woman's voice, a gentle, middle-aged man's voice, and a young boy's voice. This case is not illustrated in the diagram.

[0052] Therefore, through this embodiment, one or more expression style information can be selected from the expression style information of the virtual characters matched in the first text according to the determined number of virtual characters. Audio configuration information is generated based on the determined number of virtual characters and the selected expression style information. By combining them, the generative module can conveniently and quickly generate audio configuration information.

[0053] In other embodiments, the generative model can also pre-define the number of virtual characters. For each pre-defined number of virtual characters, based on the content type, select the same number of expression style information that matches the first text from the existing expression style information of each virtual character. Combine the number of virtual characters and the selected expression style information into an audio configuration information. For example, if the pre-definement allows for one or two virtual characters, then for the case of one virtual character, based on the content type, select one expression style information that matches the first text and generate an audio configuration information indicating that the first audio has one virtual character and its expression style; and for the case of two virtual characters, based on the content type, select two expression style information that matches the first text and generate an audio configuration information indicating that the first audio has two virtual characters and the expression style of each virtual character.

[0054] After generating and displaying each audio configuration information, in step S106 above, in response to the user's selection instruction for the first configuration information in each audio configuration information, a first audio is output. The first audio is generated based on the text content of the first text and the first configuration information. The first audio expresses the text content of the first text through the first configuration information.

[0055] For example, in Figure 2b and Figure 3a In the example, if the user selects the third audio configuration information, the first audio is generated according to the third audio configuration information. In the first audio, the text content of the first text is narrated from the perspectives of the father, mother and child respectively. Figure 3b This is a schematic diagram of the output of a first audio signal provided in an embodiment of the present disclosure, as shown below. Figure 3b As shown, this is an interface that outputs the first audio to the user during a conversation with the smart assistant, and the user can play the first audio through this interface.

[0056] In some embodiments, the first audio is generated based on the text content of the first text and the first configuration information in the following manner: Based on the text content and the first configuration information, a character script is generated; the number of character scripts is the same as the number of virtual characters represented by the first configuration information; the content of the character script matches the expression style information of the virtual characters represented by the first configuration information. Based on the expression style information of the virtual character as indicated by the first configuration information, the character script is expressed to obtain the first audio.

[0057] In this embodiment, a generative model can be invoked by an intelligent assistant to generate the first audio based on the text content of the first text and the first configuration information. First, the generative model generates character scripts based on the text content and the first configuration information. The number of character scripts is the same as the number of virtual characters represented by the first configuration information. That is, each character script corresponds one-to-one with a virtual character in the first audio, and each virtual character has its own character script. The character script can be in text format. The content of the character script matches the expression style information of the virtual character represented by the first configuration information. For example, in the previous example, the content of the character script for a middle-aged woman matches the expression style of a cheerful and joyful middle-aged woman's voice; the content of the character script for a middle-aged man matches the expression style of a gentle and kind middle-aged man's voice; and the content of the character script for a boy matches the expression style of a childish boy's voice.

[0058] Then, using a generative model, each virtual character's script is expressed according to the expression style information represented by the first configuration information, resulting in the first audio. In other words, the audio generated by expressing the respective scripts of each virtual character in the first audio is the first audio.

[0059] Therefore, through this embodiment, the text content of the first text can be decomposed to obtain the character script of each virtual character in the first audio, and the text content of the character script can be matched with the expression style information of the virtual character in the first audio. Then, according to the expression style information of the virtual character in the first audio, the character script of each virtual character is expressed to obtain the first audio, thereby ensuring that the first audio can express the text content of the first text through the first configuration information.

[0060] In some embodiments, a character script is generated based on text content and first configuration information, including: The text content is divided according to the number of virtual characters represented by the first configuration information to obtain one or more sub-text contents; Based on the expression style information of the virtual character represented by the first configuration information, the sub-text content is adjusted to obtain the character manuscript.

[0061] First, the text content can be divided into one or more sub-text contents according to the number of virtual characters represented by the first configuration information using a generative model. The number of sub-text contents is the same as the number of virtual characters represented by the first configuration information. The sub-text contents correspond one-to-one with the virtual characters in the first audio. It can be considered that the sub-text contents obtained through this step are the initial character scripts for each virtual character in the first audio.

[0062] Then, based on the expression style information of the virtual character represented by the first configuration information, the sub-text content is adjusted to obtain the character script. For each sub-text content, the sub-text content is adjusted according to the expression style information of the virtual character to which it belongs to, to obtain the character script for that virtual character. Adjusting the sub-text content may involve adding, deleting, or modifying the wording and phrases, so that the character script conforms to the expression style information of the corresponding virtual character.

[0063] For example, for the virtual character of a boy, sentences in the corresponding subtext that do not conform to the virtual character of a boy are adjusted so that the text of the resulting character script conforms to the virtual character of a boy and the expressive style of a young boy's voice.

[0064] Through this embodiment, the sub-text content can be adjusted according to the expression style information of the virtual character represented by the first configuration information. For example, the wording and sentences in the sub-text content can be deleted or modified, thereby ensuring that the text content of the obtained character script matches the expression style information of the virtual character in the first audio, improving the adaptability of each virtual character in the first audio to its respective expression content (character script), and improving the audio effect of the first audio.

[0065] In other embodiments, the first audio can also be generated by inputting first configuration information and first text into a generative model, and inputting prompts such as "Please generate a first audio that matches the first configuration information based on this text. The first audio needs to express the text content according to the expression style information of each virtual character." Through the generative model, the text content of the first text can be decomposed and modified in the above manner to obtain the character script for each virtual character represented by the first configuration information, or, based on the text content of the first text, the character script for each virtual character represented by the first configuration information can be regenerated. The character script for each virtual character matches the expression style information of each virtual character. Then, through the generative model, the character script is expressed according to the expression style information of each virtual character to obtain the first audio.

[0066] The above describes the process of generating one or more audio configuration information based on the text content of a first text file, and generating a first audio file based on the first configuration information selected by the user. The audio configuration information is used to represent the number of virtual characters in the first audio file and the expressive style information of the virtual characters. In one case, it can be as follows: Figure 3aAs shown, each audio configuration information displays the number of virtual characters and their expressive style. Alternatively, the audio configuration information can be displayed as "Smart Match 1," "Smart Match 2," or similar intelligent matching information. Regardless of how the audio configuration information is displayed, when a user selects an audio configuration information, the intelligent assistant can play a sample audio of the selected information. Playing the sample audio helps the user understand the number of virtual characters and their expressive style within the selected audio configuration information.

[0067] In some embodiments, in addition to displaying audio configuration information, the above process may also include: Displays the style information of the specified object; In response to a selection instruction for the expression style information of a specified object, a second audio is output; the second audio is generated based on the text content of the first text and the expression style information of the specified object; the text content of the first text is expressed in the second audio through the expression style information of the specified object.

[0068] The designated object can be an object pre-created within the intelligent assistant. For example, based on user-provided creation information, an object matching that creation information can be created as the designated object. The creation information describes the expressive style of the designated object, such as "Please create the voice of a lively and cute boy around 5 years old." Alternatively, with user authorization, audio of the user reading a specified text can be obtained, and an object matching the user's expressive style can be created based on that audio as the designated object. The number of designated objects can be one or multiple.

[0069] The interface displays the expression style information of a specified object. Responding to a selection instruction for the expression style information of the specified object, a second audio is generated and output based on the text content of the first text and the expression style information of the specified object. The second audio expresses the text content of the first text through the expression style information of the specified object. Users can select one or more expression style information of specified objects in the interface and create a second audio. The number of virtual characters in the second audio is the number of specified objects selected by the user; each selected specified object is a virtual character. After creating the first audio, the user can select expression style information of one or more specified objects again to create a second audio. The audio content of the first audio and the audio content of the second audio can be the same or different, but both are used to represent the text content of the first text.

[0070] Therefore, through this embodiment, in addition to intelligently matching the audio configuration information of the first text, it can also provide the expression style information of the specified object for the user to choose from, thereby providing the user with multiple ways to generate audio and improving the diversity of generated audio.

[0071] Figure 4 This is a schematic diagram illustrating the display of audio configuration information provided in yet another embodiment of this disclosure, such as... Figure 4 As shown, taking the generation of audio configuration information based on the text content of the first text as an example, the audio configuration information can be displayed as "intelligent matching" information. When the user selects the audio configuration information, the intelligent assistant can play a sample audio of the audio configuration information. By playing the sample audio, the user can easily understand the number of virtual characters and the expression style information of the virtual characters in the audio configuration information. Figure 4 The system also displays the expressive style information of a specified object, such as "the voice of specified object 1". When the user selects the expressive style information of specified object 1, the smart assistant can play sample audio of specified object 1. By playing sample audio, the user can easily understand the expressive style information of specified object 1. Figure 4 The system also displays several other configuration options, such as "Double - A + B - Friendly and Lively" and "Double - C + D - Professional and Serious." These configuration options are pre-created by the intelligent assistant and are used to indicate the number and expression style of the virtual characters. When users are not satisfied with the intelligently recommended audio configuration options or the expression style of a specific character, they can select the desired configuration options from these pre-created options. When a user selects a pre-created configuration option, the intelligent assistant can play a sample audio file for that configuration option. By playing the sample audio, users can easily understand the number and expression style of the virtual characters in the pre-created configuration option.

[0072] In some embodiments, the following can also be performed: In response to a style switching command for the first audio, display the configuration information for each audio. In response to the selection instruction for the second configuration information in each audio configuration information, the first audio after style switching is output; the first audio after style switching is generated based on the text content of the first text and the second configuration information; the text content of the first text is expressed in the first audio after style switching through the second configuration information.

[0073] When multiple audio configuration information items are generated based on the text content of the first text, the user can switch the expression style of the first audio after listening to it. For example, by triggering the "Timbre" option, the intelligent assistant responds to the trigger by generating a style switching command for the first audio and, in response to the command, displays each audio configuration information item. After the user selects a second configuration information item from each audio configuration information item, the intelligent assistant responds to the selection command for the second configuration information item by generating and outputting the first audio after the style switch. The first audio after the style switch is generated based on the text content of the first text and the second configuration information. The text content of the first text is expressed in the first audio after the style switch through the second configuration information. The audio content of the first audio after the style switch can be the same as or different from the audio content of the first audio before the switch, but both are used to represent the text content of the first text.

[0074] refer to Figure 4 In addition to displaying audio configuration information, it can also display the expression style information of a specified object, such as "the voice of specified object 1", as well as the configuration information pre-created by each smart assistant, such as "two people, A + B, friendly and lively" and "two people, C + D, professional and serious". Figure 5a This is a schematic diagram of the output of a first audio signal provided in another embodiment of this disclosure. Figure 5b This is a schematic diagram of the output of a first audio signal provided in yet another embodiment of this disclosure. The user triggers... Figure 3b The first audio interface can be navigated to... Figure 5a The page shown is for playing the first audio file, as follows: Figure 5a As shown, the page playing the first audio also displays a "Sound" component. In other ways, this "Sound" component can be invoked by the user triggering the "More" component (not shown in the image) on the page. If the user triggers the "Sound" component, it displays... Figure 5b The page shown is in Figure 5b In China, with Figure 4 For example, it shows that Figure 4 The system provides audio configuration information, expression style information for a specified object, and pre-created configuration information for various intelligent assistants. Users can select and switch between the provided information, and the intelligent assistant responds to the user's switching operation by regenerating a new first audio. This achieves the goal of switching the expression style after the first audio is generated. Figure 5b This is just an illustrative example. Figure 5b In the middle, it can also display Figure 3a The audio configuration information is not shown here.

[0075] The various models of the intelligent assistant mentioned above, such as generative models and AI models, can be deployed on terminal devices or on the server corresponding to the intelligent assistant; this embodiment does not impose any limitations. When deployed on a terminal device, the terminal device generates various audio configuration information, a first audio, a second audio, and the first audio after style switching. When deployed on a server, the server generates various audio configuration information, a first audio, a second audio, and the first audio after style switching.

[0076] In summary, through this embodiment, at least one audio configuration information corresponding to the first audio can be determined based on the text content of the first text, and the first audio can be generated based on the first configuration information selected from each audio configuration information. Therefore, the first audio can express the text content of the first text according to the number of virtual characters and the expression style information of the virtual characters represented by the first configuration information. In this way, by generating one or more selectable audio configuration information based on the text content, the problem of limited audio effects of audio generated according to the default configuration information is alleviated, and the diversity of generated audio effects is improved.

[0077] Figure 6 This is a schematic diagram of the structure of an audio generation apparatus provided in an embodiment of the present disclosure, as shown below. Figure 6 As shown, the device includes: The text determination unit 61 is configured to determine a first text in response to an instruction to generate audio based on the text; the first text is used to generate the first audio. The information display unit 62 is used to display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; The audio output unit 63 is configured to output the first audio in response to a selection instruction for the first configuration information in each of the audio configuration information; the first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

[0078] Optionally, the text determination unit 61 is specifically configured to: respond to an instruction to generate audio based on text, and if the instruction is associated with a selected dialogue message, use the dialogue message as the first text; the dialogue message includes a dialogue message sent to a smart assistant or includes a dialogue message sent by the smart assistant; the smart assistant is used to generate the audio configuration information and the first audio.

[0079] Optionally, the audio configuration information is generated based on the text content in the following manner: determining the content type of the text content, and determining the number of narrative perspectives included in the text content; determining the number of virtual characters based on the number of narrative perspectives, and determining the expression style information of the virtual characters matched by the first text based on the content type; generating the audio configuration information based on the determined number of virtual characters and the expression style information of the virtual characters matched by the first text.

[0080] Optionally, determining the number of virtual characters based on the number of narrative perspectives includes: determining that the number of virtual characters is less than or equal to the number of narrative perspectives.

[0081] Optionally, determining the expression style information of the virtual character matched by the first text based on the content type includes: selecting the expression style information of the virtual character matched by the first text from the existing expression style information of various virtual characters, based on the content type.

[0082] Optionally, generating the audio configuration information based on the determined number of virtual characters and the expression style information of the virtual characters matched with the first text includes: selecting one or more expression style information from the expression style information of the virtual characters matched with the first text according to the determined number of virtual characters; and generating the audio configuration information based on the determined number of virtual characters and the selected expression style information.

[0083] Optionally, the first audio is generated based on the text content and the first configuration information in the following manner: generating character scripts based on the text content and the first configuration information; the number of character scripts is the same as the number of virtual characters represented by the first configuration information; the content of the character scripts matches the expression style information of the virtual characters represented by the first configuration information; and expressing the character scripts according to the expression style information of the virtual characters represented by the first configuration information to obtain the first audio.

[0084] Optionally, generating the character script based on the text content and the first configuration information includes: dividing the text content into one or more sub-text contents according to the number of virtual characters represented by the first configuration information; and adjusting the sub-text contents according to the expression style information of the virtual characters represented by the first configuration information to obtain the character script.

[0085] Optionally, it further includes a designated object display unit, configured to: display the expression style information of the designated object; output a second audio in response to a selection instruction for the expression style information of the designated object; the second audio is generated based on the text content and the expression style information of the designated object; and the text content is expressed in the second audio through the expression style information of the designated object.

[0086] Optionally, it further includes a style switching unit, configured to: display each of the audio configuration information in response to a style switching instruction for the first audio; output the style-switched first audio in response to a selection instruction for the second configuration information in each of the audio configuration information; the style-switched first audio is generated based on the text content and the second configuration information; the style-switched first audio expresses the text content through the second configuration information.

[0087] Optionally, the expressive style information includes at least one of timbre information, expressive speed information, and expressive emotion information.

[0088] The audio generation apparatus in this embodiment can implement the various processes of the above-described audio generation method embodiment and achieve the same effect and function, which will not be repeated here.

[0089] One embodiment of this disclosure also provides an electronic device. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, as shown below. Figure 7 As shown, electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 701 and memories 702, with the memory 702 storing one or more application programs or data. The memory 702 can be temporary or persistent storage. The application programs stored in the memory 702 may include one or more modules (not shown), each module including a series of computer-executable instructions from the electronic device. Furthermore, the processor 701 may be configured to communicate with the memory 702, executing the series of computer-executable instructions stored in the memory 702 on the electronic device. The electronic device may also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input or output interfaces 705, one or more keyboards 706, etc.

[0090] In one specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following process: In response to an instruction to generate audio from text, a first text is determined; the first text is used to generate the first audio. Display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; In response to a selection instruction for the first configuration information in each of the audio configuration information, the first audio is output; the first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

[0091] The electronic device in this embodiment can implement the various processes of the above-described audio generation method embodiment and achieve the same effect and function, which will not be repeated here.

[0092] Another embodiment of this disclosure also provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process: In response to an instruction to generate audio from text, a first text is determined; the first text is used to generate the first audio. Display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; In response to a selection instruction for the first configuration information in each of the audio configuration information, the first audio is output; the first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

[0093] The computer-readable storage medium in this disclosure embodiment can implement the various processes of the above-described audio generation method embodiment and achieve the same effect and function, which will not be repeated here.

[0094] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process: In response to an instruction to generate audio from text, a first text is determined; the first text is used to generate the first audio. Display at least one audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; In response to a selection instruction for the first configuration information in each of the audio configuration information, the first audio is output; the first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

[0095] The computer program product in this disclosure embodiment can implement the various processes of the above-described audio generation method embodiment and achieve the same effect and function, which will not be repeated here.

[0096] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.

[0097] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0098] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0099] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0100] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.

[0101] Those skilled in the art will understand that one or more embodiments of this disclosure can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0102] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0103] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0104] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0105] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0106] One or more embodiments of this disclosure can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0107] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0108] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. An audio generation method, characterized in that, include: In response to an instruction to generate audio from text, determine the first text; The first text is used to generate the first audio; Display at least two audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; the expression style information of the virtual characters includes at least one of timbre information, expression speed information, and expression emotion information; In response to a selection instruction for the first configuration information in each of the audio configuration information, the first audio is output; The first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

2. The method according to claim 1, characterized in that, The step of determining the first text in response to an instruction to generate audio based on text includes: In response to an instruction to generate audio from text, if the instruction is associated with a selected dialogue message, the dialogue message is used as the first text; the dialogue message includes a dialogue message sent to a smart assistant or a dialogue message sent by the smart assistant; the smart assistant is used to generate the audio configuration information and the first audio.

3. The method according to claim 1, characterized in that, The audio configuration information is generated based on the text content in the following way: Determine the content type of the text content, and determine the number of narrative perspectives included in the text content; Based on the number of narrative perspectives, determine the number of virtual characters; based on the content type, determine the expressive style information of the virtual characters matched by the first text. The audio configuration information is generated based on the determined number of virtual characters and the expression style information of the virtual characters matched with the first text.

4. The method according to claim 3, characterized in that, Determining the number of virtual characters based on the number of narrative perspectives includes: The number of virtual characters is determined to be less than or equal to the number of narrative perspectives.

5. The method according to claim 3, characterized in that, The step of determining the expression style information of the virtual character matched by the first text based on the content type includes: Among the existing expression style information of various virtual characters, the expression style information of the virtual character that matches the first text is selected according to the content type.

6. The method according to claim 3, characterized in that, The step of generating the audio configuration information based on the determined number of virtual characters and the expression style information of the virtual characters matched with the first text includes: Based on the determined number of virtual characters, select one or more of the expression style information from the expression style information of the virtual characters matched in the first text; The audio configuration information is generated based on the determined number of virtual characters and the selected expression style information.

7. The method according to claim 1, characterized in that, The first audio is generated based on the text content and the first configuration information in the following manner: Based on the text content and the first configuration information, a character script is generated; the number of character scripts is the same as the number of virtual characters represented by the first configuration information; the content of the character script matches the expression style information of the virtual characters represented by the first configuration information. The character script is expressed according to the expression style information of the virtual character represented by the first configuration information, and the first audio is obtained.

8. The method according to claim 7, characterized in that, The step of generating a character script based on the text content and the first configuration information includes: The text content is divided according to the number of virtual characters represented by the first configuration information to obtain one or more sub-text contents; Based on the expression style information of the virtual character represented by the first configuration information, the sub-text content is adjusted to obtain the character manuscript.

9. The method according to claim 1, characterized in that, Also includes: Displays the style information of the specified object; In response to a selection instruction for the expressive style information of the specified object, a second audio is output; The second audio is generated based on the text content and the expression style information of the specified object; the text content is expressed in the second audio through the expression style information of the specified object.

10. The method according to claim 1, characterized in that, Also includes: In response to a style switching command for the first audio, display the configuration information for each of the audio components; In response to a selection instruction for the second configuration information in each of the aforementioned audio configuration information, the first audio after style switching is output; The first audio after style switching is generated based on the text content and the second configuration information; the text content is expressed in the first audio after style switching through the second configuration information.

11. An audio generation apparatus, characterized in that, include: A text determination unit is used to determine a first text in response to an instruction to generate audio based on the text; The first text is used to generate the first audio; An information display unit is used to display at least two audio configuration information corresponding to the first audio; the audio configuration information is used to indicate the number of virtual characters in the first audio and the expression style information of the virtual characters; the audio configuration information is generated based on the text content of the first text; the expression style information of the virtual characters includes at least one of timbre information, expression speed information, and expression emotion information; An audio output unit is configured to output the first audio in response to a selection instruction for the first configuration information in each of the audio configuration information. The first audio is generated based on the text content and the first configuration information; the text content is expressed in the first audio through the first configuration information.

12. An electronic device, characterized in that, include: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions that, when executed by a processor, implement the method described in any one of claims 1-10.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-10.

Citation Information

Patent Citations

  • Dubbing interaction method and device, computer equipment and storage medium

    CN117075839A

  • Speech synthesis method, device, equipment, medium and program product

    CN118553229A