Audio program generation method, audio program generation device, storage medium and electronic equipment
By processing topic information and resource information through the generation model, and generating audio program content text information, the problem of high complexity in audio program production in the prior art is solved, and efficient audio program generation is achieved.
Patent Information
- Application Number
- CN202510384853.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
AI Technical Summary
The existing audio program production requires professional requirements and hardware support, which consumes a lot of time and is inefficient in production.
The first audio is obtained by processing the subject information through the first generation model, and the second generation model processes the resource information and program style type information, generates program content text information, and finally synthesizes the target audio program.
It simplifies the audio program production process, reduces the professional requirements of the producers, and realizes convenient and fast audio program generation.
Smart Images

Figure CN120260524A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology. More specifically, embodiments of the present disclosure relate to an audio program generation method, an audio program generation device, a computer-readable storage medium, and an electronic device. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure recited in the claims. The description herein is not admitted to be prior art merely by including it in this section.
[0003] With the increasing demand of users for multimedia information, in addition to video programs, a multimedia form of audio programs has also been proposed. Generally, an audio program can play a specific audio while playing an explanation about specific content, enabling users to immerse themselves in listening to the audio and information related to the audio. Existing methods for producing audio programs usually require producers to manually perform production processes such as selecting music, writing copy, and recording. The production of audio programs not only has certain professional requirements and hardware requirements for producers, but also requires a large amount of time cost, and the production efficiency is low. Summary of the Invention
[0004] However, the production of existing audio programs has certain professional requirements and hardware requirements for producers, and also requires a large amount of time cost.
[0005] Therefore, there is a great need for an audio program generation method that can generate audio programs quickly, conveniently, and effectively.
[0006] In this context, embodiments of the present disclosure are expected to provide an audio program generation method, an audio program generation device, a computer-readable storage medium, and an electronic device.
[0007] According to a first aspect of the present disclosure, there is provided an audio program generation method, including: obtaining theme information of an audio program to be generated; processing the theme information through a first generation model to obtain one or more first audios; obtaining resource information corresponding to the first audio and program style type information; processing the resource information and the program style type information through a second generation model to obtain program content text information; and generating a target audio program according to the first audio and the program content text information.
[0008] In one embodiment, the obtaining theme information of the audio program to be generated includes: obtaining keywords of the audio program to be generated, and constructing a first prompt word according to the keywords of the audio program to be generated; and processing the first prompt word through a third generation model to obtain the theme information of the audio program to be generated.
[0009] In one embodiment, processing the first prompt by the third generation model to obtain the theme information of the to-be-generated audio program includes: processing the first prompt by the third generation model to obtain multiple candidate theme information; in response to receiving a theme selection operation input by the user, determining the theme information of the to-be-generated audio program from the multiple candidate theme information.
[0010] In one embodiment, processing the theme information by the first generation model to obtain one or more first audios includes: constructing a second prompt according to the theme information of the to-be-generated audio program; processing the second prompt by the first generation model to obtain one or more first audios.
[0011] In one embodiment, the method further includes: obtaining audio creator information; the constructing the second prompt according to the theme information of the to-be-generated audio program includes: constructing the second prompt according to the theme information of the to-be-generated audio program and the audio creator information.
[0012] In one embodiment, processing the resource information and the program style type information by the second generation model to obtain program content text information includes: constructing a third prompt according to the resource information and the program style type information; processing the third prompt by the second generation model to obtain the program content text information.
[0013] In one embodiment, generating a target audio program according to the first audio and the program content text information includes: obtaining a second audio based on the program content text information; synthesizing the first audio and the second audio to obtain the target audio program.
[0014] In one embodiment, the obtaining the second audio based on the program content text information includes: obtaining the timbre information of the target object; processing the program content text information and the timbre information by a text-to-speech model to obtain the second audio.
[0015] In one embodiment, the synthesizing the first audio and the second audio to obtain the target audio program includes: synthesizing the first audio and the second audio, and gradually increasing the volume of the first audio within a first time period starting from the beginning of the mixing time period of the first audio and the second audio; and / or, synthesizing the first audio and the second audio, and gradually decreasing the volume of the first audio within a second time period at the end of the mixing time period of the first audio and the second audio.
[0016] In one embodiment, synthesizing the first audio and the second audio to obtain the target audio program includes: splicing the first audio and the second audio in sequence, and gradually reducing the volume of the first audio within a third time period at the end of the first audio.
[0017] In one embodiment, the method further includes: constructing a fourth prompt word according to the theme information, the first audio, and the program content text information; processing the fourth prompt word through a fourth generation model to obtain the description information of the target audio program; and publishing the target audio program and the description information of the target audio program.
[0018] According to a second aspect of the present disclosure, there is provided an audio program generation device, including: a theme information acquisition module, configured to acquire the theme information of the audio program to be generated; a first audio acquisition module, configured to process the theme information through a first generation model to obtain one or more first audios; an audio information acquisition module, configured to acquire the resource information corresponding to the first audio and the program style type information; a text information acquisition module, configured to process the resource information and the program style type information through a second generation model to obtain the program content text information; and an audio program generation module, configured to generate a target audio program according to the first audio and the program content text information.
[0019] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the audio program generation method according to the first aspect and its possible implementation manners are implemented.
[0020] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory, configured to store the executable instructions of the processor; wherein the processor is configured to execute the audio program generation method according to the first aspect and its possible implementation manners by executing the executable instructions.
[0021] In the solution of the present disclosure, on the one hand, this exemplary embodiment proposes a new method for generating an audio program. By processing the obtained theme information through a first generation model, a first audio can be obtained. By processing the resource information and the program style type information through a second generation model, program content text information can be obtained. Compared with the prior art, in this exemplary embodiment, there is no need for the producer to manually select the audio or write the program content. The corresponding information can be obtained through the generation model, which simplifies the production complexity of the audio program and improves the production efficiency of the audio program. On the other hand, in this exemplary embodiment, one or more first audios and program content text information are obtained through the generation model, and a target audio program can be generated according to the first audio and the program content text information. The production process of the audio program has relatively low professional requirements for the producer, enabling more users to conveniently and quickly produce the desired audio program according to their customized needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 FIG. shows a schematic diagram of a system architecture in this exemplary embodiment;
[0023] Figure 2 FIG. shows a flowchart of a method for generating an audio program in this exemplary embodiment;
[0024] Figure 3 FIG. shows a schematic diagram of a framework process for generating program content text information in this exemplary embodiment;
[0025] Figure 4 FIG. shows a schematic diagram of synthesizing a first audio and a second audio and setting a fade-in effect for the first audio in this exemplary embodiment;
[0026] Figure 5 FIG. shows a schematic diagram of synthesizing a first audio and a second audio and setting a fade-out effect for the first audio in this exemplary embodiment;
[0027] Figure 6 FIG. shows another schematic diagram of setting a fade-out effect for the first audio in this exemplary embodiment;
[0028] Figure 7 FIG. shows a framework flowchart of a method for generating an audio program in this exemplary embodiment;
[0029] Figure 8 FIG. shows a schematic diagram of the structure of an audio program generation device in this exemplary embodiment;
[0030] Figure 9 FIG. shows a schematic diagram of the structure of an electronic device in this exemplary embodiment.
[0031] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation Modes
[0032] The principles and spirit of the present disclosure will be described below with reference to several exemplary implementation modes. It should be understood that these implementation modes are provided only to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these implementation modes are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0033] The implementation modes of the present disclosure can be implemented as a system, a device, an equipment, a method or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0034] The principles and spirit of the present disclosure will be elaborated in detail below with reference to several representative implementation modes of the present disclosure. Summary of the Invention
[0036] The inventor of the present disclosure has found that the production of existing audio programs has certain professional requirements and hardware requirements for the production party, and also requires a large amount of time cost. Specifically, in the related art, the producer needs to first select a theme to be explained, and select one or more songs according to the theme. Then, according to the selected theme and songs, collect materials, write copy, and further, start recording the program through professional equipment, and finally process the recorded program to obtain the audio program. Each stage requires the producer to execute manually, which takes more time and has certain requirements for the producer's production level.
[0037] In view of the above, the present disclosure provides an audio program generation method, an audio program generation device, a computer-readable storage medium and an electronic device. On the one hand, this exemplary embodiment proposes a new audio program generation method. By processing the obtained theme information through a first generation model, a first audio can be obtained. By processing the resource information and the program style type information through a second generation model, program content text information can be obtained. Compared with the prior art, this exemplary embodiment does not require the producer to manually select the audio or write the program content, and the corresponding information can be obtained through the generation model, which simplifies the production complexity of the audio program and improves the production efficiency of the audio program. On the other hand, this exemplary embodiment obtains one or more first audios and program content text information through the generation model, and can generate a target audio program according to the first audio and the program content text information. The production process of the audio program has relatively low professional requirements for the producer, and can enable more users to conveniently and quickly produce the desired audio program according to their customized needs.
[0038] The following specifically introduces various non-limiting embodiments of the present disclosure.
[0039] Overview of Application Scenarios
[0040] It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not restricted in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0041] The embodiments of the present disclosure can be applied to related scenarios of audio program production. The following specifically describes the application scenarios in combination with the system architecture.
[0042] Figure 1 A schematic diagram of a system architecture for audio program production is shown. The system architecture includes a user device 110 and a server 120. The user device 110 can be an electronic device such as a mobile phone or a computer. The user can perform interactive operations on the user device 110 to execute related functions. For example, the user can input theme information on the user device 110 to cause the user device 110 to send the theme information to the server 120 for generating an audio program; or the server 120 can return the intermediate processing process to the user device 110 for the user to confirm it. For example, after obtaining one or more first audios, the server 120 can return them to the user device 110, and the user can select or confirm them, etc. In the present exemplary embodiment, a first generation model and a second generation model can be configured in the server 120. The first generation model is used to process the theme information to obtain one or more first audios, and the second generation model can be used to process the resource information and the program style type information to obtain the program content text information. The server 120 can generate a target audio program according to the first audio and the program content text information and return it to the user device 110.
[0043] Exemplary Method
[0044] An exemplary embodiment of the present disclosure provides an audio program generation method. Referring to Figure 2 as shown, the method may include steps S210 to S250. The following makes a specific description of Figure 2 each step in.
[0045] Referring to Figure 2 , in step S210, obtain the theme information of the audio program to be generated.
[0046] Among them, the theme information refers to the main idea information used to determine the first audio included in the to-be-generated audio program. The theme information can be about people, objects, events, etc. For example, the theme of the to-be-generated audio program can be "xx Singer: The Shining Star in the Music Industry" or "xx Song: The Story Behind the Creation", etc. The theme information can be input by the user himself / herself. For example, the user inputs the theme information in the form of text input or voice input, etc. Or the theme information can also be obtained by the server through analysis of the relevant information input by the user. For example, the user inputs a text paragraph or relevant keywords, etc., and the generation model configured by the server processes the input information to determine the theme information of the to-be-generated audio program, etc.
[0047] In an exemplary embodiment, the above-mentioned obtaining the theme information of the to-be-generated audio program may include:
[0048] Obtain the keywords of the to-be-generated audio program, and construct a first prompt word according to the keywords of the to-be-generated audio program;
[0049] Process the first prompt word through a third generation model to obtain the theme information of the to-be-generated audio program.
[0050] The keywords of the to-be-generated audio program refer to the relevant words used to generate the audio program. They can be people, objects, events, styles, etc. The keywords can be obtained from the information input by the user. For example, if the user directly inputs "xx Singer", "xx Singer" can be used as the keyword. Or if the user inputs "90s music style", then "90s music style" is used as the keyword. Or the user inputs a text, and by performing word segmentation and analysis on the text, the keywords are determined therefrom. The keywords can also be obtained from the acquired files. For example, a file or a screenshot including text information can be uploaded, and by analyzing the text information in the file and the screenshot, the keywords are determined therefrom.
[0051] In this exemplary embodiment, a first prompt word can be constructed according to the keywords of the to-be-generated audio program. For example, the following shows an example information of a first prompt word:
[0052] "You are a music radio host. You plan to record 10 music radio programs with the theme of [{{keyword}}]. Please diverge based on this theme and output {{limit}} eye-catching music radio program titles. Each title needs to be between 10 and 30 Chinese characters, and ideally around 15 Chinese characters. There should be no 'from...to...' in the title copy.
[0053] The output format is: json string list"
[0054] Among them, keyword is the keyword.
[0055] The third generation model refers to a model used to generate topic information, which can be an LLM (Large Language Model). The third generation model can process the first prompt word generated according to keywords to generate topic information. The generated topic information can be one or more. When multiple pieces of topic information are generated, the user can select the topic information as needed, etc.
[0056] In an exemplary embodiment, processing the first prompt word through the third generation model to obtain the topic information of the audio program to be generated may include:
[0057] Processing the first prompt word through the third generation model to obtain multiple candidate topic information;
[0058] In response to receiving the topic selection operation input by the user, determine the topic information of the audio program to be generated from the multiple candidate topic information.
[0059] Candidate topic information refers to the alternative topic information generated by the third generation model according to the first prompt word. For example, according to the first prompt word, multiple candidate topic information such as "xx Singer: The Shining Star of Pop Music", "xx Singer in the Song, the Perfect Fusion of Emotion and Skill", "The Music Journey of xx Singer: From Innocence to Maturity", "The Stories and Inspirations behind xx Singer's Creations", "Exploring xx Singer: The Artistic Life of a Musical Talent" can be generated. The candidate topic information can be directly used as the topic information for producing the audio program. When there are multiple pieces of candidate topic information, one or more pieces of topic information that are more in line with the conditions or user preferences can also be selected from the multiple candidate topic information for producing the audio program.
[0060] Specifically, when the third generation model generates multiple candidate topic information, the user can perform a topic selection operation to determine the preferred or interesting topic information from the multiple candidate topic information. For example, the user can determine the topic information of the audio program to be generated through a selection operation among the options of the multiple candidate topic information; or the user can determine the topic information of the audio program to be generated by inputting the serial number or name of the selected candidate topic information; or alternatively, selection can also be made through voice input, etc. The present disclosure does not make specific limitations in this regard.
[0061] In step S220, process the topic information through the first generation model to obtain one or more first audios.
[0062] Among them, the first generation model refers to the model used to generate the first audio, which can be an LLM model. The first audio refers to the basic audio for producing audio programs. For example, in a music interview program, the music involved in the interview, or the background music of the interview, etc. The first audio can be one or more. For example, in an interview program for a certain piece of music, one piece of background music can be used, or in a medley program, multiple pieces of music can be spliced together, and different music intervals can correspond to different explanatory contents.
[0063] In an exemplary embodiment, processing the theme information through the first generation model to obtain one or more first audios may include:
[0064] Constructing a second prompt word according to the theme information of the audio program to be generated;
[0065] Processing the second prompt word through the first generation model to obtain one or more first audios.
[0066] In this exemplary embodiment, a second prompt word can be constructed according to the theme information. The second prompt word can be a prompt word generated only according to the theme information. For example, the second prompt word can be "Please recommend songs with the theme of {topic}", where topic is the theme information; it can also be a prompt word generated according to the theme information and extended information. That is, when constructing the second prompt word according to the theme information, other information can be added to generate a more comprehensive second prompt word. Further, processing the second prompt word by the first generation model can obtain one or more first audios.
[0067] In an exemplary embodiment, the above audio program generation method may further include:
[0068] Obtaining audio creator information;
[0069] Then, constructing the second prompt word according to the theme information of the audio program to be generated may include:
[0070] Constructing a second prompt word according to the theme information of the audio program to be generated and the audio creator information.
[0071] Among them, the audio creator information can be the creator information of the first audio, including the creator name, creation time, creator profile, and other information related to the creator, etc. In this exemplary embodiment, a second prompt word can be constructed according to the theme information of the audio program and the audio creator information. For example, the second prompt word can be "Please recommend songs by the singer {artist} with the theme: {topic}", where artist is the audio creator information, such as the singer's name, and topic is the theme information.
[0072] In step S230, obtain the resource information corresponding to the first audio and the program style type information.
[0073] Among them, the resource information corresponding to the first audio refers to the resource data related to the first audio, such as audio data of music songs, singer names, popular cover version data, lyrics, music styles, creation backgrounds, award-winning situations, and other related information. The program style type information refers to the program style classification of the audio program to be generated, such as warm style type, popular science style type, lively style type, or steady style type, etc.
[0074] The resource information corresponding to the first audio can be retrieved by the server in the audio resource library after obtaining the relevant information of the first audio. For example, retrieve the corresponding audio data according to the name of the song, and the resource information can also be determined according to the retrieval result and user operations. For example, first retrieve the alternative resource information, and determine the resource information corresponding to the first audio according to the user's selection operation. The program style type can be determined by the system itself. For example, determine the corresponding style type according to the retrieved song; it can also be determined by the user. For example, the user can input the number, name, or description information of the program style type, so that the server can determine the corresponding program style type according to the information input by the user.
[0075] In step S240, process the resource information and the program style type information through the second generation model to obtain the program content text information.
[0076] Among them, the program content text information refers to the program description information of the audio program to be generated, which is the main information used to constitute the audio program. For example, in an interview program, the host can determine what to say in the interview according to the manuscript content. When the first audio content is multiple first audios, the corresponding generated program content text information can be the program content text information related to the multiple first audios. For example, retrieve multiple music of a certain singer according to a certain theme. For the medley audio of multiple music, the program content text information can be the explanation of the music appearing in the medley audio, or the psychological journey and related stories of the singer when creating different music.
[0077] In an exemplary embodiment, the above process of processing the resource information and the program style type information through the second generation model to obtain the program content text information may include:
[0078] Construct a third prompt word according to the resource information and the program style type information;
[0079] Process the third prompt word through the second generation model to obtain the program content text information.
[0080] The third prompt refers to the prompt constructed based on resource information and program style type information. In actual applications, the construction of the third prompt can also include other pre-configured information, such as the number of words in the text of the program content, the transition rules and transition words between different first audio, and so on. The following shows an example of the third prompt generated according to resource information and program style type information:
[0081] "My name is {{anchorName}}, and I am a music program host. I have a program called 《{{podcastName}}》. Today, I plan to prepare a program with the theme of 《{{title}}》. Please write an attractive script for me based on my theme and selected songs. The script needs to meet the following requirements:
[0082] 1. {{style}};
[0083] 2. At the beginning of the script, quickly introduce the program theme in about 50 words, and then start introducing each song. For each song, first write an introduction copy of about 150 words, and then write the song name. The order cannot be reversed. After all the songs are introduced, return to the theme and encourage the audience to comment and collect the program;
[0084] 3. The content of the song introduction needs to be based on facts and cannot be made up;
[0085] 4. The transition words between songs need to be professionally designed. You can design the transition words from the connection points (which can be theme, emotion, style, or artist) between the current song and the next song to be played. Do not use: Next, The next song;
[0086] 5. For the convenience of my recording, please mark the position where the song is inserted in the middle of the script. The marking format is "
Song Name - Singer Name
[0087] The following are the playlists I have prepared for this program:
[0088] {{songInfo}}”
[0089] Among them, anchorName is the name of the host, podcastName is the name of the audio program, and title is the name of the program. The specific content can be configured in the generation model or input by the user. Style is the information of the program style type, which can be a simple indication information, such as "warm style", or the specific information converted from the simple indication information input by the user, such as the specific definition of warmth converted according to the style type warmth; songInfo is the resource information corresponding to the first retrieved audio. It should be noted that in the above prompt words, the content other than the program style type information and the resource information can be pre-configured or modified according to actual needs.
[0090] In addition, in this exemplary embodiment, the specific rules of the program style type information can be pre-configured in the generation model. The user only needs to input the indication information of the style type, such as name, code or serial number, etc., to determine which specific style type information is required. Table 1 below shows an example of the mapping relationship between a style type and the specific description information configured in the generation model:
[0091]
[0092]
[0093] Figure 3 shows a schematic diagram of the framework process for generating program content text information, which may specifically include the following steps: After obtaining multiple first audios 310, the corresponding resource information 320 can be retrieved according to the first audios 310. The resource information may include creator information, audio style, lyrics, creation background, award-winning situation, etc.; then the program style type information 330 to be executed can be obtained; and then a third prompt word 340 can be constructed according to the resource information and the program style type information; the program content text information 350 can be obtained by processing the third prompt word through a second generation model.
[0094] In step S250, a target audio program is generated according to the first audio and the program content text information.
[0095] Finally, a target audio program can be generated according to the first audio and the program content text information. Specifically, the program content text information can be first converted into audio data, for example, by the user's recording operation to record the recording audio related to the program content text information, or by the model to convert the program content text information into audio data, etc. Then, the first audio can be synthesized with the audio corresponding to the program content text information to generate the target audio program.
[0096] In an exemplary embodiment, the above generation of the target audio program according to the first audio and the program content text information may include:
[0097] Obtain a second audio based on the program content text information;
[0098] Synthesize the first audio and the second audio to obtain a target audio program.
[0099] The second audio refers to audio data generated based on the program text information. For example, the program text information can be converted into audio data through a text-to-speech conversion model to obtain the second audio, or the recorder can perform voice recording on the program text information through a recording device to obtain the second audio, etc. After obtaining the second audio, the first audio and the second audio can be synthesized, such as fusing or splicing the first audio and the second audio to obtain a target audio program.
[0100] In an exemplary embodiment, the above-mentioned obtaining of the second audio based on the program content text information may include:
[0101] Obtain the timbre information of the target object;
[0102] Process the program content text information and the timbre information through a text-to-speech conversion model to obtain the second audio.
[0103] The timbre information refers to information that can reflect the voice characteristics of the target object. Different objects have different timbres. The target object can be any user object. For example, the program producer can be the target object. If a specific person is desired to appear in the audio program, that specific person can also be the target object. The method of obtaining the timbre information of the target object can be, for example, inputting the voice information of the target object so that the server can extract the corresponding timbre information from the voice information. For example, the target object can be an animated character, and by inputting the line voice of the animated character, the timbre information of the animated character can be extracted; in addition, the target object can be the program producer, and the program producer can randomly say a few sentences of voice, or read the corresponding voice according to a preset text template, so as to obtain the timbre information from the input voice, etc.
[0104] The text-to-speech conversion model (Text-to-Speech, TTS) refers to a model used to convert text information into voice information. In this exemplary embodiment, the program content text information and the timbre information can be input into the text-to-speech conversion model to obtain a second audio including the program content text information and having a specific timbre.
[0105] In an exemplary embodiment, the above-mentioned synthesizing the first audio and the second audio to obtain a target audio program may include:
[0106] Synthesize the first audio and the second audio, and gradually increase the volume of the first audio within the first time period starting from the beginning of the mixing time period of the first audio and the second audio;
[0107] And / or, synthesize the first audio and the second audio, and gradually decrease the volume of the first audio within a second time period at the end of the mixing time period of the first audio and the second audio.
[0108] When synthesizing the first audio and the second audio in this exemplary embodiment, in order to improve the listening effect of the audio program, mechanisms of "fade-in" and "fade-out" can be set.
[0109] Specifically, during the process of synthesizing the first audio and the second audio, within a first time period at the beginning of the mixing time period of the first audio and the second audio, gradually increase the volume of the first audio. Here, the first time period can be a preset time period when the first audio and the second audio start to be mixed, and the specific time range can be customized according to user needs. For example, it can be set to 3s (seconds), 5s, etc. The increasing rule of the volume of the first audio can also be adjusted according to user needs. For example, it can be linearly increased or non-linearly increased, etc. As Figure 4 , shows a schematic diagram of synthesizing the first audio and the second audio and setting the fade-in effect of the first audio. Among them, 410 is the first audio, such as a retrieved song segment, and 420 is the second audio, such as a recorded or model-generated voice segment of the anchor. The first audio 410 and the second audio 420 are synthesized to show a partially mixed effect in terms of time. Within the first time period 430 at the beginning of the mixing time period, the volume of the first audio 410 can be gradually increased.
[0110] During the process of synthesizing the first audio and the second audio, within a second time period at the end of the mixing time period of the first audio and the second audio, gradually decrease the volume of the first audio. Here, the second time period can be a preset time period near the end of the mixing of the first audio and the second audio, and the specific time range can be customized according to user needs. For example, it can be set to 3s, 5s, etc. The decreasing rule of the volume of the first audio can also be adjusted according to user needs. For example, it can be linearly decreased or non-linearly decreased, etc. As Figure 5 , shows a schematic diagram of synthesizing the first audio and the second audio and setting the fade-out effect of the first audio. Among them, 510 is the first audio and 520 is the second audio. The first audio 510 and the second audio 520 are synthesized to show a partially mixed effect in terms of time. Within the second time period 530 at the end of the mixing time period, the volume of the first audio 510 can be gradually increased.
[0111] It should be noted that the curves of increasing and decreasing the volume of the first audio, as well as the ranges of the first time period and the second time period, etc. can be set according to specific values or determined according to the user's interaction operations. For example, the user can change the shape of the curve by performing sliding operations such as pulling up or down on the fade-in curve or fade-out curve, thereby changing the rules of volume increase and decrease.
[0112] In an exemplary embodiment, the above-mentioned synthesizing the first audio and the second audio to obtain a target audio program may include:
[0113] Sequentially splice the first audio and the second audio, and gradually decrease the volume of the first audio within a third time period at the end of the first audio.
[0114] This exemplary embodiment may also synthesize the first audio and the second audio in a splicing manner, and during the synthesis process, gradually decrease the volume of the first audio within a third time period at the end of the first audio to achieve the "fade-out" effect of the first audio. Figure 6 Fig. shows another schematic diagram of setting the fade-out effect of the first audio, where 610 is the first audio and 620 is the second audio. Splice the first audio 610 and the second audio 620 in the order of the first audio first and the second audio second, and within a third time period 630 at the end of the first audio, the volume of the first audio 610 can be gradually increased.
[0115] In an exemplary embodiment, the above-mentioned audio program generation method may further include:
[0116] Construct a fourth prompt word according to the theme information, the first audio, and the program content text information;
[0117] Process the fourth prompt word through a fourth generation model to obtain the description information of the target audio program;
[0118] Publish the target audio program and the description information of the target audio program.
[0119] Among them, the description information of the target audio program can be used to summarize and briefly introduce the target audio program. This exemplary embodiment can construct a fourth prompt word based on the theme information, the first audio, and the program content text information, and use a fourth generation model to process the fourth prompt word to obtain the description information of the target audio program. Finally, publish the target audio program and the description information of the target audio program, and the audience can preliminarily understand the content of the target audio program through the description information of the target audio program.
[0120] It should be noted that the first generation model, the second generation model, the third generation model, and the fourth generation model can be different models. For example, corresponding models are used for processing in each stage. The first generation model, the second generation model, the third generation model, and the fourth generation model can also be the same model. For example, the second generation model, the third generation model, and the fourth generation model can be the same LLM model, etc.
[0121] Figure 7 The framework flowchart of an audio program generation method in this exemplary embodiment is shown, which may specifically include the following steps:
[0122] Step S702: Obtain the keywords input by the producer for the audio program to be generated, and construct a first prompt word according to the keywords of the audio program to be generated;
[0123] Step S704: Process the first prompt word through the third generation model to obtain multiple candidate theme information;
[0124] Step S706: In response to receiving the theme selection operation input by the user, determine the theme information of the audio program to be generated from the multiple candidate theme information;
[0125] Step S708: Construct a second prompt word according to the theme information of the audio program to be generated, and process the second prompt word through the first generation model to obtain one or more first audios;
[0126] Step S710: Process the resource information corresponding to the first audio and the program style type information through the second generation model to obtain program content text information;
[0127] Step S712: Based on the program content text information, the producer records the second audio through a recording device;
[0128] Step S714: Obtain the timbre information of the target object; and process the program content text information and the timbre information through a text-to-speech conversion model to obtain a second audio;
[0129] Step S716: Synthesize the first audio and the second audio to obtain the target audio program;
[0130] Step S718: Construct a fourth prompt word according to the theme information, the first audio, and the program content text information, and process the fourth prompt word through the fourth generation model to obtain the description information of the target audio program;
[0131] Step S720: Publish the target audio program and the description information of the target audio program.
[0132] Among them, step S712 and step S714 are two different ways to generate the second audio. The user can choose to use the self-recording method or the model generation method by themselves, and the present disclosure does not make specific limitations on this.
[0133] Exemplary Apparatus
[0134] An exemplary embodiment of the present disclosure also provides an audio program generation device. Refer to Figure 8 As shown, the audio program generation device 800 may include the following program modules: a theme information acquisition module 810 for acquiring the theme information of the audio program to be generated; a first audio acquisition module 820 for processing the theme information through a first generation model to obtain one or more first audios; an audio information acquisition module 830 for acquiring the resource information corresponding to the first audio and the program style type information; a text information acquisition module 840 for processing the resource information and the program style type information through a second generation model to obtain the program content text information; and an audio program generation module 850 for generating a target audio program according to the first audio and the program content text information.
[0135] In one embodiment, the theme information acquisition module 810 includes: a first prompt word construction unit for acquiring the keywords of the audio program to be generated and constructing a first prompt word according to the keywords of the audio program to be generated; and a third model processing unit for processing the first prompt word through a third generation model to obtain the theme information of the audio program to be generated.
[0136] In one embodiment, the third model processing unit includes: a candidate theme acquisition subunit for processing the first prompt word through a third generation model to obtain a plurality of candidate theme information; and a theme information selection subunit for determining the theme information of the audio program to be generated from the plurality of candidate theme information in response to receiving a user input theme selection operation.
[0137] In one embodiment, the first audio acquisition module 820 includes: a second prompt word construction unit for constructing a second prompt word according to the theme information of the audio program to be generated; and a first model processing unit for processing the second prompt word through a first generation model to obtain one or more first audios.
[0138] In one embodiment, the audio program generation device further includes: a creator information acquisition unit for acquiring audio creator information; and a second prompt word construction unit for constructing a second prompt word according to the theme information of the audio program to be generated and the audio creator information.
[0139] In one implementation, the text information acquisition module 840 includes: a third prompt word construction unit for constructing a third prompt word according to the resource information and the program style type information; and a second model processing unit for processing the third prompt word through a second generation model to obtain program content text information.
[0140] In one implementation, the audio program generation module 850 includes: a second audio acquisition unit for acquiring a second audio obtained based on the program content text information; and an audio synthesis unit for synthesizing the first audio and the second audio to obtain a target audio program.
[0141] In one implementation, the second audio acquisition unit includes: a timbre acquisition subunit for acquiring the timbre information of the target object; and a text-to-speech conversion subunit for processing the program content text information and the timbre information through a text-to-speech conversion model to obtain the second audio.
[0142] In one implementation, the audio synthesis unit includes: a first synthesis subunit for synthesizing the first audio and the second audio and gradually increasing the volume of the first audio within a first time period at the start of the mixing time period of the first audio and the second audio; and a second synthesis subunit for and / or synthesizing the first audio and the second audio and gradually decreasing the volume of the first audio within a second time period at the end of the mixing time period of the first audio and the second audio.
[0143] In one implementation, the audio synthesis unit includes: a third synthesis subunit for sequentially splicing the first audio and the second audio and gradually decreasing the volume of the first audio within a third time period at the end of the first audio.
[0144] In one implementation, the audio program generation device may further include: a fourth prompt word construction unit for constructing a fourth prompt word according to the theme information, the first audio, and the program content text information; a fourth model processing unit for processing the fourth prompt word through a fourth generation model to obtain the description information of the target audio program; and an information publishing unit for publishing the target audio program and the description information of the target audio program.
[0145] In addition, the other specific details of the embodiments of the present disclosure have been described in detail in the embodiments of the above method and will not be elaborated here.
[0146] Exemplary Storage Medium
[0147] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above methods of the present disclosure are implemented. The above methods can be implemented by a program product, such as a portable compact disc read-only memory (CD-ROM) including program code and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0148] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0149] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0150] The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0151] Program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0152] Exemplary Electronic Device
[0153] Exemplary embodiments of the present disclosure also provide an electronic device, which can be any device among Figure 1 The electronic device includes a processor and a memory, and the memory is used to store executable instructions of the processor. The processor is configured to execute the above-mentioned method of the present disclosure by executing the executable instructions.
[0154] Refer to Figure 9 to describe the electronic device of the exemplary embodiments of the present disclosure. Figure 9 The displayed electronic device 900 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0155] As Figure 9 shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one processing unit 910, at least one storage unit 920, and a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910).
[0156] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 910 can execute method steps as Figure 2 shown, etc.
[0157] The storage unit 920 may include a volatile storage unit, such as a random access storage unit (RAM) 921 and / or a cache storage unit 922, and may further include a read-only storage unit (ROM) 923.
[0158] The storage unit 920 may also include a program / utilities 924 having a set (at least one) of program modules 925, such program modules 925 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of these examples or some combination thereof may include an implementation of a network environment.
[0159] The bus 930 may include a data bus, an address bus, and a control bus.
[0160] The electronic device 900 may also communicate with one or more external devices 1000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and such communication may be carried out through an input / output (I / O) interface 940. The electronic device 900 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 950. As shown in the figure, the network adapter 950 communicates with other modules of the electronic device 900 through the bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0161] It should be noted that although several modules or sub-modules of the apparatus are mentioned in the above detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units / modules may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.
[0162] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0163] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. Such a division is only for the convenience of description. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. An audio program generation method, characterized in that, It includes: Obtain the theme information of the audio program to be generated; Process the theme information through a first generation model to obtain one or more first audios; Obtain the resource information corresponding to the first audio and the program style type information; Process the resource information and the program style type information through a second generation model to obtain program content text information; Generate a target audio program according to the first audio and the program content text information.
2. The method according to claim 1, characterized in that The obtaining the theme information of the audio program to be generated includes: Obtain the keywords of the audio program to be generated, and construct a first prompt word according to the keywords of the audio program to be generated; Process the first prompt word through a third generation model to obtain the theme information of the audio program to be generated.
3. The method according to claim 2, wherein The processing the first prompt word through a third generation model to obtain the theme information of the audio program to be generated includes: Process the first prompt word through a third generation model to obtain multiple candidate theme information; In response to receiving a theme selection operation input by the user, determine the theme information of the audio program to be generated from the multiple candidate theme information.
4. The method according to claim 1, characterized in that, The processing the theme information through a first generation model to obtain one or more first audios includes: Construct a second prompt word according to the theme information of the audio program to be generated; Process the second prompt word through the first generation model to obtain one or more first audios.
5. The method according to claim 4, wherein The method further includes: Obtain audio creator information; The constructing the second prompt word according to the theme information of the audio program to be generated includes: Construct the second prompt word according to the theme information of the audio program to be generated and the audio creator information.
6. The method according to claim 1, wherein The processing the resource information and the program style type information through a second generation model to obtain program content text information includes: Construct a third prompt word according to the resource information and the program style type information; Process the third prompt word through the second generation model to obtain the program content text information.
7. The method according to claim 1, wherein The generating a target audio program according to the first audio and the program content text information includes: Obtain a second audio obtained based on the program content text information; Synthesize the first audio and the second audio to obtain the target audio program.
8. An audio program generation device, characterized in that, It includes: A theme information acquisition module, configured to obtain the theme information of the audio program to be generated; A first audio acquisition module, configured to process the theme information through a first generation model to obtain one or more first audios; An audio information acquisition module, configured to obtain the resource information corresponding to the first audio and the program style type information; A text information acquisition module, configured to process the resource information and the program style type information through a second generation model to obtain program content text information; An audio program generation module, configured to generate a target audio program according to the first audio and the program content text information.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 7 by executing the executable instructions.