Generation method and device of music broadcasting program, equipment, medium and program product
By acquiring song lists and broadcast style information, generating podcast transcripts using a large language model, and synthesizing speech and audio, the problem of low production efficiency and content homogenization in music podcast programs is solved, achieving efficient and personalized content production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-05
AI Technical Summary
Existing music podcast production relies on manual operation, resulting in low production efficiency and a lack of variety and homogeneity in content style, making it difficult to meet the demand for high-frequency content updates.
By acquiring a list of target songs and descriptions of broadcast style, a large language model is used to generate podcast transcripts, which are then converted into audio and synthesized with the song audio to achieve an automated production process.
It improves the production efficiency of music podcast programs, ensures the consistency of content logic and listening style, and avoids the rigidity and homogenization of content caused by traditional template-based generation methods.
Smart Images

Figure CN121983008A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and specifically to methods, apparatus, devices, media, and program products for generating music podcast programs. Background Technology
[0002] With the rapid development of audio platforms, music podcasts have become an important channel for users to access music content. Currently, music podcast production mainly relies on a manual operation model, where operators manually select songs, write scripts, record vocals, and perform post-production audio editing and splicing. This approach not only results in long production cycles and high labor costs due to tedious manual operations, making it difficult to meet the high-frequency content update demands, but also, while some automated content production solutions exist, most are limited to fill-in-the-blank patterns based on fixed sentence templates. For example, they simply embed metadata such as song titles and artists into preset recommendation texts and song information slots, or they can only generate short reviews for individual songs using simple rules. They cannot automatically generate complete and logically coherent long podcast scripts that include opening remarks, natural transitions between songs, and an ending. This leads to highly homogenized program content, lacking the unique personal style and emotional depth of the host, and failing to meet listeners' needs for more diverse and richer content.
[0003] Therefore, there is an urgent need for a method to generate music podcast programs in order to solve the problems of low production efficiency, monotonous content style, and serious homogenization in related technologies. Summary of the Invention
[0004] This disclosure provides a method, apparatus, device, medium, and program product for generating music podcast programs, in order to solve the problems of low production efficiency and monotonous content style and serious homogenization in related technologies.
[0005] In a first aspect, this disclosure provides a method for generating music podcast programs, the method comprising:
[0006] Obtain the target song list and the description information of the broadcast style used to broadcast the target song list; Based on the song information and broadcast style description information in the target song list, a podcast script is generated; the podcast script contains song lyrics with the style corresponding to the broadcast style description information. The podcast transcript is converted into audio, and then the audio is combined with the audio of the songs in the target song list to create a music podcast program.
[0007] In one alternative implementation, obtaining the target song list includes: Based on the broadcast style description information, candidate songs that match the broadcast style description information are selected from the preset music database; The candidate songs are filtered according to the preset filtering rules to obtain the target song list.
[0008] In one alternative implementation, after obtaining the target song list, the method further includes: Query whether a music podcast containing candidate songs has been generated based on the broadcast style description information within a preset historical time period; If so, the candidate song will be removed from the target song list.
[0009] In one optional implementation, a podcast transcript is generated based on song information from the target song list and broadcast style description information, including: Construct input instructions, which include broadcast style description information, text format constraints, and metadata for each song in the target song list. The metadata includes at least the song name and the artist name. Input the commands into the text generation model to obtain the podcast transcript.
[0010] In one alternative implementation, the input instructions are fed into a text generation model to obtain a podcast transcript, including: Encapsulate the metadata of all songs in the target song list into a single input command; A podcast transcript containing the lyrics of all songs in the target song list is generated using a text generation model.
[0011] In one optional implementation, the speech audio is synthesized with the song audio corresponding to the target song list, including: The audio is synthesized sequentially according to the playback order of the songs in the target song list, and audio transition processing is performed on at least one song audio to achieve a smooth transition between the speech audio and the song audio.
[0012] In one optional implementation, temporal synthesis is performed according to the playback order of songs in the target song list, and audio transition processing is performed on at least one song audio, including: Determine the entry and exit times for each song's audio based on the playback order of the songs in the target song list. Apply volume gradation processing to the song audio during the entry or exit period to create a volume increase or decrease effect. At least a portion of the speech audio is superimposed onto the time period corresponding to the volume gradation feature, and the speech audio is distributed only within the range of the entry time period and / or exit time period.
[0013] In one alternative implementation, the method further includes: Based on podcast transcripts and broadcast style descriptions, generate at least one candidate program title; Select the target title from the candidate program titles; Upload the music podcast program, target title, and cover image associated with the broadcast style description to the content publishing platform.
[0014] In one alternative implementation, after generating the podcast transcript and before converting the podcast transcript into audio, the method further includes: In response to modification or regeneration instructions for the podcast transcript, adjust the broadcast style description or transcript format constraints in the input instructions, and re-execute the steps to generate the podcast transcript.
[0015] Secondly, this disclosure provides an apparatus for generating music podcast programs, the apparatus comprising: The acquisition module is used to acquire the target song list and the playback style description information used to play the target song list; The generation module is used to generate podcast transcripts based on song information and broadcast style description information in the target song list; the podcast transcripts contain song lyrics with the style corresponding to the broadcast style description information. The synthesis module is used to convert podcast transcripts into audio, and then synthesize the audio with the corresponding song audio from the target song list to obtain a music podcast program.
[0016] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the method for generating a music podcast program as described in the first aspect or any corresponding embodiment.
[0017] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method for generating a music podcast program according to the first aspect or any corresponding embodiment described above.
[0018] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to execute a method for generating a music podcast program according to the first aspect or any corresponding embodiment described above.
[0019] The method for generating music podcast programs provided in this disclosure obtains a list of target songs and broadcast style description information for defining broadcast characteristics, and generates a podcast script containing song lyrics with corresponding styles based on these two. This transforms standardized song information into personalized narration content with specific style tendencies, thereby avoiding the content rigidity and homogenization problems caused by traditional template-based generation methods. Furthermore, by converting the generated script into audio and synthesizing it with song audio, this method realizes an automated production process from text creation to audio finished product. While significantly improving the efficiency of music podcast program production, it ensures the consistency of the final generated program in terms of content logic and listening style. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this disclosure; Figure 2 This is a schematic flowchart of a first method for generating a music podcast program according to an embodiment of the present disclosure; Figure 3 This is a second flowchart illustrating a method for generating a music podcast program according to an embodiment of the present disclosure; Figure 4 This is a flowchart illustrating an automated production method for a music podcast recommendation category according to an embodiment of this disclosure. Figure 5 This is a structural block diagram of a music podcast program generation apparatus according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0023] It should be noted that the information (including but not limited to user input information, such as information entered into input boxes), data (including but not limited to data used for analysis, stored data, and displayed data, such as context code, all code of the current project, service pressure corresponding to operations performed on all code of the current project, and code development status of the current project), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. For example, the context code, operations performed on all code of the current project, the corresponding service pressure, and code development status involved in this disclosure were all obtained with full authorization.
[0024] Before providing a detailed description of the embodiments of this disclosure, some of the nouns and terms involved in the embodiments of this disclosure will be explained.
[0025] (1) Recommendation-type music podcasts: an audio program format centered on "recommendation and sharing". Its typical structure is "the host's voice introduction + song playback", aiming to guide users to listen to and collect songs by introducing the background of the songs and sharing their listening experience.
[0026] (2) Virtual Human Account: refers to the automated operation account created. Each account is not just an ID, but is bound to a unique TTS tone (such as a gentle female voice, anime voice, etc.), a specific music category (such as rock, ancient style) and a detailed personality style (such as a sharp music critic or a girl next door).
[0027] (3) AI script generation: refers to the use of large language model (LLM) to automatically generate a complete program script containing the opening remarks, the lyrics of each song and the closing remarks in one go through preset prompts and dynamically injected song metadata.
[0028] (4) Prompt: refers to the instruction text input to the large language model.
[0029] (5) TTS (Text-To-Speech): A technology that converts text files into audio.
[0030] (6) Gradual volume increase during song entry / exit: an audio signal processing technique that refers to gradually increasing the volume from zero when a song starts playing (entry) and gradually decreasing the volume to zero when the song ends or switches (exit), in order to achieve a smooth transition between audio segments.
[0031] (7) Program synthesis: refers to the process of splicing and processing the human voice segments generated by TTS and the selected song audio files according to the timeline logic to generate the final playable audio file.
[0032] As one optional application scenario of this disclosure embodiment, such as Figure 1 As shown, the method for generating music podcast programs provided in this disclosure can be applied to an audio processing system, which may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.
[0033] The terminal device can be a smartphone, tablet, laptop, PDA, or desktop computer, etc. The terminal device can be used to send generation instructions to server 103, configure playback style description information, or upload a list of target songs. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranets, local area networks, wide area networks, mobile communication networks, and combinations thereof.
[0034] In related technologies, the production of music podcasts mainly relies on manual processes such as song selection, scriptwriting, recording, and post-editing, or on simple templates for mechanical filling. This approach not only suffers from low production efficiency and high labor costs due to its deep interface hierarchy and cumbersome steps, but also struggles to quickly produce program content with a consistent style and personalized characteristics when faced with a massive amount of songs, resulting in a monotonous listening experience for users. This disclosure provides a method for generating music podcasts that can efficiently and in batches produce music podcasts with distinct personalities and high-quality listening experiences.
[0035] According to an embodiment of this disclosure, a method for generating a music podcast program is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0036] This embodiment provides a method for generating music podcast programs, which can be used in the aforementioned terminal devices, such as laptops and desktop computers. Figure 2 This is a flowchart of a method for generating a music podcast program according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the target song list and the broadcast style description information used to broadcast the target song list.
[0037] The target song list refers to a set of songs that the podcast plans to play or recommend in this episode. This list can contain identifiers for one or more songs (such as song IDs, song titles, etc.). For example, the target song list could be "This Week's Top 5 Popular Songs," which includes five songs such as "Song A," "Song B," "Song C," etc.
[0038] Broadcast style description information refers to parameters or descriptive data used to define the specific characteristics presented by the "narration" in a podcast program. It covers all the configuration information that can determine the tone of the text, word choice, emotional color, and voice timbre.
[0039] For example, the broadcast style description can be a combination of tags, such as "Style: lighthearted and humorous + Voice: energetic male voice"; or it can be a natural language description, such as "Analyzing songs seriously and in-depth in the tone of a senior music critic".
[0040] Of course, in one optional application scenario, this can also correspond to the persona information of a "virtual person" with specific characteristics (such as name, age, and background). A virtual person refers to a digital image or account character generated by a computer with specific identity settings and personality traits, usually given the identity of a host simulating a human. For example, a rock-style virtual host named "Virtual Anchor A," or a healing-style virtual anchor named "Virtual Anchor B." Persona information refers to the data set that defines the characteristics of a virtual person, including but not limited to personality traits, language style, tone of voice, background settings, and preferred areas. For example, for "Virtual Anchor A," its persona information might include "outgoing personality, fast speaking speed, likes to use trendy slang, and has unique insights into rock music"; for "Virtual Anchor B," its persona information might include "gentle voice, slow speaking speed, likes to quote poetry, and has delicate emotions."
[0041] The action of "acquiring" refers to the process of reading or receiving data from internal storage, external databases, user input, or third-party interfaces. For example, a user selects the "Happy Mood" mode on a music software interface (i.e., acquires the corresponding style description information) and confirms 5 recommended songs (i.e., acquires the target song list).
[0042] This step is the preparation stage for generating the program, which clarifies "what to play" (song list) and "performance style" (style description information).
[0043] In practical applications, the methods for obtaining broadcast style description information are extremely flexible. In one implementation, it can be statically bound, for example, a "rock channel" is pre-defined to correspond to "passionate style"; in another implementation, it can be dynamically matched, for example, first analyzing the overall beat count and emotional information of the target song list, and if the songs are generally sad, then generating or retrieving matching "sad and healing" broadcast style description information.
[0044] Step S202: Generate podcast transcript based on the song information and broadcast style description information in the target song list; the podcast transcript contains song lyrics with the style corresponding to the broadcast style description information.
[0045] Song information refers to the specific attribute data of each song in the target song list. In addition to the song title and artist, it can also include the meaning of the lyrics, the background of the creation, the genre, the year of release, related stories, etc.
[0046] For example, the song information for "Song A" might include "The singer is a famous rock band, expressing a yearning for freedom."
[0047] A podcast script is the text content used to generate podcast audio, also known as a "script." It typically includes an opening, song introductions, transitions between songs, and a closing.
[0048] Song introductions refer to the language used to introduce, evaluate, or elaborate on a song before or after it is played.
[0049] Corresponding style refers to the consistency between the rhetoric, tone, word choice, and broadcast style description information in the text.
[0050] For example, if the style is "humorous," the script will contain jokes or banter; if the style is "professional," the script will contain music theory knowledge or industry background.
[0051] The action of "generation" refers to the process of creating text based on input data using natural language processing techniques (such as large language models).
[0052] For example, inputting "information about Song A" and "humor style" as prompts into a pre-defined large language model will output a text.
[0053] This step aims to transform objective song information into readable text with subjective coloring and specific personal charm.
[0054] This step can be implemented in various ways. It could be a simple template fill, such as filling the song title into a preset stylized template; more preferably, it can utilize generative artificial intelligence to allow the model to "understand" the style description information and engage in creative writing. For example, when the style description information is "history teacher," the model could explore the historical background of the song's release year and incorporate it into the text.
[0055] For example, assuming the broadcast style description is "sentimental" or "healing," the generated script for a song like "Song B" might be: "In every sleepless night, perhaps only this song, 'Song B,' can understand the regret in your heart. May this song soothe your pain." This script carries a comforting and empathetic tone, fitting the "sentimental" style requirement. In this way, the generated podcast script is no longer a mechanical announcement but possesses distinct personalized characteristics, achieving a unique content experience for each individual.
[0056] Step S203: Convert the podcast transcript into audio, and synthesize the audio with the audio of the songs corresponding to the target song list to obtain a music podcast program.
[0057] Voice audio refers to the waveform data or audio file obtained by converting the text-based podcast transcript, i.e., the "voice" part. For example, a 30-second WAV file contains the host reading the script generated in the previous step.
[0058] Song audio refers to the actual music file corresponding to the song in the target song list.
[0059] A music podcast is a final, complete streaming audio file that users can listen to. It consists of two parts: narration and music playback.
[0060] The "conversion" action refers to using text-to-speech (TTS) technology to turn text into sound. Typically, a specific voice tone is selected based on the broadcast style (such as a deep male voice or a sweet child's voice). For example, a TTS engine might be invoked, using a "rock male voice" voice library to read aloud the text generated in step S202 and save it as audio.
[0061] The act of “synthesis” refers to mixing, splicing, or editing vocal tracks and music tracks to create a continuous audio stream.
[0062] For example, multiple audio segments can be connected in the order of "opening audio, first song audio, first song interlude audio, second song audio"; or human voice audio can be superimposed during a specific time period of song playback (such as the intro or outro).
[0063] This step is the output stage of the finished product. The text is converted into audio, and the vocals and singing are blended to form a complete radio program, which users can listen to or publish directly.
[0064] In summary, the music podcast generation method provided in this disclosure, by acquiring a list of target songs and broadcast style description information used to define broadcast characteristics, and generating a podcast script containing song lyrics with corresponding styles based on these two, can transform standardized song information into personalized narration content with specific style tendencies, thereby avoiding the content rigidity and homogenization problems caused by traditional template-based generation methods. Furthermore, by converting the generated script into audio and synthesizing it with song audio, this method realizes an automated production process from text creation to audio finished product, significantly improving the efficiency of music podcast production while ensuring the consistency of the final generated program in terms of content logic and listening style.
[0065] This embodiment provides a method for generating music podcast programs, which can be used in the aforementioned terminal devices, such as laptops and desktop computers. Figure 3 This is a flowchart of a method for generating a music podcast program according to an embodiment of the present disclosure, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain the target song list and the broadcast style description information used to broadcast the target song list.
[0066] This step is the preparation stage for generating the program, which mainly involves determining "what to play" (songs) and "how to play" (style).
[0067] Based on this, in order to achieve accurate content matching, this step specifically includes: Based on the broadcast style description information, candidate songs that match the broadcast style description information are selected from the preset music database; and the candidate songs are filtered according to the preset filtering rules to obtain the target song list.
[0068] Broadcast style description information is a set of data used to define the overall tone, emotional color, or anthropomorphic characteristics of a podcast. For example, it can be a combination of tags (such as late at night, healing, deep male voice) or a natural language description (such as an artsy young man taking a walk on the beach, speaking in a relaxed tone, and liking jazz).
[0069] A pre-defined music database refers to a database or music library that stores a large number of song files and their attribute information.
[0070] Candidate songs refer to a set of alternative songs that initially meet the style requirements but have not yet been finalized for selection.
[0071] The preset filtering rules refer to the logical conditions used to further select songs in addition to style matching, such as "Top 50 most played songs in the past week", "released within the past year", or "more than 1,000 comments".
[0072] Filtering refers to the process of selecting data that meets specific criteria from a massive dataset. For example, if the style description information includes "traditional Chinese style," then songs of the "hip-hop" and "rock" types will be automatically filtered out from the database, leaving only "traditional Chinese style" songs as candidates.
[0073] In practice, the filtering process here goes beyond simple tag matching. In a preferred implementation, if the style description is in natural language (e.g., "songs suitable for listening to while drinking coffee on a rainy day"), semantic analysis techniques can be used to map it into the multi-dimensional feature vector space of a music database, thereby finding songs that match the atmosphere, rather than relying solely on the explicit "rainy day" tag. Furthermore, pre-defined filtering rules can include hard indicators such as "recent popularity" and "user ratings" to ensure that the selected songs are not only stylistically correct but also of high quality.
[0074] This step ensures that the selected songs are highly consistent with the broadcast style description, avoiding the incongruity caused by style conflicts (such as broadcasting children's songs in a rock style), while using the filtering rules to ensure the quality and popularity of the recommended songs.
[0075] Furthermore, after obtaining the target song list, the process also includes deduplication: Check if a music podcast containing candidate songs has been generated based on the broadcast style description information within a preset historical time period; if so, remove the candidate songs from the target song list.
[0076] The preset historical time period refers to the time window for tracing past records, such as "the past 30 days" or "the past 24 hours".
[0077] The query and removal function refers to checking the historical records. If it is found that "Song A" was played in this style mode two weeks ago, it will be automatically removed and replaced with the next song.
[0078] For example, suppose the initial list contains "Song A", but a search reveals that "Song A" was already recommended in a program three days ago in this broadcast style. In order to keep the content fresh, "Song A" will be removed from the current list and replaced with the next song that meets the criteria.
[0079] This mechanism avoids repeatedly recommending the same song with the same broadcast style, ensuring the frequency of podcast content updates and keeping the content fresh for listeners.
[0080] Step S302: Generate a podcast script based on the song information and broadcast style description information in the target song list; the podcast script contains song lyrics with the style corresponding to the broadcast style description information.
[0081] This step aims to generate a personalized script that matches the broadcasting style.
[0082] Specifically, the process of generating podcast transcripts includes: constructing input instructions, which contain broadcast style description information, transcript format constraints, and metadata for each song in the target song list, with the metadata including at least the song name and artist name; inputting the input instructions into the text generation model to obtain the podcast transcript.
[0083] Input instructions, usually referred to as prompts, are text instructions that guide artificial intelligence models in generating content.
[0084] Metadata refers to data that describes the basic attributes of a song, such as song title, artist, album name, and release year.
[0085] Text generation models refer to artificial intelligence models trained using deep learning techniques (such as the Transformer architecture) that possess the ability to understand and generate natural language. Specifically, these can be general-purpose large language models (LLMs, such as GPT) or specialized writing models fine-tuned for specific vertical domains. These models can understand complex Prompt instructions and simulate specific human language styles for writing.
[0086] The action of "construction" refers to assembling scattered information (character design, rules, playlists) into a complete instruction text according to a specific template.
[0087] For example, a Prompt could be created by combining the broadcast style "sharp-tongued music critic", the format requirement "each paragraph no more than 100 words", and the song data "Song B and Singer C": "You are a sharp-tongued music critic. Please use no more than 100 words to evaluate singer C's Song B..."
[0088] By using structured input instructions, the capabilities of large language models can be fully utilized to produce high-quality manuscripts that meet specific requirements in a batch and in a standardized manner, greatly reducing the cost of manual writing.
[0089] In terms of specific generation strategy, this embodiment adopts a holistic generation approach: The metadata of all songs in the target song list is encapsulated into a single input instruction; a podcast transcript containing the strings of words from all songs in the target song list is generated using a text generation model.
[0090] Encapsulation refers to packaging multiple data items together. In this context, it means putting the information for the 5 or 10 songs to be played in this episode into the same Prompt file all at once.
[0091] Instead of calling the text generation model separately for each song, the system sends the instruction: "Please generate a complete podcast transcript for the following 5 songs: 1. Song D, 2. Song E...". The large language model then outputs a complete text containing introductions to all the songs at once.
[0092] This holistic generation method not only improves processing efficiency, but more importantly, it allows the model to coordinate the context, making the transitions between songs more natural and coherent, and avoiding the fragmented feeling of fragmented generation.
[0093] The specific content and format of the manuscript are as follows: The podcast transcript includes an opening, song lyrics for each song arranged in the order of the target song list, and an ending; transcript format constraints limit the word count range for the opening, song lyrics, and ending, and specify that the podcast transcript is written in the first person and includes a contextual description of the songs.
[0094] Contextual description refers to creating a specific image or atmosphere through text, rather than a dry introduction.
[0095] The first-person perspective uses pronouns such as "I" and "we" to simulate the tone of a direct conversation between the anchor and the audience.
[0096] The script was no longer a simple announcement like "Next up is 'Song F'", but something like: "(Opening) Hello everyone, I'm your late-night DJ. (Transition) It's raining outside, and this is the perfect time to listen to 'Song F'. When the singer's hoarse voice came on, it was as if I was taken back to that damp afternoon..."
[0097] The standardized structure ensures the integrity of the program, while the first-person perspective and scene-based requirements give the script a strong emotional color and a sense of immersion, enhancing the appeal of the content.
[0098] In addition, to ensure the quality of the manuscript, this embodiment also introduces a manual intervention mechanism: After generating the podcast transcript and before converting it into audio, in response to modification or regeneration instructions for the podcast transcript, adjust the broadcast style description information or transcript format constraints in the input instructions, and re-execute the steps to generate the podcast transcript.
[0099] Adjustment refers to modifying parameters in the Prompt. For example, if a user feels the generated document is too long, they can change the constraint to "reduce word count by 20%".
[0100] If a user reviewer finds the generated text to be too serious, they can modify the broadcast style description in the Prompt and call the text generation model again until the text is satisfactory. This allows for effective control of content quality and correction of potential deviations within an automated production process.
[0101] Step S303: Convert the podcast transcript into audio, and synthesize the audio with the audio of the songs corresponding to the target song list to obtain a music podcast program.
[0102] This step is the synthesis stage that transforms the text into the final listenable program.
[0103] In this step, the audio synthesis is performed as follows: The audio is synthesized sequentially according to the playback order of the songs in the target song list, and audio transition processing is performed on at least one song audio to achieve a smooth transition between the speech audio and the song audio.
[0104] Timing composition arranges footage according to a timeline. For example, 0:00 to 0:30 is the opening, and 0:30 to 3:30 is the first song.
[0105] Audio transition processing refers to technical processing performed at the connection point between two audio segments, such as mixing and dissolve. It can avoid abrupt song cuts and improve the listening experience of the program.
[0106] The specific transition processing method is as follows: determine the entry time period and exit time period of each song audio according to the playback order of the songs in the target song list; perform volume gradation processing on the song audio during the entry time period or exit time period to create the effect of volume gradually increasing or decreasing; superimpose at least a part of the voice audio onto the time period corresponding to the volume gradation feature, and make the voice audio only distributed within the range of the entry time period and / or exit time period.
[0107] The entry and exit time periods usually refer to the intro and outro of a song, or artificially set segments at the beginning and end of a song (such as the first 15 seconds and the last 15 seconds).
[0108] Volume gradient processing includes gradually increasing and decreasing volume.
[0109] Overlay refers to the mixing operation that combines vocal tracks and music tracks on the same timeline.
[0110] The vocals are only distributed within the entry and / or exit time periods, meaning that the vocals will not cover the main body of the song (i.e., the vocal performance part), thus avoiding vocal conflict.
[0111] For example, if the intro to "Song G" is identified as 10 seconds long, then that 10 seconds is set as the entry period, and the volume is gradually increased from 0 to 100%. During this time, the intro audio (8 seconds long) is placed within this 10-second interval. When the song enters the verse, the vocals stop; as the song nears its end (exit period), the volume gradually decreases, and the closing remarks or the introduction to the next song then begin.
[0112] This synthesis method cleverly utilizes the beginning and end of the song as background music, eliminating the need to find additional background music, achieving a smooth transition in the program flow, and enhancing the listening experience.
[0113] Step S304: Post-processing and release of the program.
[0114] After generating the complete audio file, in order to complete content distribution, this embodiment also includes: generating at least one candidate program title based on the podcast transcript and broadcast style description information; selecting a target title from the candidate program titles; and uploading the music podcast program, the target title, and the cover image associated with the broadcast style description information to the content publishing platform.
[0115] The candidate program titles are multiple title options automatically extracted by an artificial intelligence model based on the content of the script.
[0116] The cover image associated with the broadcast style description information is a pre-configured image that matches the style (e.g., a cover image of a guitar corresponding to "rock style").
[0117] Content distribution platforms refer to terminal channels such as music streaming platforms.
[0118] For example, after analyzing the manuscript, the AI model provides three titles: "Rainy Day Playlist," "Healing Moments," and "Hearing the Sound of Rain." Selecting "Hearing the Sound of Rain" as the target title generates a corresponding image for the cover, and this information, along with the generated audio file, can be published to a music streaming platform with a single click. This eliminates the need for tedious manual form filling and file uploading in the backend, significantly improving publishing efficiency.
[0119] In summary, the music podcast generation method provided in this disclosure firstly ensures the consistency and freshness of the program content's style by selecting and deduplicating songs based on broadcast style description information, avoiding character collapse or content repetition caused by inappropriate song selection. Secondly, by constructing input instructions containing metadata and style commands, combined with the one-time interactive capability of the text generation model, it is possible to generate logically coherent, emotionally rich, and specific character-style long podcast scripts, significantly improving the anthropomorphism of the content and production efficiency. Thirdly, through refined audio transition processing and temporal synthesis technology, smooth transition effects are achieved using the song's own entrance / exit segments, significantly improving the user's listening experience without the need for additional background music resources. Finally, by combining a human feedback mechanism with an automated publishing process, it ensures that content quality can still be flexibly controlled while automating production, balancing efficiency and controllability.
[0120] To better illustrate the method for generating music podcast programs according to embodiments of this disclosure, a preferred embodiment will be provided below. This embodiment is intended to detail the implementation process of this disclosure, but is not intended to limit the scope of protection of this disclosure.
[0121] In the field of music podcast content production, especially for programs focused on song recommendations or product seeding, the industry primarily relies on manual operation or simple template-based production. While the manual approach yields high-quality content, the entire process of selection, scriptwriting, and recording is extremely costly and inefficient, resulting in programs only being released after the popularity of a hit song has faded. Furthermore, related automated technologies are mostly based on fixed templates (such as song title + recommendation), leading to highly homogenized programs lacking personalized character styles. To address the aforementioned problems of high cost, low efficiency, and style uniformity, this embodiment proposes an automated production method for product seeding music podcasts.
[0122] The process of the automated production method for music podcasts related to product recommendations in this embodiment is as follows: Figure 4 As shown, the process includes the following: Step 1: Construction and initialization of the virtual human account system.
[0123] To address the issue of a monotonous content style, several virtual accounts were first created, and differentiated personas were assigned to them.
[0124] Virtual characters are categorized according to their musical style. For example, virtual character A is created and set to a "light and cheerful" style, specializing in the "popular hits" category; virtual character B is created and set to an "electronic rock" style, specializing in the "trendy electronic music" category; virtual character C is created and set to a "sad and healing" style, specializing in the "love songs" category; as well as category D, which covers anime, comics, and games (ACG), and category E, which covers film and television OSTs, etc.
[0125] Each virtual character is matched with a unique TTS voice library, such as matching a deep and gentle voice to a sad and healing account, and a high-pitched and energetic voice to an anime account.
[0126] Upload a fixed program cover image to each account and associate it with the publishing channel.
[0127] Step two: Intelligent song combination of popular songs.
[0128] Based on a pre-defined song selection logic, automated song selection is performed periodically (e.g., every 24 hours) to ensure that the content keeps up with current trends. The specific process is as follows: Obtain popular playlists such as the Top 500 on the platform, and then distribute and map them to virtual human accounts of the corresponding style based on the playlist's genre tags (such as "rock" and "folk").
[0129] Based on the category positioning of virtual humans, songs are finely filtered according to dimensions such as genre matching, artist popularity, and ranking of plays in the past 28 days.
[0130] To ensure the freshness of recommended content, a deduplication logic is implemented. If a virtual account has recommended "Song A" within a certain period of time (e.g., 30 days), that song will be removed, but this does not affect different virtual accounts.
[0131] Ultimately, each virtual character will receive a list of 5 to 10 songs to be recommended.
[0132] Step 3: Dynamically generate text based on the character design.
[0133] After obtaining the song list, a large AI model is used to generate a complete script that matches the virtual human's persona based on the selected song list.
[0134] First, a unique prompt is configured for each virtual human. This template contains three parts: Static character description: For example, "You are a magical conch shell living on the beach, and you speak with humor and philosophical insights."
[0135] Dynamic data injection: The metadata (song title, artist, and optional description or creative background) of 5 to 10 selected songs is embedded into the Prompt.
[0136] Formatting constraints: For example, the output should include "opening remarks (30 to 50 words)," "song lyrics in sequence (50 to 80 words per song)," and "closing remarks," and the content should include an introduction to the background or highlights of the songs.
[0137] The prepared prompt is sent to the AI model. Through a one-time interaction, the AI outputs a complete podcast script that is logically coherent and in a tone consistent with the speaker's persona. The script includes not only introductions to the songs but also natural transitions between them.
[0138] Users can review the AI-generated text. If the text is not lively enough, users can adjust the parameters in the Prompt (such as "increase liveliness") to regenerate the text, but keep the playlist unchanged.
[0139] After approval, the AI model is called again to generate 3 to 5 attractive alternative titles based on the content of the text and the user's personality style, for the user to confirm.
[0140] Step 4: TTS audio synthesis and splicing with padless music.
[0141] Use the proprietary voice to convert the text into a high-quality human voice file.
[0142] This embodiment executes specific audio splicing logic. Unlike traditional radio stations that use background music, this embodiment uses serial splicing: first, a short introductory vocal segment is played, and after the vocal segment ends, the full song file is seamlessly connected.
[0143] To achieve a natural sound, crescendo and diminuendo effects are applied during the song's entrance and exit phases. For example, the next vocal introduction can be inserted the instant the volume of the previous song fades out, or the next song can gradually increase in volume at the end of the vocal introduction. This technique avoids frequency conflicts between the vocals and the song's intro (especially for songs with vocals that start without an intro).
[0144] Finally, the synthesized complete program audio is associated with the selected title and virtual character cover, and automatically uploaded and published to the podcast account corresponding to the virtual character, completing the entire production process of a single episode.
[0145] This embodiment also provides a music podcast program generation apparatus, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0146] This embodiment provides a device for generating music podcast programs, such as... Figure 5 As shown, it includes: The acquisition module 501 is used to acquire the target song list and the playback style description information for playing the target song list; The generation module 502 is used to generate a podcast script based on the song information and the broadcast style description information in the target song list; the podcast script contains song lyrics with the style corresponding to the broadcast style description information. The synthesis module 503 is used to convert the podcast transcript into speech audio and synthesize the speech audio with the song audio corresponding to the target song list to obtain a music podcast program.
[0147] In some alternative implementations, the acquisition module 501 is used for: Based on the broadcast style description information, candidate songs that match the broadcast style description information are selected from the preset music database; The candidate songs are filtered according to the preset filtering rules to obtain the target song list.
[0148] In some alternative implementations, the acquisition module 501 is further configured to: Query whether a music podcast containing candidate songs has been generated based on the broadcast style description information within a preset historical time period; If so, the candidate song will be removed from the target song list.
[0149] In some alternative implementations, the generation module 502 is used for: Construct input instructions, which include broadcast style description information, text format constraints, and metadata for each song in the target song list. The metadata includes at least the song name and the artist name. Input the commands into the text generation model to obtain the podcast transcript.
[0150] In some alternative implementations, the generation module 502 is used for: Encapsulate the metadata of all songs in the target song list into a single input command; A podcast transcript containing the lyrics of all songs in the target song list is generated using a text generation model.
[0151] In some alternative implementations, the synthesis module 503 is used for: The audio is synthesized sequentially according to the playback order of the songs in the target song list, and audio transition processing is performed on at least one song audio to achieve a smooth transition between the speech audio and the song audio.
[0152] In some alternative implementations, the synthesis module 503 is used for: Determine the entry and exit times for each song's audio based on the playback order of the songs in the target song list. Apply volume gradation processing to the song audio during the entry or exit period to create a volume increase or decrease effect. At least a portion of the speech audio is superimposed onto the time period corresponding to the volume gradation feature, and the speech audio is distributed only within the range of the entry time period and / or exit time period.
[0153] In some alternative implementations, the synthesis module 503 is further configured to: Based on podcast transcripts and broadcast style descriptions, generate at least one candidate program title; Select the target title from the candidate program titles; Upload the music podcast program, target title, and cover image associated with the broadcast style description to the content publishing platform.
[0154] In some alternative implementations, the generation module 502 is further configured to: In response to modification or regeneration instructions for the podcast transcript, adjust the broadcast style description or transcript format constraints in the input instructions, and re-execute the steps to generate the podcast transcript.
[0155] The music podcast program generation apparatus provided in this disclosure can execute the music podcast program generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0156] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0157] The following is a detailed reference. Figure 6 This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0158] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0159] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the method for generating a music podcast program according to embodiments of this disclosure.
[0160] Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0161] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the method for generating music podcast programs shown in the above embodiments.
[0162] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0163] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for generating a music podcast program, characterized in that, The method includes: Obtain the target song list and the broadcast style description information used to broadcast the target song list; Based on the song information in the target song list and the broadcast style description information, a podcast script is generated; the podcast script contains song lyrics with the style corresponding to the broadcast style description information. The podcast transcript is converted into audio, and the audio is combined with the audio of the songs corresponding to the target song list to obtain a music podcast program.
2. The method according to claim 1, characterized in that, The process of obtaining the target song list includes: Based on the broadcast style description information, candidate songs that match the broadcast style description information are selected from a preset music database; The candidate songs are filtered according to preset filtering rules to obtain the target song list.
3. The method according to claim 2, characterized in that, After obtaining the target song list, the method further includes: Query whether a music podcast containing the candidate songs has been generated based on the broadcast style description information within a preset historical time period; If so, the candidate song will be removed from the target song list.
4. The method according to claim 1, characterized in that, The step of generating a podcast script based on the song information in the target song list and the broadcast style description information includes: Construct an input instruction, which includes the broadcast style description information, text format constraints, and metadata of each song in the target song list, wherein the metadata includes at least the song name and the artist name; The input instructions are fed into the text generation model to obtain the podcast transcript.
5. The method according to claim 4, characterized in that, The step of inputting the input command into the text generation model to obtain the podcast transcript includes: The metadata of all songs in the target song list is encapsulated into the same input instruction; The text generation model generates a podcast transcript containing the strings of words from all the songs in the target song list.
6. The method according to claim 1, characterized in that, The step of synthesizing the speech audio with the song audio corresponding to the target song list includes: The audio is synthesized in a temporal sequence according to the playback order of the songs in the target song list, and audio transition processing is performed on at least one of the song audios to achieve a smooth transition between the speech audio and the song audio.
7. A device for generating music podcast programs, characterized in that, The device includes: The acquisition module is used to acquire a list of target songs and a description of the playback style for broadcasting the list of target songs; The generation module is used to generate a podcast script based on the song information in the target song list and the broadcast style description information; the podcast script contains song lyrics with the style corresponding to the broadcast style description information; The synthesis module is used to convert the podcast transcript into audio and synthesize the audio with the audio of the songs corresponding to the target song list to obtain a music podcast program.
8. An electronic device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method for generating a music podcast program according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method for generating a music podcast program according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the method for generating a music podcast program as described in any one of claims 1 to 6.