Song splicing method and device, electronic equipment and storage medium
By matching playlist and song information with target templates, filtering and segmenting songs into musical phrases, and generating spliced songs, the problem of users finding it difficult to quickly understand the features of playlists is solved, improving music discovery efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-01
AI Technical Summary
The current playlist is displayed with user-defined text descriptions and song lists, making it difficult for users to quickly understand the characteristics of the songs in the playlist, resulting in low music discovery efficiency and a poor user experience.
By determining the playlist and song information, matching the target template, filtering and segmenting the songs into musical phrases, and generating spliced songs based on the template parameters, the recombination and splicing of songs is achieved.
It helps users quickly understand the song composition and theme of a playlist, creating a novel auditory experience and improving music discovery efficiency and user experience.
Smart Images

Figure CN121958601A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to a song splicing method, apparatus, electronic device, and storage medium. Background Technology
[0002] In today's rapidly developing digital music landscape, playlists have become a core platform for organizing and sharing music on major music platforms. They are typically created by users who select a series of songs based on personal preferences or specific themes. However, current playlists are usually displayed with user-defined text descriptions paired with a song list. For users unfamiliar with the songs in a playlist, the brief text descriptions and song titles make it difficult to quickly grasp the characteristics of the songs. This forces users to spend a significant amount of time listening to each song individually to determine if the music in the playlist suits their needs, severely impacting music discovery efficiency and user experience. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a song splicing method, apparatus, electronic device and storage medium to help users understand the characteristics of songs in a playlist in a short time and improve music discovery efficiency and user experience.
[0004] In a first aspect, embodiments of the present invention provide a song splicing method, the method comprising: determining playlist information of a target playlist and song information of songs in the target playlist; wherein, the playlist information includes: a playlist name and descriptive text for describing the theme style of the playlist; the song information includes: song background information and song music attribute information; determining a target template from candidate theme type templates based on the playlist information and the song information; wherein, the theme type template includes song constraint information corresponding to a preset theme type and structural splicing parameters of the spliced song to be generated; selecting target songs from the target playlist based on the song constraint information and song information in the target template, and dividing the target songs into multiple musical phrases; generating a spliced song based on the multiple musical phrases, the song information of the target songs, and the structural splicing parameters in the target template.
[0005] Secondly, embodiments of the present invention also provide a song splicing device, the device comprising: a first determining module, configured to determine playlist information of a target playlist and song information of songs in the target playlist; wherein the playlist information includes: a playlist name and descriptive text describing the theme style of the playlist; the song information includes: song background information and song music attribute information; a second determining module, configured to determine a target template from candidate theme type templates based on the playlist information and the song information; wherein the theme type template includes song constraint information corresponding to a preset theme type and structural splicing parameters of the spliced song to be generated; a first segmentation module, configured to filter out target songs from the target playlist based on the song constraint information and song information in the target template, and segment the target songs into multiple musical phrase fragments; and a first splicing module, configured to generate a spliced song based on the multiple musical phrase fragments, the song information of the target song, and the structural splicing parameters in the target template.
[0006] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned song splicing method.
[0007] Fourthly, embodiments of the present invention provide a storage medium storing machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above-mentioned song splicing method.
[0008] The embodiments of the present invention bring the following beneficial effects: This invention provides a song splicing method, apparatus, electronic device, and storage medium. The method includes: determining playlist information of a target playlist and song information of the songs in the target playlist; wherein, the playlist information includes: a playlist name and descriptive text describing the playlist's theme style; the song information includes: song background information and song musical attribute information; based on the playlist information and the song information, determining a target template from candidate theme type templates; wherein, the theme type template contains song constraint information corresponding to a preset theme type and structural splicing parameters of the spliced song to be generated; based on the song constraint information and song information in the target template, selecting target songs from the target playlist and dividing the target songs into multiple musical phrases; and generating a spliced song based on the multiple musical phrases, the song information of the target songs, and the structural splicing parameters in the target template.
[0009] This method matches target templates that fit the playlist theme based on playlist information and song information within the playlist. Target songs are then selected based on the song constraints within the target templates, and these target songs are divided into multiple musical phrases. Finally, the musical phrases are spliced together based on the musical phrases, the song information of the target songs, and the structural splicing parameters in the target templates to obtain the spliced song. This method, by extracting and recombining musical phrases from songs within a playlist, helps users quickly perceive the song composition and theme of the playlist, creates a novel auditory experience, effectively enhances the fun of browsing playlists, improves music discovery efficiency, and enhances the user experience.
[0010] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 A flowchart of a song splicing method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of an audio segment corresponding to a target song provided in an embodiment of the present invention; Figure 3 A schematic diagram illustrating a first result provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a song splicing device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] Based on this, the song splicing method, apparatus, electronic device, and storage medium provided in this embodiment of the invention can be applied to music splicing scenarios. It is worth noting that the executing entity of this invention can be a system device, a terminal, or a server; no specific limitation is made here.
[0016] To facilitate understanding of this embodiment, a song splicing method disclosed in this invention will first be described in detail, such as... Figure 1 As shown, this method includes the following steps: Step S102: Determine the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text used to describe the theme and style of the playlist; the song information includes: song background information and song music attribute information.
[0017] The aforementioned target playlists can be created by music users themselves or organized by the music platform operator. The descriptive text above describes the playlist's theme and style, summarizing its core theme and emotional tone. The aforementioned song background information includes lyrics, release year, artist information, and creative background information. This type of information belongs to the song's static attributes and is usually stored in a pre-built resource library, which can be directly retrieved by song name. The aforementioned song musical attribute information refers to quantitative and qualitative data that characterizes the song's acoustic and performance features, including song style, emotional atmosphere, BPM data, instrument information, key, and temporal distribution of beats. This type of information can be extracted by analyzing the song's audio signal using a deep learning model, and then verified by manual correction by operational editors or user feedback.
[0018] For example, a music user created a playlist called "Weekend Relaxation," selected 10 songs such as "Imagine" and "Yesterday," added a descriptive text to the playlist: "Classic melodies to accompany you through a relaxing weekend," and submitted it.
[0019] After obtaining the playlist and description text submitted by the user, the system determines the song information corresponding to the playlist and outputs and stores the playlist information and song information in a unified structured format, such as JSON format.
[0020] Step S104: Based on the playlist information and song information, determine the target template from the candidate theme type templates; wherein, the theme type template contains the song constraint information corresponding to the preset theme type and the structural splicing parameters of the spliced song to be generated.
[0021] The aforementioned alternative theme type templates contain song constraint information corresponding to the preset theme type and structural splicing parameters for the generated spliced song. Song constraint information can be understood as the rules used to filter songs that meet the requirements within the playlist, matching the preset theme type. This serves as the basis for subsequently selecting target songs from the target playlist and includes: BPM constraint information, key constraint information, genre constraint information, and mood / atmosphere constraint information. The aforementioned structural splicing parameters can be understood as the rules and parameter configurations for splicing musical phrases to achieve the generation of spliced songs with the preset theme style. These are key technical parameters guiding the selection, sorting, and combination of musical phrases.
[0022] Here, the target template can be determined by a large language model. Specifically, playlist information, song information, and alternative theme type templates can be filled into a preset prompt word template to generate prompt words. By inputting the prompt words into the large language model, the large language model is guided to determine the theme type of the target playlist based on the playlist information and song information, and to determine the target template that matches the theme type of the target playlist from the alternative theme type templates.
[0023] Step S106: Based on the song constraint information and song information in the target template, select the target song from the target playlist and divide the target song into multiple musical phrases.
[0024] The aforementioned musical phrases can be understood as the smallest musical structural units that constitute a song. They usually contain complete melodic phrases or rhythmic phrases, such as melodic fragments corresponding to a single line of lyrics, or several consecutive musical measures. They possess independent musical expressive attributes and can be used as basic units for splicing operations.
[0025] After determining the target template, the song constraints in the target template are compared item by item with the song information of each song in the target playlist. Combined with preset filtering conditions, target songs that highly match the theme type of the target template are selected. To ensure the selection space for subsequent musical phrase splicing, the number of target songs must meet a preset threshold. If the number of selected songs does not reach this threshold, the filtering condition adjustment mechanism can be automatically triggered to relax the initial screening threshold and re-filter.
[0026] Subsequently, for each selected target song, an audio processing model, such as the open-source All-In-One tool, can be used to identify audio segments of different functional types in its musical structure, including typical functional types such as intro, verse, chorus, interlude, and coda. The audio segments of the above different functional types are then segmented to obtain musical phrases with independent musical expression attributes.
[0027] Step S108: Generate a spliced song based on multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template.
[0028] In this step, after obtaining the musical phrase fragments, it is necessary to determine the corresponding musical phrase information for each musical phrase fragment, specifically including the identification information of the musical phrase fragment, the storage path information of the musical phrase fragment, the corresponding musical form structure function type, lyrics information, audio energy score and other data.
[0029] Subsequently, the phrase information, the target song information, and the structural splicing parameters from the target template are input into the phrase splicing model. Guided by preset prompts, the model analyzes and processes the data based on the structural splicing parameters and preset constraints, outputting the target phrase fragments to be spliced, as well as the connection methods between these fragments. Finally, based on the selected target phrase fragments and the generated connection methods, audio preprocessing and splicing operations are performed to generate the final spliced song.
[0030] This application innovatively transforms traditional static playlists into a medley format, creating a more vivid, interesting, and emotionally rich music playback experience for users. At the same time, it relies on artificial intelligence technology to build a fully automated process from song information acquisition to audio synthesis, enabling the generation of high-quality song medleys with one click, greatly improving the efficiency and convenience of audio content creation.
[0031] The above-mentioned song splicing method includes: determining the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text used to describe the theme style of the playlist; the song information includes: song background information and song music attribute information; based on the playlist information and the song information, determining the target template from the candidate theme type templates; wherein, the theme type template contains song constraint information corresponding to the preset theme type and structural splicing parameters of the spliced song to be generated; based on the song constraint information and song information in the target template, selecting the target song from the target playlist and dividing the target song into multiple musical phrases; and generating the spliced song based on the multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template.
[0032] This method matches target templates that fit the playlist theme based on playlist information and song information within the playlist. Target songs are then selected based on the song constraints within the target templates, and these target songs are divided into multiple musical phrases. Finally, the musical phrases are spliced together based on the musical phrases, the song information of the target songs, and the structural splicing parameters in the target templates to obtain the spliced song. This method, by extracting and recombining musical phrases from songs within a playlist, helps users quickly perceive the song composition and theme of the playlist, creates a novel auditory experience, effectively enhances the fun of browsing playlists, improves music discovery efficiency, and enhances the user experience.
[0033] The following embodiments provide specific implementation methods for determining playlist information and song information.
[0034] Specifically, a target playlist is created, and the playlist information is determined. Background information of the songs in the target playlist is retrieved from a pre-set resource library. The background information of the songs includes lyrics, release year, artist, and creative background information. The audio signals of the songs in the target playlist are obtained, and the musical attribute information of the songs is determined based on the audio signals. The musical attribute information of the songs includes at least the following: song style, song mood and atmosphere, BPM data, instrument information, song key, and temporal distribution information of the beat.
[0035] After a user creates a target playlist, add a playlist name and descriptive text to describe the playlist's theme and style, and submit it. After obtaining the user's submitted playlist and descriptive text, determine the playlist information and song information.
[0036] The aforementioned background information about the songs is stored in the database as static attributes of the songs, including but not limited to the song title, artist, album, release year, lyrics, and creative background. This information can be obtained by searching the resource library. For example, searching for detailed information about the song "Bohemian Rhapsody" will reveal that it is a work by the band Queen, included in the album "A Night at the Opera," released in 1975, as well as the complete lyrics.
[0037] The aforementioned musical attribute information includes: song style (e.g., rock, jazz, classical, pop, hip-hop, electronic, country, etc.), song mood / atmosphere (e.g., celebratory, melancholic, romantic, longing, etc.), BPM data (e.g., 60 beats / minute, 80 beats / minute), instrument information (e.g., piano, guitar, drums, bass), song key (e.g., C major, A minor), and temporal distribution information of the beats. The temporal distribution information of the beats records the specific time points and temporal distribution of each strong beat and weak beat node in the song's audio stream.
[0038] The extraction of musical attribute information from songs can be achieved using deep learning technology. First, the song's audio signal can be preprocessed, including noise reduction and normalization, to improve the accuracy of subsequent analysis. Second, key musical features, such as amplitude spectrum and Mel spectrum, can be extracted from the preprocessed audio signal; these features characterize the core attributes of the musical signal. Then, deep learning models such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) are used to learn and classify the extracted features, enabling automatic identification of the song's musical attribute information. Furthermore, manual correction by editors or user feedback can be combined to verify and correct the model's results, further improving the accuracy of the song's musical attribute information.
[0039] Finally, the acquired playlist information, along with the determined song information, can be uniformly output and stored in a structured format, such as JSON.
[0040] The following examples provide a specific implementation method for determining the target template.
[0041] Specifically, playlist information, song information, and alternative theme type templates are filled into a preset prompt word template to generate prompt words; the prompt words are then input into a large language model to guide the large language model to determine the theme type of the target playlist based on the playlist information and song information, and to determine the target template that matches the theme type of the target playlist from the alternative theme type templates.
[0042] In other words, the playlist information, song background information, song music attribute information, and alternative theme type templates are first filled into a preset prompt word template to generate a clear and instructive input prompt word. Then, the prompt word is input into a large language model. Based on its semantic understanding and association analysis capabilities, the large language model first extracts the core theme type of the target playlist from the playlist information and song information. Finally, the identified playlist theme type is matched with the alternative theme type templates to select the template with the highest theme fit as the target template, providing a basis for subsequent song selection and musical phrase splicing.
[0043] The following embodiments provide a specific implementation method for determining the target song.
[0044] In one approach, the song constraint information includes: BPM constraint information, mode constraint information, genre constraint information, and mood constraint information. A specified BPM interval corresponding to the BPM constraint information is obtained. Based on preset filtering parameters and the specified BPM interval, a BPM filtering interval is determined. Songs in the target playlist whose BPM data falls within the BPM filtering interval are identified as the first song. The mode of the first song is obtained. Songs in the first song whose mode matches the specified mode in the mode constraint information, and songs whose mode modulates to the specified mode with a modulation span not exceeding a preset whole tone, are identified as the second song. From the second song, a third song whose genre matches the specified genre in the genre constraint information is identified. The mood and lyrics information of the third song are obtained. Songs whose mood matches the specified mood in the mood constraint information, or whose lyrics match the specified mood, are identified as target songs. The number of target songs meets a preset threshold.
[0045] This method employs a progressive filtering logic, narrowing down the candidate range layer by layer based on the four dimensions of song constraint information, ultimately selecting target songs whose quantity meets a preset threshold. The specific filtering logic is as follows: 1) BPM Dimension Filtering: Based on the specified BPM range corresponding to the BPM constraint information in the target template, and combined with preset filtering parameters, the final BPM filtering range is determined. Songs in the target playlist whose BPM data falls within this range are selected as the first song. For example, songs whose BPM data is no more than 5% different from the specified BPM range can be kept as the first song, while songs whose BPM data exceeds the range are deleted.
[0046] 2) Mode Dimension Filtering: Extract the mode information of the first song and filter out two types of songs as the second song. One type is a song whose mode is completely consistent with the specified mode in the mode constraint information. The other type is a song whose mode transposition to the specified mode does not exceed the preset whole tone number transposition span calculation formula. Here, the specified mode can include multiple modes.
[0047] 3) Music style dimension filtering: Select songs from the second song whose music style is the same as the music style specified in the music style constraint information, and determine them as the third song. Here, the specified music style can include multiple types.
[0048] 4) Final screening based on emotional atmosphere dimension: Obtain the emotional atmosphere characteristics and lyrics information of the third song, and select songs whose emotional atmosphere is consistent with the specified emotional atmosphere in the emotional atmosphere constraint information, or whose lyrics emotional tendency matches the specified emotional atmosphere as target songs; here, the specified emotional atmosphere can include multiple types.
[0049] The following embodiments provide a specific implementation method for dividing a target song into multiple musical phrases.
[0050] For each target song, the target audio signal corresponding to the target song is obtained, and the target audio signal is input into a pre-trained audio processing model. The model outputs multiple audio segments and their association information. Each audio segment corresponds to a different function type in the musical structure of the target song. The association information of the audio segments includes: the start time, end time, and function type of the audio segment. The audio segments are then segmented to obtain multiple musical phrases.
[0051] The aforementioned audio processing model can employ an open-source All-In-One model, which possesses efficient music structure analysis capabilities and can accurately identify audio segments of various functional types within the corresponding musical structure, such as intro, verse, chorus, interlude, and outro. This model is built on a deep learning architecture and, after training on large-scale music data, can learn and recognize the characteristics and patterns of different musical segments. When the audio signal corresponding to the target song is input into this model, it will output multiple audio segments of different functional types within the target song, along with the associated information for each audio segment, specifically including the start time, end time, and functional type of the audio segment. For example... Figure 2 As shown, Figure 2 The target song is divided into 7 audio segments according to the timeline. The functional types of the audio segments are: Intro, Verse, Chorus, Bridge, Verse, Chorus, and Outro, which correspond to the common song structure of popular music.
[0052] Subsequently, the audio segments of all functional types in the target song were finely segmented to obtain multiple musical phrases with independent musical expression attributes. The core objectives of this operation are twofold: first, to accurately anchor the start and end points of vocals, avoiding the problem of vocals being truncated during splicing and ensuring the smoothness of the audio; and second, to achieve finer-grained material splitting, providing ample choice for flexible splicing of subsequent songs and enhancing the diversity and adaptability of splicing schemes.
[0053] In one approach, the audio segment is an audio segment with lyrics; the lyrics segment corresponding to the audio segment is obtained, and the start and end time points of the lyrics corresponding to each line of lyrics in the lyrics segment are determined; wherein, the start and end time points of the lyrics include: the start time of the lyrics and the end time of the lyrics; for each line of lyrics, the strong beat time points of the audio segment within a preset duration before and after the start and end time points of the lyrics are determined, and the audio segment is cut at the first strong beat time point closest to the start time of the lyrics and the second strong beat time point closest to the end time of the lyrics, respectively, to obtain the musical phrase segment corresponding to each line of lyrics.
[0054] The strong beats mentioned above refer to the accented beats in a song's beat sequence that conform to the time signature rules, such as the beat number 1 in 4 / 4 time. The timing of strong beats can be obtained from the temporal distribution information of the beats in the song's musical attribute information.
[0055] For audio segments with lyrics, the timestamps corresponding to each word in the lyrics can be obtained from the word-by-word lyrics of the song. The word-by-word lyrics timestamps are parsed, and the lyrics are divided into individual lines by combining punctuation marks (such as commas, periods, exclamation marks, etc.) and semantic pause rules (such as pauses between complete sentences of subject, verb, and object). The start time and end time of each line of lyrics are then output.
[0056] Subsequently, for each lyric, the strong beat time points within 0.5 seconds before and after the start and end times of the lyrics can be found. Audio segments are then cut at the first strong beat time point closest to the start time and the second strong beat time point closest to the end time of the lyrics to obtain the musical phrase segment corresponding to each lyric. This ensures that the cutting position falls within the rhythm gap and avoids an abrupt sound after splicing.
[0057] In another approach, the audio segment is a wordless audio segment; the duration of a single measure of the target song is determined based on the BPM data of the target song to which the audio segment belongs and the preset number of beats per measure; the duration of the musical phrase is determined based on the preset number of measures contained in the musical phrase; the audio segment is divided according to the duration of the musical phrase to obtain multiple musical phrase segments.
[0058] In other words, when the audio segment is a wordless segment, the duration of a single measure in the target song is calculated by combining the BPM data of the target song to which the audio segment belongs with the preset number of beats per measure. A formula for calculating the duration of a single measure is provided below: N*60 / B; where... The duration of a single segment; N represents the number of beats per measure; B represents the BPM data. Then, based on the preset number of measures in a musical phrase, the duration of a single musical phrase is further calculated. For example, if a musical phrase is preset to contain 4 measures, the duration of a single musical phrase can be calculated based on the duration of a single measure. Finally, using the calculated duration of the musical phrase as the unit of division, the audio segment without lyrics is divided evenly and regularly, resulting in multiple musical phrase segments with complete rhythmic structures.
[0059] In one approach, before splicing musical phrases, the phrase fragments can be preprocessed based on the mode constraints in the target template song's constraint information. Specifically, the average BPM data of multiple musical phrases is determined, and the BPM data of the musical phrases is adjusted to the average value; the song mode of the musical phrases is determined, and the song mode of the musical phrases is converted to the specified mode in the mode constraint information; the audio format of multiple musical phrases is converted to a target format with a specified sampling rate.
[0060] Here, to ensure the rhythmic consistency, tonality and mode constraint information of the subsequently spliced songs are consistent, and the sampling is consistent, three standardized preprocessing operations need to be performed on all the selected musical phrase fragments: Rhythm standardization, which means changing the speed but not the key, first calculates the average BPM data of all musical phrases, and then adjusts the BPM data of each musical phrase to this average value to ensure that the rhythm of the spliced song is continuous and without breaks.
[0061] Then, the key of each musical phrase is extracted. This key matches the functional type of the audio segment to which the phrase belongs in the target song's musical structure. An audio modulation algorithm is then called to convert the musical phrase's key to the specified key in the key constraint information. It should be noted that there can be multiple specified keys in the target template. For each specified key, the original key of the musical phrase can be converted to its corresponding key version, forming multiple sets of tonalally unified musical phrases.
[0062] Finally, the audio format of all musical phrases is uniformly converted to a target format with a specified sampling rate, such as WAV format with a sampling rate of 44100Hz and a bit depth of 16 bits, to ensure format compatibility when splicing different musical phrases and avoid audio distortion caused by parameter mismatch.
[0063] The following embodiments provide a specific implementation method for generating spliced songs.
[0064] In one approach, the splicing structure parameters include: a sequence of reference function types for the song to be spliced, the duration percentage of each reference function type in the sequence, and the audio energy score range of the corresponding audio segments of each reference function type; obtaining phrase information corresponding to multiple phrase segments; wherein the phrase information includes: the identifier of the phrase segment, the storage path information of the phrase segment, the function type corresponding to the phrase segment in the musical structure of the target song, the lyrics information of the phrase segment, and the audio energy score of the phrase segment; the audio energy score is used to indicate the energy intensity of the audio signal; inputting the phrase information, the song information of the target song, the target template, and preset prompts into the phrase splicing model, and guiding the phrase splicing model based on the structural splicing parameters and preset constraints through the preset prompts, outputting a first result; wherein the first result includes: the target phrase segments participating in the splicing, and the connection method between the target phrase segments; and generating a spliced song based on the target phrase segments and the connection method.
[0065] The aforementioned reference functional type sequence refers to the order of the musical structure of the song to be generated, such as the distribution of functional type segments like intro → verse → chorus → interlude → chorus → coda, encompassing the functional types and distribution of the spliced song in terms of musical structure. The aforementioned duration percentage of the reference functional types refers to the required proportion of each functional type segment in the total duration of the spliced song, such as 10% for the intro segment, 25% for the verse segment, and 30% for the chorus segment, ensuring the structural balance of the spliced song. The aforementioned audio energy score indicates the energy intensity of the audio signal. The higher the audio energy score, the greater the energy intensity of the audio signal. The aforementioned phrase splicing model can be a large language model. The aforementioned preset constraints indicate the core judgment criteria for selecting and connecting target phrase segments, ensuring that the target segments selected from the phrase segments, and the splicing methods between segments, fully conform to the structural and auditory requirements of the spliced song.
[0066] Here, the musical phrase information corresponding to all musical phrase fragments is extracted, specifically including the fragment's identifier, storage path, musical form / function type, lyrics, and audio energy score. This musical phrase information, along with the target song's information, the target template, and preset prompts, are then input into the musical phrase splicing model. Guided by the preset prompts, the model outputs a first result based on structural splicing parameters and preset constraints. This first result includes the target musical phrase fragments involved in the splicing, and the connection methods between these fragments. For example, connection methods include fade-in / fade-out, hard cut on a strong beat, and EQ transition.
[0067] Finally, based on the connection method between the target musical phrases, a coherent and harmonious spliced song is generated.
[0068] In one approach, the preset constraints include: eliminating musical phrases whose function type does not match the reference function type; determining musical phrases according to structural splicing parameters and preset splicing song duration, and determining the target musical phrases to be spliced; when determining the target musical phrases, prioritizing the selection of musical phrases with the same key, and marking the target musical phrases that need to be transposed if musical phrases with different keys need to be selected; arranging the target musical phrases of the same function type in ascending order of audio energy score; and ensuring that the difference in audio energy score between interconnected target musical phrases does not exceed a preset difference threshold.
[0069] Specifically, the first step is to select musical phrases from multiple phrases whose functional types completely match the reference functional types in the target template. For example, the reference functional type sequence in the target template "deeply emotional template" only includes four functional types: intro, verse, chorus, and coda. Therefore, musical phrases whose functional types do not belong to these functional types are eliminated.
[0070] Then, based on the structural splicing parameters and the preset splicing song duration, musical phrase segments are determined, and target musical phrase segments for splicing are identified. Specifically, the musical phrase information of the target musical phrase segments selected for the reference function type sequence must completely match the splicing structural parameters. For example, if the target template sets the duration of the intro function type segment to account for 10% and the splicing song duration to be 200 seconds, then the total duration of the selected intro type musical phrase segments must precisely match 20 seconds, and the audio energy score must fall within the audio energy score range corresponding to the intro function segment.
[0071] On the one hand, when determining the target musical phrase, priority is given to selecting phrases in the same key. If phrases in different keys need to be selected, the target musical phrases that require modulation are marked with modulation marks. On the one hand, target musical phrases of the same functional type are arranged in order of increasing audio energy score. For example, for the structural splicing parameters in the target template that include "the duration of the verse functional type segment is 30-90 seconds, and the audio energy score range is 4-6 points", three phrases are selected from the screened target musical phrase segments of the verse functional type, and the musical phrase segments are arranged in the order of audio energy score 4 points → 5 points → 6 points.
[0072] On the one hand, the difference in audio energy scores between interconnected target musical phrases does not exceed a preset difference threshold. For example, the audio energy score difference is checked on adjacent musical phrases after sorting, requiring that the difference in audio energy scores between interconnected target musical phrases is no greater than 2 points. If it exceeds the threshold, it is replaced with a musical phrase with a closer audio energy score to ensure the smoothness of energy transition after splicing.
[0073] The pre-defined prompt-guided musical phrase splicing model, based on structural splicing parameters and the aforementioned pre-defined constraints, can output the first result. For example, the first result is as follows: Figure 3 As shown, the first result is a list of target musical phrases and a suggested transition table. For example... Figure 3 (a) in the diagram represents the list of target musical phrases. This list displays the identifier of each target musical phrase segment (i.e., the phrase ID in the diagram), the functional type of the target musical phrase segment (i.e., the structural type in the diagram), the audio energy score of the target musical phrase segment (i.e., the energy value in the diagram), the duration of the target musical phrase segment, the key of the target musical phrase segment, and the key change markings. For example... Figure 3 (b) is a connection suggestion table, which contains the difference in audio energy scores of the target musical phrase segments that are connected to each other, as well as connection suggestions.
[0074] Finally, based on the specific schemes in the suggested connection method table, the corresponding connection processing is performed on the musical phrase fragments in the target musical phrase list to generate a spliced song that meets the structural and auditory requirements.
[0075] In one approach, the audio energy score of a musical phrase can be obtained by: discretely sampling the audio signal corresponding to the musical phrase to obtain multiple sampled signals; determining the root mean square (RMS) value of the amplitude values of the multiple sampled signals; and mapping the RMS value to a specified audio energy score interval to obtain the audio energy score of the musical phrase.
[0076] In this method, firstly, discrete sampling is performed on the audio signal corresponding to the musical phrase fragment to obtain several equally spaced sampled signals; then, the root mean square (RMS) value of the amplitude of these sampled signals is calculated using the following formula:
[0077] in, The total number of sampled signals, For the first The square of the amplitude value of each sampled signal.
[0078] Then, the root mean square value is mapped to the specified audio energy score interval to obtain the audio energy score of the musical phrase.
[0079] For example, the audio energy score range is set from 1 to 10, and the mapping rules are as follows: When the root mean square value is less than the first critical threshold, the audio energy score is mapped to 1 point; when the root mean square value is greater than the second critical threshold, the audio energy score is mapped to 10 points; the first critical threshold is less than the second critical threshold.
[0080] For the root mean square value between the first and second critical thresholds, the audio energy score is calculated using a linear mapping formula: first, calculate the first difference between the root mean square value and the first critical threshold; then, calculate the second difference between the second and first critical thresholds; calculate the ratio of the first difference to the second difference; multiply the ratio by 9 and add 1 to get the result; round the result to the nearest integer, which is the audio energy score corresponding to the musical phrase.
[0081] For the corresponding method embodiments described above, see [link to relevant documentation]. Figure 4 The diagram shows a song splicing device, which includes: The first determining module 402 is used to determine the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text used to describe the theme and style of the playlist; the song information includes: song background information and song music attribute information; The second determining module 404 is used to determine the target template from the candidate theme type templates based on the playlist information and song information; wherein, the theme type template contains song constraint information corresponding to the preset theme type and structural splicing parameters of the spliced song to be generated; The first segmentation module 406 is used to filter out target songs from the target playlist based on the song constraint information and song information in the target template, and to segment the target songs into multiple musical phrases. The first splicing module 408 is used to generate a spliced song based on multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template.
[0082] This method matches target templates that fit the playlist theme based on playlist information and song information within the playlist. Target songs are then selected based on the song constraints within the target templates, and these target songs are divided into multiple musical phrases. Finally, the musical phrases are spliced together based on the musical phrases, the song information of the target songs, and the structural splicing parameters in the target templates to obtain the spliced song. This method, by extracting and recombining musical phrases from songs within a playlist, helps users quickly perceive the song composition and theme of the playlist, creates a novel auditory experience, effectively enhances the fun of browsing playlists, improves music discovery efficiency, and enhances the user experience.
[0083] The first determining module is used to create a target playlist and determine the playlist information of the target playlist; retrieve the background information of the songs in the target playlist from a preset resource library; wherein, the background information of the songs includes lyrics information, release year information, artist information, and creation background information; obtain the audio signal of the songs in the target playlist, and determine the music attribute information of the songs based on the audio signal; wherein, the music attribute information of the songs includes at least: song style, song mood atmosphere, BPM data, instrument information, song key, and temporal distribution information of the beat.
[0084] The second determining module fills the playlist information, song information, and alternative theme type templates into the preset prompt word template to generate prompt words; the prompt words are input into the large language model to guide the large language model to determine the theme type of the target playlist based on the playlist information and song information, and to determine the target template that matches the theme type of the target playlist from the alternative theme type templates.
[0085] The aforementioned song constraint information includes: BPM constraint information, mode constraint information, genre constraint information, and mood / atmosphere constraint information. The first segmentation module is used to obtain the specified BPM interval corresponding to the BPM constraint information, determine the BPM filtering interval based on preset filtering parameters and the specified BPM interval, and identify songs in the target playlist whose BPM data falls within the BPM filtering interval as the first song; obtain the mode of the first song, and identify songs in the first song whose mode matches the specified mode in the mode constraint information, as well as songs whose mode changes to the specified mode with a modulation span not exceeding a preset whole tone, as the second song; from the second song, identify the third song whose genre matches the specified genre in the genre constraint information; obtain the mood / atmosphere and lyrics information of the third song, and identify songs whose mood / atmosphere matches the specified mood / atmosphere in the mood / atmosphere constraint information, or whose lyrics match the specified mood / atmosphere, as the target song; wherein, the number of target songs meets a preset number threshold.
[0086] The first segmentation module described above is used to acquire the target audio signal corresponding to each target song, input the target audio signal into a pre-trained audio processing model, and output multiple audio segments and their associated information. Each audio segment corresponds to a different functional type in the musical structure of the target song. The associated information of the audio segments includes: the start time, end time, and functional type of the audio segment. The audio segments are segmented to obtain multiple musical phrase segments.
[0087] The aforementioned audio segment is an audio segment with lyrics; the aforementioned first segmentation module is used to obtain the lyric segment corresponding to the audio segment and determine the start and end time points of the lyrics corresponding to each line of lyrics in the lyric segment; wherein, the start and end time points of the lyrics include: the start time of the lyrics and the end time of the lyrics; for each line of lyrics, the strong beat time points of the audio segment within a preset duration before and after the start and end time points of the lyrics are determined, and the audio segment is segmented at the first strong beat time point closest to the start time of the lyrics and the second strong beat time point closest to the end time of the lyrics, respectively, to obtain the musical phrase segment corresponding to each line of lyrics.
[0088] The above audio segment is an audio segment without lyrics; the above first segmentation module is used to determine the duration of a single measure of the target song based on the BPM data of the target song to which the audio segment belongs and the preset number of beats per measure; to determine the duration of the musical phrase corresponding to the musical phrase based on the preset number of measures contained in the musical phrase; and to segment the audio segment according to the duration of the musical phrase to obtain multiple musical phrase segments.
[0089] The aforementioned song constraint information includes: mode constraint information; the aforementioned device includes a first conversion module, used to determine the average value of BPM data of multiple musical phrase segments, adjust the BPM data of the musical phrase segments to the average value; determine the song mode of the musical phrase segments, convert the song mode of the musical phrase segments to the specified mode in the mode constraint information; and convert the audio format of multiple musical phrase segments to a target format with a specified sampling rate.
[0090] The aforementioned splicing structure parameters include: a sequence of reference function types for the song to be spliced, the duration percentage of each reference function type in the sequence, and the audio energy score range of the corresponding audio segments of each reference function type; the aforementioned first splicing module is used to acquire the phrase information corresponding to multiple phrase segments; wherein, the phrase information includes: the identifier information of the phrase segment, the storage path information of the phrase segment, the function type corresponding to the phrase segment in the musical structure of the target song, the lyrics information of the phrase segment, and the audio energy score of the phrase segment; the audio energy score is used to indicate the energy intensity of the audio signal; the phrase information, the song information of the target song, the target template, and the preset prompt words are input into the phrase splicing model, and the preset prompt words guide the phrase splicing model to output a first result based on the structural splicing parameters and preset constraints; wherein, the first result includes: the target phrase segments participating in the splicing, and the connection method between the target phrase segments; the spliced song is generated based on the target phrase segments and the connection method.
[0091] The aforementioned preset constraints include: eliminating musical phrases whose function type does not match the reference function type; determining musical phrases according to the structural splicing parameters and the preset splicing song duration, and determining the target musical phrases to be spliced; when determining the target musical phrases, prioritizing musical phrases with the same key, and if musical phrases with different keys need to be selected, marking the target musical phrases that need to be transposed; arranging target musical phrases of the same function type in ascending order of audio energy score; and ensuring that the difference in audio energy score between interconnected target musical phrases does not exceed a preset difference threshold.
[0092] The aforementioned device includes a second acquisition module, used to discretely sample the audio signal corresponding to the musical phrase fragment to obtain multiple sampled signals; determine the root mean square value of the amplitude values of the multiple sampled signals; and map the root mean square value to a specified audio energy score interval to obtain the audio energy score of the musical phrase fragment.
[0093] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-described song splicing method. This electronic device can be a server or a terminal device.
[0094] See Figure 5As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100. The processor 100 executes the computer-executable instructions to implement the above-described song splicing method.
[0095] Furthermore, Figure 5 The illustrated electronic device also includes a bus 102 and a communication interface 103. The processor 100, communication interface 103, and memory 101 are connected via the bus 102. The memory 101 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk drive. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless). The interface can use the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 5The diagram uses only a single bidirectional arrow, but this does not imply a single bus or a single type of bus. Processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 100 or by instructions in software form. Processor 100 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101, and the processor 100 reads the information from memory 101 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0096] The processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the song splicing method described above: determining the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text used to describe the theme style of the playlist; the song information includes: song background information and song music attribute information; based on the playlist information and the song information, determining the target template from the candidate theme type templates; wherein, the theme type template contains song constraint information corresponding to the preset theme type and structural splicing parameters of the spliced song to be generated; based on the song constraint information and song information in the target template, selecting the target song from the target playlist and dividing the target song into multiple musical phrases; generating the spliced song based on the multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template.
[0097] This method matches target templates that fit the playlist theme based on playlist information and song information within the playlist. Target songs are then selected based on the song constraints within the target templates, and these target songs are divided into multiple musical phrases. Finally, the musical phrases are spliced together based on the musical phrases, the song information of the target songs, and the structural splicing parameters in the target templates to obtain the spliced song. This method, by extracting and recombining musical phrases from songs within a playlist, helps users quickly perceive the song composition and theme of the playlist, creates a novel auditory experience, effectively enhances the fun of browsing playlists, improves music discovery efficiency, and enhances the user experience.
[0098] The processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the aforementioned song splicing method: creating a target playlist and determining the playlist information; retrieving background information of songs in the target playlist from a preset resource library; wherein, the background information of the songs includes lyrics information, release year information, artist information, and creation background information; acquiring the audio signals of the songs in the target playlist, and determining the song's musical attribute information based on the audio signals; wherein, the song's musical attribute information includes at least: song style, song mood and atmosphere, BPM data, instrument information, song key, and temporal distribution information of the beat.
[0099] The processor in the aforementioned electronic device can execute computer-executable instructions to perform the following operations of the song splicing method: filling playlist information, song information, and alternative theme type templates into a preset prompt word template to generate prompt words; inputting the prompt words into a large language model to guide the large language model to determine the theme type of the target playlist based on the playlist information and song information, and to determine the target template that matches the theme type of the target playlist from the alternative theme type templates.
[0100] The aforementioned song constraint information includes: BPM constraint information, mode constraint information, genre constraint information, and mood / atmosphere constraint information. The processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the aforementioned song splicing method: obtain the specified BPM interval corresponding to the BPM constraint information; determine the BPM filtering interval based on preset filtering parameters and the specified BPM interval; and identify songs in the target playlist whose BPM data falls within the BPM filtering interval as the first song; obtain the mode of the first song; identify songs in the first song whose mode matches the specified mode in the mode constraint information, and songs whose mode modulates to the specified mode with a modulation span not exceeding a preset whole tone, as the second song; from the second song, identify the third song whose genre matches the specified genre in the genre constraint information; obtain the mood / atmosphere and lyrics information of the third song; and identify songs whose mood / atmosphere matches the specified mood / atmosphere in the mood / atmosphere constraint information, or whose lyrics match the specified mood / atmosphere, as the target song; wherein the number of target songs meets a preset quantity threshold.
[0101] The processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the song splicing method: for each target song, acquire the target audio signal corresponding to the target song, input the target audio signal into a pre-trained audio processing model, and output multiple audio segments and their associated information. Each audio segment corresponds to a different functional type within the musical structure of the target song. The associated information of the audio segments includes: the start time, end time, and functional type of the audio segment. The audio segments are then segmented to obtain multiple musical phrases.
[0102] The aforementioned audio segment is an audio segment with lyrics; the processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the aforementioned song splicing method: obtaining the lyric segment corresponding to the audio segment, and determining the start and end time points of the lyrics corresponding to each line of lyrics in the lyric segment; wherein, the start and end time points of the lyrics include: the start time of the lyrics and the end time of the lyrics; for each line of lyrics, determining the strong beat time points of the audio segment within a preset duration before and after the start and end time points of the lyrics, and cutting the audio segment at the first strong beat time point closest to the start time of the lyrics and the second strong beat time point closest to the end time of the lyrics, respectively, to obtain the musical phrase segment corresponding to each line of lyrics.
[0103] The aforementioned audio segment is an audio segment without lyrics; the processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the aforementioned song splicing method: determining the duration of a single measure of the target song based on the BPM data of the target song to which the audio segment belongs and the preset number of beats per measure; determining the duration of the musical phrase corresponding to the musical phrase based on the preset number of measures contained in the musical phrase; and segmenting the audio segment according to the musical phrase duration to obtain multiple musical phrase segments.
[0104] The aforementioned song constraint information includes: mode constraint information; the processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the aforementioned song splicing method: determining the average value of the BPM data of multiple musical phrases, adjusting the BPM data of the musical phrases to the average value; determining the song mode of the musical phrases, converting the song mode of the musical phrases to the specified mode in the mode constraint information; and converting the audio format of multiple musical phrases to a target format with a specified sampling rate.
[0105] The aforementioned splicing structure parameters include: a sequence of reference function types for the song to be spliced, the duration percentage of each reference function type in the sequence, and the audio energy score range of the corresponding audio segments of the reference function types. The processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations of the song splicing method: acquiring phrase information corresponding to multiple phrase segments; wherein, the phrase information includes: the identifier information of the phrase segment, the storage path information of the phrase segment, the function type corresponding to the phrase segment in the musical structure of the target song, the lyrics information of the phrase segment, and the audio energy score of the phrase segment; the audio energy score is used to indicate the energy intensity of the audio signal; inputting the phrase information, the song information of the target song, the target template, and preset prompts into the phrase splicing model, and guiding the phrase splicing model based on the structural splicing parameters and preset constraints through the preset prompts, outputting a first result; wherein, the first result includes: the target phrase segments participating in the splicing, and the connection method between the target phrase segments; generating a spliced song based on the target phrase segments and the connection method.
[0106] The aforementioned preset constraints include: eliminating musical phrases whose function type does not match the reference function type; determining musical phrases according to the structural splicing parameters and the preset splicing song duration, and determining the target musical phrases to be spliced; when determining the target musical phrases, prioritizing musical phrases with the same key, and if musical phrases with different keys need to be selected, marking the target musical phrases that need to be transposed; arranging target musical phrases of the same function type in ascending order of audio energy score; and ensuring that the difference in audio energy score between interconnected target musical phrases does not exceed a preset difference threshold.
[0107] The processor in the aforementioned electronic device can perform the following operations of the song splicing method by executing computer-executable instructions: discretely sample the audio signal corresponding to the musical phrase fragment to obtain multiple sampled signals; determine the root mean square value of the amplitude values of the multiple sampled signals; and map the root mean square value to a specified audio energy score interval to obtain the audio energy score of the musical phrase fragment.
[0108] This embodiment also provides a storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the above-described song splicing method.
[0109] The computer-executable instructions stored in the aforementioned storage medium can be executed to perform the following operations in the song splicing method: determining the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text used to describe the theme style of the playlist; the song information includes: song background information and song music attribute information; based on the playlist information and the song information, determining the target template from the candidate theme type templates; wherein, the theme type template contains song constraint information corresponding to the preset theme type and structural splicing parameters of the spliced song to be generated; based on the song constraint information and song information in the target template, selecting the target song from the target playlist and dividing the target song into multiple musical phrases; generating the spliced song based on the multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template.
[0110] This method matches target templates that fit the playlist theme based on playlist information and song information within the playlist. Target songs are then selected based on the song constraints within the target templates, and these target songs are divided into multiple musical phrases. Finally, the musical phrases are spliced together based on the musical phrases, the song information of the target songs, and the structural splicing parameters in the target templates to obtain the spliced song. This method, by extracting and recombining musical phrases from songs within a playlist, helps users quickly perceive the song composition and theme of the playlist, creates a novel auditory experience, effectively enhances the fun of browsing playlists, improves music discovery efficiency, and enhances the user experience.
[0111] The computer-executable instructions stored in the aforementioned storage medium can be executed to perform the following operations in the song splicing method: creating a target playlist and determining the playlist information; retrieving background information of songs in the target playlist from a preset resource library; wherein the background information of the songs includes lyrics information, release year information, artist information, and creation background information; acquiring the audio signals of the songs in the target playlist and determining the song's musical attribute information based on the audio signals; wherein the song's musical attribute information includes at least: song style, song mood and atmosphere, BPM data, instrument information, song key, and temporal distribution information of the beat.
[0112] The computer-executable instructions stored in the aforementioned storage medium can be executed to perform the following operations in the song splicing method: filling the playlist information, song information, and alternative theme type templates into a preset prompt word template to generate prompt words; inputting the prompt words into a large language model to guide the large language model to determine the theme type of the target playlist based on the playlist information and song information, and to determine the target template that matches the theme type of the target playlist from the alternative theme type templates.
[0113] The aforementioned song constraint information includes: BPM constraint information, mode constraint information, genre constraint information, and mood / atmosphere constraint information; the computer-executable instructions stored in the aforementioned storage medium, by executing the computer-executable instructions, can realize the following operations in the aforementioned song splicing method: obtaining the specified BPM interval corresponding to the BPM constraint information, determining the BPM filtering interval based on preset filtering parameters and the specified BPM interval, and determining the songs in the target playlist whose BPM data is within the BPM filtering interval as the first songs; obtaining the mode of the first songs, and determining the songs in the first songs whose mode is consistent with the specified mode in the mode constraint information, and the songs whose mode modulates to the specified mode with a modulation span not exceeding a preset whole tone as the second songs; from the second songs, determining the third songs whose genre is the same as the specified genre in the genre constraint information; obtaining the mood / atmosphere and lyrics information of the third songs, and determining the songs whose mood / atmosphere is consistent with the specified mood / atmosphere in the mood / atmosphere constraint information, or whose lyrics information matches the specified mood / atmosphere, as the target songs; wherein, the number of target songs meets a preset number threshold.
[0114] The computer-executable instructions stored in the aforementioned storage medium can be executed to perform the following operations in the song splicing method: For each target song, obtain the target audio signal corresponding to the target song, input the target audio signal into a pre-trained audio processing model, and output multiple audio segments and their associated information. Each audio segment corresponds to a different functional type in the musical structure of the target song. The associated information of the audio segments includes: the start time, end time, and functional type of the audio segment; and the audio segments are segmented to obtain multiple musical phrases.
[0115] The aforementioned audio segment is an audio segment with lyrics; the computer-executable instructions stored in the aforementioned storage medium, by executing the computer-executable instructions, can realize the following operations in the aforementioned song splicing method to obtain the lyric segment corresponding to the audio segment, and determine the start and end time points of the lyrics corresponding to each line of lyrics in the lyric segment; wherein, the start and end time points of the lyrics include: the start time of the lyrics and the end time of the lyrics; for each line of lyrics, determine the strong beat time points of the audio segment within a preset duration before and after the start and end time points of the lyrics, and cut the audio segment at the first strong beat time point closest to the start time of the lyrics and the second strong beat time point closest to the end time of the lyrics, respectively, to obtain the musical phrase segment corresponding to each line of lyrics.
[0116] The aforementioned audio segment is an audio segment without lyrics; the computer-executable instructions stored in the aforementioned storage medium, by executing the computer-executable instructions, can realize the following operations in the aforementioned song splicing method: determining the duration of a single measure of the target song based on the BPM data of the target song to which the audio segment belongs and the preset number of beats per measure; determining the duration of the musical phrase corresponding to the musical phrase based on the preset number of measures contained in the musical phrase; and segmenting the audio segment according to the musical phrase duration to obtain multiple musical phrase segments.
[0117] The aforementioned song constraint information includes: mode constraint information; computer-executable instructions stored in the aforementioned storage medium, by executing the computer-executable instructions, the following operations in the aforementioned song splicing method can be implemented: determining the average value of BPM data of multiple musical phrase segments, adjusting the BPM data of the musical phrase segments to the average value; determining the song mode of the musical phrase segments, converting the song mode of the musical phrase segments to the specified mode in the mode constraint information; and converting the audio format of multiple musical phrase segments to a target format with a specified sampling rate.
[0118] The aforementioned splicing structure parameters include: a sequence of reference function types for the song to be spliced, the duration percentage of each reference function type in the sequence, and the audio energy score range of the corresponding audio segments of each reference function type; computer-executable instructions stored in the aforementioned storage medium, which, by executing these instructions, enable the following operations in the song splicing method to obtain the phrase information corresponding to multiple phrase segments; wherein, the phrase information includes: the identifier information of the phrase segment, the storage path information of the phrase segment, the function type corresponding to the phrase segment in the musical structure of the target song, the lyrics information of the phrase segment, and the audio energy score of the phrase segment; the audio energy score is used to indicate the energy intensity of the audio signal; the phrase information, the song information of the target song, the target template, and the preset prompt words are input into the phrase splicing model, and the preset prompt words guide the phrase splicing model to output a first result based on the structural splicing parameters and preset constraints; wherein, the first result includes: the target phrase segments participating in the splicing, and the connection method between the target phrase segments; and a spliced song is generated based on the target phrase segments and the connection method.
[0119] The aforementioned preset constraints include: eliminating musical phrases whose function type does not match the reference function type; determining musical phrases according to the structural splicing parameters and the preset splicing song duration, and determining the target musical phrases to be spliced; when determining the target musical phrases, prioritizing musical phrases with the same key, and if musical phrases with different keys need to be selected, marking the target musical phrases that need to be transposed; arranging target musical phrases of the same function type in ascending order of audio energy score; and ensuring that the difference in audio energy score between interconnected target musical phrases does not exceed a preset difference threshold.
[0120] The computer-executable instructions stored in the aforementioned storage medium can be executed to perform the following operations in the song splicing method: discretely sample the audio signal corresponding to the musical phrase segment to obtain multiple sampled signals; determine the root mean square value of the amplitude values of the multiple sampled signals; and map the root mean square value to a specified audio energy score interval to obtain the audio energy score of the musical phrase segment.
[0121] The computer program products of the song splicing method, apparatus, electronic device and storage medium provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0123] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0124] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0125] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0126] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for splicing songs, characterized in that, The method includes: Determine the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text used to describe the theme and style of the playlist; the song information includes: song background information and song music attribute information; Based on the playlist information and the song information, a target template is determined from the candidate theme type templates; wherein, the theme type template contains song constraint information corresponding to the preset theme type and structural splicing parameters of the spliced song to be generated; Based on the song constraint information in the target template and the song information, target songs are selected from the target playlist and the target songs are divided into multiple musical phrases. Based on the multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template, a spliced song is generated.
2. The method according to claim 1, characterized in that, The step of determining the playlist information of the target playlist and the song information of the songs in the target playlist includes: Create the target playlist and determine the playlist information of the target playlist; Retrieve background information of songs in the target playlist from a pre-set resource library; wherein, the background information of the songs includes lyrics, release year, artist information, and creative background information; The audio signals of the songs in the target playlist are obtained, and the music attribute information of the songs is determined based on the audio signals; wherein the music attribute information of the songs includes at least: song style, song mood and atmosphere, BPM data, instrument information, song key, and temporal distribution information of the beat.
3. The method according to claim 1, characterized in that, The step of determining the target template from the candidate theme type templates based on the playlist information and the song information includes: The playlist information, the song information, and the alternative theme type templates are filled into the preset prompt word template to generate prompt words; The prompt words are input into the large language model, which guides the large language model to determine the theme type of the target playlist based on the playlist information and the song information, and to determine the target template that matches the theme type of the target playlist from the candidate theme type templates.
4. The method according to claim 1, characterized in that, The song constraint information includes: BPM constraint information, mode constraint information, genre constraint information, and emotional atmosphere constraint information; The step of selecting target songs from the target playlist based on the song constraint information in the target template and the song information includes: Obtain the specified BPM interval corresponding to the BPM constraint information, determine the BPM filtering interval based on the preset filtering parameters and the specified BPM interval, and determine the songs in the target playlist whose BPM data is within the BPM filtering interval as the first songs. Obtain the key of the first song, and determine the songs in the first song whose key is consistent with the specified key in the key constraint information, and the songs whose key is transposed to the specified key with a transposition span not exceeding a preset number of whole tones, as the second songs; From the second song, determine a third song whose musical style is the same as the specified musical style in the musical style constraint information; The emotional atmosphere and lyrics information of the third song are obtained, and the songs whose emotional atmosphere is consistent with the specified emotional atmosphere in the emotional atmosphere constraint information, or whose lyrics information matches the specified emotional atmosphere, are determined as target songs; wherein the number of target songs meets a preset number threshold.
5. The method according to claim 1, characterized in that, The step of dividing the target song into multiple musical phrases includes: For each target song, the target audio signal corresponding to the target song is obtained, and the target audio signal is input into a pre-trained audio processing model. Multiple audio segments and their association information are output. Each audio segment corresponds to a different function type in the musical structure of the target song. The association information of the audio segments includes: the start time, end time, and function type of the audio segment. The audio segment is divided into multiple musical phrase segments.
6. The method according to claim 1, characterized in that, The splicing structure parameters include: a reference function type sequence of the spliced song to be generated, the duration percentage of the reference function type in the reference function type sequence, and the audio energy score range of the functional audio segment corresponding to the reference function type; The step of generating a spliced song based on the multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template includes: Obtain the musical phrase information corresponding to each of the multiple musical phrase fragments; wherein, the musical phrase information includes: the identification information of the musical phrase fragment, the storage path information of the musical phrase fragment, the function type corresponding to the musical phrase fragment in the musical structure of the target song, the lyrics information of the musical phrase fragment, and the audio energy score of the musical phrase fragment; the audio energy score is used to indicate the energy intensity of the audio signal; The musical phrase information, the song information of the target song, the target template, and the preset prompt words are input into the musical phrase splicing model. The preset prompt words guide the musical phrase splicing model to output a first result based on the structural splicing parameters and preset constraints. The first result includes: the target musical phrase fragments participating in the splicing, and the connection method between the target musical phrase fragments. Based on the target musical phrase fragment and the connection method, a spliced song is generated.
7. The method according to claim 6, characterized in that, The preset constraints include: Remove musical phrases whose function type does not match the reference function type; The musical phrase fragments are determined according to the structure splicing parameters and the preset splicing song duration, and the target musical phrase fragments to be spliced are determined. When determining the target musical phrase fragment, musical phrase fragments with the same key are selected first. If musical phrase fragments with different keys need to be selected, the target musical phrase fragments that need to be transposed are marked with transposition. The target musical phrase fragments of the same functional type are arranged in ascending order of the audio energy score; The difference in audio energy scores between the interconnected target musical phrase segments does not exceed a preset difference threshold.
8. A song splicing device, characterized in that, The device includes: The first determining module is used to determine the playlist information of the target playlist and the song information of the songs in the target playlist; wherein, the playlist information includes: the playlist name and descriptive text for describing the theme and style of the playlist; the song information includes: song background information and song music attribute information; The second determining module is used to determine a target template from the candidate theme type templates based on the playlist information and the song information; wherein, the theme type template includes song constraint information corresponding to a preset theme type and structural splicing parameters of the spliced song to be generated; The first segmentation module is used to filter out target songs from the target playlist based on the song constraint information in the target template and the song information, and to segment the target songs into multiple musical phrases. The first splicing module is used to generate a spliced song based on the multiple musical phrases, the song information of the target song, and the structural splicing parameters in the target template.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the song splicing method according to any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the song splicing method according to any one of claims 1-7.