Methods for generating cover songs, computer equipment, storage media, and software products
By identifying and replacing reverb and harmony segments in the original vocals, and combining this with the user's vocal characteristics to generate cover songs, the problem of low-quality cover songs has been solved, achieving high-quality and personalized cover effects.
Patent Information
- Application Number
- CN202411474524.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-10-22
AI Technical Summary
The existing technology produces cover songs with quality problems such as muffled notes, strange sounds, or unbalanced volume, resulting in low quality cover songs.
By identifying the vocal segments to be repaired in the original singer's voice, replacing them with target vocal segments that do not have reverb and harmony, extracting singing skills features, and combining them with the user's timbre features, a cover song is generated.
It improves the quality of cover songs, achieves personalized timbre and high-quality cover song effects, and solves the sound quality problem of cover songs in existing technologies.
Smart Images

Figure CN119296497B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a method for generating cover songs, a computer device, a storage medium, and a program product. Background Art
[0002] With the rapid development of artificial intelligence technology, the music production field has also ushered in a new era, especially in the areas of song generation and cover songs. Previously, cover songs relied on the singer's vocal skills and complex audio processing techniques, requiring a significant investment of time and effort. Traditionally, AI technology could process the original song's audio to generate cover versions simulating other singers; however, these generated cover versions might suffer from issues such as muffled vocals, unusual sounds, or unbalanced volume, resulting in lower quality. Summary of the Invention
[0003] Therefore, it is necessary to provide a method, computer equipment, storage medium, and program product for generating cover songs that can improve the quality of the generated cover songs, in response to the above-mentioned technical problems.
[0004] Firstly, this application provides a method for generating cover songs, including:
[0005] For the target song selected by the user for a cover, identify the vocal segment to be repaired in the original vocals of the target song; the vocal segment to be repaired is a vocal segment with reverb and / or harmony;
[0006] Obtain a target vocal segment that matches the vocal segment to be repaired; the target vocal segment is a vocal segment without reverb and harmony, and the target vocal segment and the vocal segment to be repaired correspond to the same singing content in the target song;
[0007] The target vocal segment is used to replace the vocal segment to be repaired in the original vocal segment to obtain the repaired original vocal segment. The singing characteristics of the repaired original vocal segment are extracted to obtain the singing characteristics.
[0008] Extract the timbre features from the user-provided audio material, and generate a cover version of the target song based on the timbre features and the vocal guide features; the cover song is a song generated by simulating the user's timbre when singing the target song.
[0009] In one embodiment, identifying the vocal segment to be repaired in the original vocals of the target song includes:
[0010] Interference segments in the original vocals of the target song were detected; the interference segments were vocal segments containing reverb and / or harmony.
[0011] For any of the aforementioned interference segments, if the interference segment partially falls within the non-singing time period of the original vocals, or if the interference segment completely falls within the singing time period of the original vocals and the endpoint timestamp of the interference segment is not aligned with the endpoint timestamp of the sound unit, the two endpoint timestamps of the interference segment are adjusted to obtain a corrected interference segment; the two endpoint timestamps of the corrected interference segment are both aligned with the endpoint timestamp of one sound unit and the corrected interference segment does not include the non-singing time period;
[0012] Based on the corrected interference fragment, the vocal fragment to be repaired in the original vocals of the target song is obtained.
[0013] In one embodiment, adjusting the timestamps of the two endpoints of the interference fragment to obtain the corrected interference fragment includes:
[0014] If the start timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit in the original vocal recording, then the start timestamp of the interference segment is corrected to the start timestamp of any vocal unit to obtain the corrected interference segment.
[0015] If the start timestamp of the interference segment falls within the non-singing time period, the start timestamp of the interference segment is corrected to the start timestamp of the next pronunciation unit adjacent to the non-singing time period, thus obtaining the corrected interference segment.
[0016] If the end timestamp of the interference segment falls between the two endpoint timestamps of any of the pronunciation units, then the end timestamp of the interference segment is corrected to the end timestamp of any of the pronunciation units to obtain the corrected interference segment.
[0017] If the end timestamp of the interference segment falls within the non-singing time period, the end timestamp of the interference segment is corrected to the end timestamp of the previous vocal unit adjacent to the non-singing time period, thus obtaining the corrected interference segment.
[0018] In one embodiment, obtaining the target voice segment that matches the voice segment to be repaired includes:
[0019] If there is a clean segment in the original vocals that matches the vocal segment to be repaired, then the clean segment is determined as the target vocal segment that matches the vocal segment to be repaired; the clean segment is a vocal segment without reverb and harmony.
[0020] If there is no pure segment in the original vocals that matches the vocal segment to be repaired, then a target vocal segment that matches the vocal segment to be repaired is determined from the candidate vocal set.
[0021] In one embodiment, determining the target voice segment that matches the voice segment to be repaired from the candidate voice set includes:
[0022] Obtain a candidate voice set; the candidate voice set includes multiple candidate voices;
[0023] Based on at least one of the singer's gender, pitch rise / fall, and loudness of the candidate voices, a target vocal segment matching the vocal segment to be repaired is determined from the plurality of candidate voices.
[0024] In one embodiment, obtaining the candidate voice set includes:
[0025] Obtain multiple dry vocals corresponding to the target song; the dry vocals are non-original vocals that do not have reverb or harmony and whose audio length is the same as the audio length of the original vocals;
[0026] Based on the sound quality index information of the dry sound, the multiple candidate sounds are determined from the multiple dry sounds to obtain the candidate sound set; the sound quality index information includes at least one of loudness, clarity, and pitch.
[0027] In one embodiment, before identifying the vocal segment to be repaired in the original vocals of the target song, the method further includes:
[0028] The accompaniment and vocals of the target song are separated to obtain the original vocals after separation;
[0029] Based on the positional information of each pronunciation unit in the lyrics of the target song, the non-singing time periods between each line in the lyrics are identified; the positional information includes the endpoint timestamp of the pronunciation unit and the line identifier information of the pronunciation unit in the lyrics;
[0030] The audio corresponding to the non-singing time period in the separated original vocals is replaced with blank audio to obtain the original vocals of the target song.
[0031] In one embodiment, replacing the audio corresponding to the non-singing time period in the separated original vocals with blank audio to obtain the original vocals of the target song includes:
[0032] Determine the beginning and end reserved time periods in the non-singing time periods between any two rows in each of the rows;
[0033] The non-singing time intervals between any two lines, excluding the reserved time intervals at the beginning and the reserved time intervals at the end, are used as the trimming time intervals;
[0034] The audio corresponding to the trimmed time segment in the separated original vocals is replaced with blank audio to obtain the original vocals of the target song.
[0035] Secondly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0036] Thirdly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0037] Fourthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0038] The aforementioned method for generating cover songs, computer equipment, storage media, and program products, for a user-selected target song for cover, identify vocal segments with reverberation and harmony from the original vocals of the target song as vocal segments to be repaired; determine a target vocal segment that matches the vocal segment to be repaired, which is a vocal segment in the target song with the same singing content as the vocal segment to be repaired, and which does not have reverberation and / or harmony; replace the vocal segment to be repaired in the original vocals with the target vocal segment to obtain the repaired original vocals, and extract singing features from the repaired original vocals to obtain the leading singing features; and based on the sound provided by the user... The timbre and vocal features extracted from the source material are used to generate a cover version of the target song. This cover version is generated by simulating the user's timbre while singing the target song. By identifying and repairing vocal segments in the original singer's voice that have reverberation and harmony issues, and replacing the vocal segments to be repaired with high-quality target vocal segments that match the vocal segments to be repaired, the original singer's voice is finely optimized by optimizing the reverberation and harmony defects in the original singer's voice. This process preserves the original singer's vocal characteristics. Furthermore, by extracting the timbre features of the user's voice, a high-quality cover version with the user's personalized timbre is generated, thus improving the overall quality of the generated cover version. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1This is an application environment diagram of a cover song generation method in one embodiment;
[0041] Figure 2 This is a flowchart illustrating a method for generating cover songs in one embodiment;
[0042] Figure 3 This is a schematic diagram of a user interface for generating cover songs in one embodiment;
[0043] Figure 4 This is a schematic diagram of another user interface for generating cover songs in another embodiment;
[0044] Figure 5 This is a flowchart illustrating the original vocal trimming process in one embodiment.
[0045] Figure 6 This is a schematic diagram of a modified interference segment in one embodiment;
[0046] Figure 7 This is a flowchart illustrating a process for determining the original vocals and identifying the vocal segments to be repaired within the original vocals, as described in one embodiment.
[0047] Figure 8 This is a flowchart illustrating a process for generating a set of candidate voices in one embodiment.
[0048] Figure 9 This is a flowchart illustrating a process for generating lead vocal features in one embodiment.
[0049] Figure 10 This is a flowchart illustrating another method for generating cover songs in one embodiment;
[0050] Figure 11 This is a structural block diagram of a cover song generation device in one embodiment;
[0051] Figure 12 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0053] The cover song generation method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server. For the target song selected by the user for a cover version, server 104 identifies the vocal segment to be repaired in the original vocals of the target song; the vocal segment to be repaired is a vocal segment with reverb and / or harmony; server 104 obtains a target vocal segment that matches the vocal segment to be repaired; the target vocal segment is a vocal segment without reverb and harmony, and the target vocal segment and the vocal segment to be repaired correspond to the same singing content in the target song; server 104 uses the target vocal segment to replace the vocal segment to be repaired in the original vocals to obtain the repaired original vocals, and extracts the singing characteristics of the repaired original vocals to obtain the guiding singing characteristics; server 104 extracts the timbre characteristics from the sound material provided by the user, and terminal 102 or server 104 generates a cover version of the target song based on the timbre characteristics and guiding singing characteristics; the cover song is a song generated by simulating the user's timbre in singing the target song. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0054] In an exemplary embodiment, Figure 2 As shown, a method for generating cover songs is provided, which can be applied to... Figure 1 Taking server 104 as an example, the explanation includes:
[0055] Step S202: For the target song selected by the user for a cover, identify the vocal segments to be repaired in the original singer's vocals of the target song.
[0056] Among them, the vocal segments to be repaired are those with reverberation and / or harmony.
[0057] The target song can refer to the original song that the user chooses to cover, which includes the complete accompaniment and vocals.
[0058] The original vocals can refer to the voice of the singer in the original recording of the target song.
[0059] The vocal segments to be repaired can be the parts of the original vocals that have been interfered with by harmonies and reverb. Reverb refers to sound reflections and delays in the original vocals caused by the recording environment or other technical reasons, which may affect the clarity of the vocals. Harmony refers to other vocal parts or tones that accompany the original vocals, usually added to increase the richness of the music, but may also affect the clarity of the vocals. The presence of harmony and reverb can prevent the subsequent generation of high-quality vocal characteristics, so it is necessary to minimize the reverb and harmony in the original vocals.
[0060] Optionally, the server can remove reverb and harmony from the original vocals using methods such as inverse filtering, spectral subtraction, and recurrent neural networks. However, even after removing reverb and harmony using these methods, reverb and harmony will still remain in the original vocals, making complete removal impossible. Therefore, a binary classification detection model can be used to detect segments of the original vocals that still retain reverb and harmony after one removal process, and these segments are designated as vocal segments to be repaired. Thus, in this embodiment, the vocal segment to be repaired can refer to a vocal segment that still retains reverb and harmony after one removal process. The binary classification detection model can be an artificial intelligence model trained using audio segments with reverb and harmony as positive samples and audio segments without reverb and harmony as negative samples.
[0061] For the convenience of those skilled in the art, Figure 3 An example of a user interface diagram for generating cover songs is provided. In specific implementation, the user can open the terminal's music application, which includes the function of creating cover songs. The terminal's music application can display a music library interface 300, which can include multiple songs 302 available for the user to listen to. The music library interface also includes song creation prompts 304, which can be used to instruct the user to create cover songs. The user can select the target song they want to cover from the multiple songs 302. The terminal can respond to the selection operation of the target song and enter the song creation interface 306. Optionally, the selection operation of the target song can include, but is not limited to, clicking, long-pressing, or other triggering operations on the control of the target song. The song creation interface 306 includes a cover song selection control 308 and a cover song generation entry 310. The cover song selection control 308 can be used to select and generate high-quality cover songs in this application embodiment. The cover song generation entry 310 can be used to trigger the creation of cover songs. High-quality cover songs are created through the cover song generation method of this application embodiment, thus meeting the user's needs for professionally sung and high-quality cover songs.
[0062] For the convenience of those skilled in the art, Figure 4An example of another user interface diagram for generating cover songs is provided. In a specific implementation, the terminal's music application can display a video recommendation interface 400. When browsing videos displayed on the video recommendation interface 400, users can create high-quality cover songs for any song in a favorite video using the cover song generation method of this embodiment. The video recommendation interface 400 may include a target song confirmation entry 402. Users can select a song from a favorite video as the target song and perform a user-triggered operation on the target song confirmation entry 402 corresponding to that song. The target song confirmation entry 402 is used to enter the cover song confirmation interface 404. The cover song confirmation interface 404 includes a cover song confirmation entry 406, which is used to trigger the generation of a cover song. The terminal can respond to the user-triggered operation on the cover song confirmation entry 406 and create high-quality cover songs using the cover song generation method of this embodiment. Therefore, users can create a cover version with their own voice for any song they like, whether from a music library or a video, using the method of this embodiment, thereby meeting users' personalized cover song creation needs.
[0063] The above solutions can provide users with a higher quality and more personalized cover song production experience.
[0064] Step S204: Obtain the target voice segment that matches the voice segment to be repaired.
[0065] The target vocal segment is a vocal segment without reverb and harmony, and the target vocal segment and the vocal segment to be repaired correspond to the same singing content in the target song. In other words, the target vocal segment can be a segment of the original singer's voice without reverb and harmony, or a segment of a non-original singer's voice without reverb and harmony. The non-original singer's voice can refer to the voice of a non-original singer performing the target song. The same singing content can include the same lyrics, the same rhythm, the same pitch, etc.
[0066] Optionally, the target vocal segment and the vocal segment to be repaired corresponding to the same singing content in the target song can mean that the target vocal segment and the vocal segment to be repaired are in the same singing time period in the target song, or it can mean that the target vocal segment and the vocal segment to be repaired are in different singing time periods in the target song but correspond to the same singing content. For example, two choruses in a song may have the same singing content. The target vocal segment of one chorus can be used to replace the vocal segment to be repaired in the other chorus, and the replaced part corresponds to the same singing content.
[0067] In practice, when the server identifies the vocal segment to be repaired in the original vocals of a target song, it can determine the corresponding performance time period of the vocal segment in the target song and the content of the song during that performance time period. Then, it acquires the original vocals and the corresponding non-singing vocals. The non-singing vocals can be sung by a skilled non-original singer, whose voice is as similar as possible to the original singer's, with good sound quality and minimal background noise. From these original and non-singing vocals, it identifies the target vocal segment with the same performance content as the vocal segment to be repaired in the target song. If multiple segments from these original and non-singing vocals match the vocal segment to be repaired, the highest quality segment with the least background noise and other desirable characteristics can be selected as the target vocal segment. Alternatively, the target vocal segment can be obtained first from the original vocals; if no suitable segment exists from the original vocals, then the target vocal segment can be obtained from the non-singing vocals.
[0068] Step S206: Replace the vocal segment to be repaired in the original vocal segment with the target vocal segment to obtain the repaired original vocal segment, and extract the singing characteristics of the repaired original vocal segment to obtain the guiding vocal characteristics.
[0069] The lead vocal features carry the original singer's vocal characteristics, preserving these characteristics without reverberation or harmonic interference. These vocal characteristics may include, but are not limited to, features reflecting the original singer's pitch, breath control, melody, rhythm, singing technique, and emotional expression. Optionally, these features may include phonetic posterior probability maps (PPG) and fundamental frequency (F0). PPG reflects the temporal distribution and phoneme category information in the voice, and can be used to reflect vocal clarity, rhythm, melody, and other vocal qualities; fundamental frequency is the basic vibration frequency of the voice, and can be used to reflect pitch, intonation, range, vibrato, and other vocal qualities.
[0070] In practice, the original vocals may contain multiple segments requiring restoration. For example, the first segment might correspond to the period from 1 minute 30 seconds to 1 minute 45 seconds in the target song, while the second segment might correspond to the period from 2 minutes 40 seconds to 2 minutes 50 seconds. Therefore, for each segment requiring restoration, a corresponding target vocal segment can be selected. This allows replacing each segment with its corresponding target vocal segment, resulting in the restored original vocals. Finally, based on the phoneme posterior probability map (PPG) or fundamental frequency (F0) extracted from the restored original vocals, the guiding vocal features are determined.
[0071] By finely optimizing the original vocals, this method can solve problems such as muffled notes, strange sounds, and unbalanced volume in existing cover song generation methods. Furthermore, users can choose any song they want to cover, satisfying their production needs in terms of both personalization and quality.
[0072] Step S208: Extract the timbre features from the audio material provided by the user, and generate a cover version of the target song based on the timbre features and the vocal features.
[0073] Among them, the cover songs are songs generated by simulating the user's voice singing the target song.
[0074] Optionally, the user-provided audio material can be an audio clip of the user singing any song.
[0075] Timbre can refer to the unique quality or characteristics of a sound. It is the main factor that distinguishes one sound from another and can be used to identify the characteristics of different human voices.
[0076] In practice, the terminal can collect user audio materials or obtain user-uploaded audio materials and store them in a Clustered File System (CFS) shared network disk so that the user's audio materials can be shared among multiple servers, reducing development complexity. The server uses Voice Activity Detection (VAD) to detect whether the effective voice duration in the user's audio materials reaches a duration threshold (e.g., 20 or 30 seconds). Audio materials with effective voice durations below the threshold can be discarded. Noise reduction processing is performed on user audio materials with effective voice durations that meet the duration threshold to minimize the interference of noise caused by environment, equipment, abnormal operations during recording, etc., on subsequent calculations. Timbre features are extracted from the user-provided audio materials using a trained timbre extraction model. Optionally, the timbre extraction model can include, but is not limited to, deep learning models such as convolutional neural networks, variational autoencoders, recurrent neural networks, and models based on self-attention mechanisms. Taking the convolutional neural network model as an example, by applying convolution operations to the spectrogram of the audio signal (such as the Mel spectrogram), local patterns and features related to timbre, such as frequency distribution and harmonic structure, can be effectively extracted.
[0077] After extracting the timbre features from the user-provided audio material, the terminal or server can generate a cover vocal for the target song based on these timbre and vocal characteristics. Then, the cover vocal and the accompaniment of the target song are professionally mixed to generate a high-quality cover song. Optionally, the terminal can input the timbre and vocal characteristics into an artificial intelligence model such as a vocal conversion model. The AI model can then generate a cover vocal that simulates the user's timbre while retaining the singing characteristics from the vocal characteristics. After acquiring the cover vocal, it can be professionally mixed with the accompaniment of the target song to generate a high-quality cover song that simulates the user's timbre.
[0078] In the aforementioned method for generating cover songs, for the target song selected by the user, the vocal segments containing reverb and harmony are identified from the original vocals of the target song as the vocal segments to be repaired; a target vocal segment matching the vocal segment to be repaired is determined, which is a vocal segment with the same singing content in the target song corresponding to the vocal segment to be repaired, and which does not contain reverb and / or harmony; the vocal segment to be repaired in the original vocals is replaced with the target vocal segment to obtain the repaired original vocals, and singing features are extracted from the repaired original vocals to obtain the leading vocal features; based on the vocal materials provided by the user... Based on timbre and vocal characteristics, a cover version of the target song is generated. This cover version is generated by simulating the user's timbre while singing the target song. By identifying and repairing vocal segments in the original singer's voice that have reverberation and harmony issues, and replacing the vocal segments to be repaired with high-quality target vocal segments that match the vocal segments to be repaired, the original singer's voice is finely optimized by optimizing the reverberation and harmony defects in the original singer's voice. This process preserves the original singer's vocal characteristics. Furthermore, by extracting the timbre characteristics of the user's voice, a high-quality cover version with the user's personalized timbre is generated, thus improving the overall quality of the generated cover version.
[0079] In an exemplary embodiment, before identifying the vocal segment to be repaired in the original vocals of the target song, the method further includes: separating the accompaniment and vocals of the target song to obtain the separated original vocals; identifying the non-singing time periods between lines in the lyrics based on the position information of each pronunciation unit in the lyrics of the target song; the position information includes the endpoint timestamp of the pronunciation unit and the line identifier information of the pronunciation unit in the lyrics; and replacing the audio corresponding to the non-singing time periods in the separated original vocals with blank audio to obtain the original vocals of the target song.
[0080] In practice, the server can perform spectral analysis on the target song, using the different features of phonemes and accompaniment on the spectrogram to separate the accompaniment and vocals, or it can use a deep learning model to separate the accompaniment and vocals, obtaining the separated original vocals. The server can then parse the lyrics in the separated original vocals to obtain the positional information of each pronunciation unit in the lyrics of the target song. Specifically, this can be done by parsing the timestamps of the lyrics word by word, obtaining the parsing results. The parsing results include the positional information of each pronunciation unit in the lyrics, where the positional information includes the endpoint timestamp of the pronunciation unit and the line identifier information of the pronunciation unit in the lyrics. The endpoint timestamp includes the start timestamp corresponding to the pronunciation start point of the pronunciation unit in the target song and the end timestamp corresponding to the pronunciation end point of the pronunciation unit in the target song.
[0081] In a specific implementation, the parsing result can be represented as the following structure. :
[0082] ;
[0083] Where start is the start timestamp of the pronunciation unit, end is the end timestamp of the pronunciation unit, word is the content of the pronunciation unit, and line is the line identifier of the pronunciation unit in the lyrics. Each group This refers to the positional information of a vocal unit.
[0084] Optionally, if the target song is a Chinese song, the pronunciation unit can be a single character; if the target song is an English song, the pronunciation unit can be a single word.
[0085] Since the lyrics of a target song can be divided into multiple lines, each line may include multiple pronunciation units or a sentence composed of multiple pronunciation units. Therefore, the line identifier information of a pronunciation unit in the lyrics is used to indicate which line the pronunciation unit is located on.
[0086] In practice, the terminal identifies the non-singing time periods between lines of lyrics based on the position information of each pronunciation unit. The non-singing time periods between lines can be the time period between the end timestamp of the last pronunciation unit in the nth line and the start timestamp of the first pronunciation unit in the (n+1)th line.
[0087] The server replaces the audio corresponding to the non-singing time periods in the separated original vocals with blank audio. This can be done by first trimming the audio corresponding to the non-singing time periods in the separated original vocals, and then using blank audio to fill in the trimmed parts of the separated original vocals. This ensures that the audio length of the original vocals obtained after trimming is consistent with that of the original vocals before trimming. This ensures that there is no reverb or harmony interference in the non-singing time periods, thereby minimizing the reverb and harmony in the original vocals in the early stages and helping to improve the quality of the subsequently generated cover songs.
[0088] Furthermore, in an exemplary embodiment, replacing the audio corresponding to the non-singing time periods in the separated original vocals with blank audio to obtain the original vocals of the target song may include: determining the beginning and end reserved time periods in the non-singing time periods between any two lines in each row; taking the other time periods in the non-singing time periods between any two lines, excluding the beginning and end reserved time periods, as the trimming time periods; and replacing the audio corresponding to the trimming time periods in the separated original vocals with blank audio to obtain the original vocals of the target song.
[0089] For the convenience of those skilled in the art, Figure 5 An example flowchart of the original vocal trimming process is provided. Optionally, the beginning and end reserved time periods can be 100 milliseconds or 200 milliseconds. By retaining the beginning and end reserved time periods in the non-singing time periods, it is ensured that these reserved time periods will not be trimmed, thereby avoiding the cutting off of some trailing notes and sustained notes. This ensures that useful vocals are not cut off when trimming non-singing time periods, improving the integrity and quality of the original vocals and contributing to the improvement of the quality of the subsequently generated cover songs.
[0090] Similarly, after trimming the audio corresponding to the time segment in the separated original vocals, blank audio can be used to fill in the trimmed part of the separated original vocals to ensure that the audio length of the original vocals obtained after trimming is consistent with that of the separated original vocals before trimming. This ensures that there is no reverb or harmony interference during non-singing time segments, and also avoids cutting out useful trailing notes, prolonged notes, and other vocals, which helps to improve the quality of the subsequently generated cover songs.
[0091] In an exemplary embodiment, identifying the vocal segment to be repaired in the original vocals of a target song includes: detecting interference segments in the original vocals of the target song; the interference segments are vocal segments with reverberation and / or harmony; for any interference segment, if the interference segment partially falls within the non-singing time period of the original vocals, or if the interference segment completely falls within the singing time period of the original vocals and the endpoint timestamp of the interference segment is not aligned with the endpoint timestamp of the sound unit, adjusting the two endpoint timestamps of the interference segment to obtain a corrected interference segment; the two endpoint timestamps of the corrected interference segment are both aligned with the endpoint timestamp of a sound unit and the corrected interference segment does not include the non-singing time period; based on the corrected interference segment, the vocal segment to be repaired in the original vocals of the target song is obtained.
[0092] In practice, the server can remove reverb and harmony from the original vocals using methods such as inverse filtering, spectral subtraction, and recurrent neural networks. However, even after removing reverb and harmony using these methods, reverb and harmony will still remain in the original vocals, making complete removal impossible. Therefore, a binary classification detection model can be used to detect interference segments in the original vocals that still retain reverb after removing one layer of reverb and harmony. The detected interference segments can be represented by the following structure. express:
[0093] ;
[0094] Where "begin" is the start timestamp of the interfering segment, and "finish" is the end timestamp of the interfering segment, for each group... This corresponds to one interference segment.
[0095] For the convenience of those skilled in the art, Figure 6 An example diagram illustrating the correction of interfering segments is provided. Here, "the sound of the sea crying" represents the original vocal performance time segment, with each word considered a single pronunciation unit. Interfering segment A partially falls within the non-singing time segment of the original vocal performance; interfering segment B does not fall within the non-singing time segment but is not aligned with the endpoint timestamps of any pronunciation units; and interfering segment C partially falls within the non-singing time segment of the original vocal performance. Because the endpoint timestamps of these interfering segments are not aligned with the endpoint timestamps of any pronunciation units, this may cause the replaced segment to not smoothly and naturally connect with the unreplaced segment in the original vocal performance when the segment to be repaired is subsequently replaced. Therefore, the interfering segments need to be corrected to obtain corrected interfering segments, which are then used as the vocal segment to be repaired.
[0096] The two endpoint timestamps of the interference segment include the start timestamp of the interference segment's starting point in the target song and the end timestamp of the interference segment's ending point in the target song.
[0097] In practice, for any interfering segment, if the interfering segment falls entirely within the non-singing time period of the original vocals, the interfering segment can be discarded. Specifically, in one of the exemplary embodiments described above, the non-singing time periods in the original vocals are cropped and replaced with blank audio. Therefore, if the interfering segment falls entirely within the non-singing time period of the original vocals, it does not need to be considered as a vocal segment to be repaired.
[0098] In one embodiment, adjusting the two endpoint timestamps of the interference segment to obtain a corrected interference segment may include: if the start timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit in the original vocal performance, then the start timestamp of the interference segment is corrected to the start timestamp of that vocal unit, resulting in a corrected interference segment; if the start timestamp of the interference segment falls within a non-singing time period, then the start timestamp of the interference segment is corrected to the start timestamp of the next vocal unit adjacent to the non-singing time period, resulting in a corrected interference segment; if the end timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit, then the end timestamp of the interference segment is corrected to the end timestamp of that vocal unit, resulting in a corrected interference segment; if the end timestamp of the interference segment falls within a non-singing time period, then the end timestamp of the interference segment is corrected to the end timestamp of the previous vocal unit adjacent to the non-singing time period, resulting in a corrected interference segment.
[0099] In practice, the original vocal recording may contain multiple vocal segments that need to be repaired. These multiple vocal segments can be used to generate a list of vocal segments to be repaired. This list can be represented by the following structure. :
[0100] ;
[0101] Wherein, `initate` is the start timestamp of the vocal segment to be repaired, and `terminate` is the end timestamp of the vocal segment to be repaired. Each group... For each vocal segment to be repaired, this list of vocal segments can be used to efficiently determine the corresponding singing time period of the vocal segment in the original vocal recording.
[0102] The above solution allows the replaced vocal segments to seamlessly and naturally connect with other unreplaced vocal segments during subsequent replacement, thereby improving the fluency and quality of the generated vocal features and ultimately enhancing the quality of the generated cover songs.
[0103] For the convenience of those skilled in the art, Figure 7This document provides an example flowchart for determining the original vocals and identifying vocal segments within them that require repair. As can be seen, when the server determines a target song a user wants to cover from the music library, it can query the database for pre-made vocal guide features corresponding to the target song. The process of pre-making vocal guide features can involve separating the accompaniment and vocals of any target song in the music library, removing reverb and harmony, and detecting any remaining reverb and harmony after removal to obtain the reverb and harmony detection results, i.e., the aforementioned interference segments. Next, the document analyzes the word-by-word timestamps of the original vocal lyrics, thereby trimming non-singing time segments from the original vocals and replacing these non-singing time segments with blank audio to obtain the original vocals. Simultaneously, it corrects the reverb and harmony detection results to obtain a list of vocal segments requiring repair, and replaces these segments with the target vocal segments to obtain the repaired original vocals. Finally, it determines the vocal guide features based on the singing characteristics extracted from the repaired original vocals.
[0104] Optionally, based on the corrected interference fragments, the vocal fragments to be repaired in the original vocals of the target song can be obtained. This can include: if any corrected interference fragment is several pronunciation units in the target line of lyrics, then the fragments corresponding to the target line of lyrics are all taken as vocal fragments to be repaired in the original vocals of the target song. The target line can be the line in the lyrics where any corrected interference fragment is located. In other words, for the sake of natural transitions in breath control, volume, etc., if the interference fragment is a few words in a certain line, the fragments corresponding to the entire line of lyrics can be taken as vocal fragments to be repaired. Subsequently, more words can be replaced or the entire line of lyrics containing those words can be replaced, thereby generating smooth, natural, and high-quality vocal characteristics, thus generating a high-quality cover song.
[0105] In another exemplary embodiment, obtaining a target vocal segment that matches the vocal segment to be repaired includes: if there is a clean segment in the original vocal recording that matches the vocal segment to be repaired, then the clean segment is determined as the target vocal segment that matches the vocal segment to be repaired; the clean segment is a vocal segment without reverb and harmony; if there is no clean segment in the original vocal recording that matches the vocal segment to be repaired, then the target vocal segment that matches the vocal segment to be repaired is determined from the candidate vocal set.
[0106] A clean vocal segment is a vocal segment in the original recording that lacks reverb and harmony. A target song may contain repeated lyrics; for example, the lyrics of the first and second choruses in the target song might be repeated. Assuming the first chorus lacks reverb and harmony, it is a clean vocal segment. The second chorus, however, contains reverb and harmony, making it a vocal segment to be repaired. Since both the first and second choruses contain the same vocal content from the target song, the first chorus can be used as a target vocal segment to match the second chorus. In other words, the first chorus (clean segment) can replace the second chorus (target segment). Because the second chorus has been replaced with the first chorus (lacking reverb and harmony), there is no need to search for a matching target vocal segment from the candidate vocal set. This allows for the maximum utilization of segments from the original vocal recording to generate lead vocal features, improving the smoothness and consistency of subsequent lead vocal feature generation, thereby enhancing the quality of the cover song.
[0107] If no clean vocal segment matching the vocal segment to be repaired exists in the original vocal recording, a target vocal segment matching the vocal segment to be repaired is determined from the candidate vocal set. The candidate vocal set may include vocals of non-original singers who have excellent singing skills and whose voices are similar to those of the original singer of the target song, and whose voices have good quality and low noise.
[0108] In an exemplary embodiment, determining a target vocal segment that matches the vocal segment to be repaired from a set of candidate vocals includes: acquiring a set of candidate vocals; the set of candidate vocals includes multiple candidate vocals; and determining a target vocal segment that matches the vocal segment to be repaired from the multiple candidate vocals based on at least one of the following information: the singer's gender, pitch rise / fall, and loudness.
[0109] Furthermore, in an exemplary embodiment, obtaining a candidate voice set includes: obtaining multiple dry voices corresponding to the target song; the dry voices are non-original vocals that do not have reverberation or harmony and whose audio length is the same as the audio length of the original vocals; determining multiple candidate voices from the multiple dry voices to obtain a candidate voice set based on the sound quality index information of the dry voices; the sound quality index information includes at least one of loudness, clarity, and pitch.
[0110] The obtained dry vocals are high-quality vocal segments obtained by non-original singers performing the target song in its entirety. These high-quality vocal segments can then replace the vocal segments in the original song that need repair.
[0111] In practice, based on the sound quality indicators of the dry vocals, multiple candidate vocals are determined from a pool of dry vocals. This can be done by performing loudness detection on any given dry vocal, assessing the volume level of each lyric segment or word to ensure consistent loudness during performance and avoid impacting overall audio quality due to volume imbalances. For example, the short-time energy (STE) of each word in the dry vocal lyrics can be calculated to measure the energy of each word within each time window, reflecting the loudness of each word in the dry vocals. Based on the average loudness of each word in a lyric segment, the volume balance of the dry vocals can be determined. Vocals with balanced volume are then selected as candidate vocals.
[0112] Alternatively, based on the sound quality indicators of the dry voice, multiple candidate voices can be identified from multiple dry voices. This can be done by performing a lyrics clarity test on any dry voice, or by using an Automatic Speech Recognition (ASR) model to convert the speech in the dry voice into text, and then comparing the converted text with the actual lyrics to determine whether the lyrics are sung correctly. Furthermore, the converted text can be refined into phoneme-level text to detect the clarity of phoneme pronunciation, thereby detecting whether easily confused phoneme pairs such as "s" and "z", or "p" and "b" are sung clearly.
[0113] Alternatively, based on the sound quality index information of the dry vocals, multiple candidate vocals can be determined from multiple dry vocals. This can be done by performing pitch detection on any dry vocal and generating a pitch score. Pitch detection can be done by comparing the pitch of the lyrics in the dry vocal with the standard pitch of the lyrics of the target song to obtain a pitch score. The pitch scores of a single line of lyrics can be filtered, retaining the dry vocal segments with pitch scores higher than the pitch threshold as candidate vocals; or the average pitch score of a section of lyrics can be filtered, discarding the entire dry vocal with an average pitch score lower than the pitch threshold and retaining the dry vocal with an average pitch score higher than the pitch threshold as candidate vocals.
[0114] Alternatively, based on the sound quality indicators of the dry voice, multiple candidate voices can be determined from multiple dry voices. This can be done by using a machine learning model to detect multiple dimensions of the dry voice, such as breath and pitch, to obtain a pleasantness score, and then using dry voices with pleasantness scores higher than the pleasantness threshold as candidate voices.
[0115] The terminal selects a target vocal segment that matches the vocal segment to be repaired from multiple candidate voices based on at least one of the following: the singer's gender, pitch, and loudness. Specifically, the terminal may select candidate voices that are in the same singing time period as the vocal segment to be repaired. Then, from the candidate voices in the same singing time period, it may prioritize selecting candidate voices with the same singer's gender as the vocal segment to be repaired, and candidate voices with the same pitch (key) as the vocal segment to be repaired. Optionally, candidate voices with different keys can be adjusted to match the key of the vocal segment to be repaired. It may also prioritize selecting candidate voices with relatively high loudness, such as those with loudness above a loudness threshold. Finally, it determines the target vocal segment that matches the vocal segment to be repaired from these prioritized candidate voices.
[0116] For the convenience of those skilled in the art, Figure 8 An example flowchart is provided for generating a candidate voice set. The collected dry voice is encoded and uploaded, and the audio length of the dry voice is verified to be consistent with the original singer's voice. Then, loudness detection, clarity detection, and filtering of pitch score and pleasantness score are performed to obtain the candidate voice set.
[0117] For the convenience of those skilled in the art, Figure 9 An example flowchart for generating vocal features is provided. It iterates through each vocal segment in the list to be repaired, first checking if there is a reusable segment in the original vocals. This means checking if there is a clean segment in the original vocals that contains the same content as the target song corresponding to the vocal segment to be repaired. If so, the reusable segment in the original vocals replaces the current vocal segment to be repaired, and the next vocal segment to be repaired is replaced. If not, a set of candidate voices is obtained, and the candidate voices are sorted according to the singer's gender, pitch, and loudness. The highest quality candidate voice (e.g., meeting all the above conditions of singer's gender, pitch, and loudness) is selected as the target vocal segment, and the target vocal segment replaces the vocal segment to be repaired. Then, the next vocal segment to be repaired is replaced, and so on.
[0118] The above method can determine candidate voices from multiple dry voices based on the sound quality index information of the dry voice, generate a set of high-quality candidate voices, and then select suitable target vocal segments from the set of high-quality candidate voices to repair the vocal segments to be repaired. This achieves refined optimization processing of the original vocals, thereby generating high-quality vocal features based on the original vocals with repaired reverb and harmony. Based on the high-quality vocal features and the timbre features in the user's voice material, a high-quality cover song with the user's personalized timbre can be generated.
[0119] Figure 10This is another method for generating cover songs according to an exemplary embodiment, such as... Figure 10 As shown, this method is applied to Figure 1 Taking server 104 as an example, the explanation includes:
[0120] Step S1002: For the target song selected by the user for a cover, detect the interfering segments in the original singer's vocals of the target song.
[0121] The interfering segment is a vocal segment with reverb and / or harmony.
[0122] Step S1004: For any interfering segment, if the interfering segment partially falls within the non-singing time period of the original vocals, or if the interfering segment completely falls within the singing time period of the original vocals and the endpoint timestamp of the interfering segment is not aligned with the endpoint timestamp of the sound unit, adjust the two endpoint timestamps of the interfering segment to obtain the corrected interfering segment.
[0123] The corrected interference fragment has two endpoint timestamps aligned with the endpoint timestamp of a single vocal unit, and the corrected interference fragment does not include non-singing time periods.
[0124] In one embodiment, adjusting the two endpoint timestamps of the interference segment to obtain a corrected interference segment includes: if the start timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit in the original vocal performance, then the start timestamp of the interference segment is corrected to the start timestamp of any vocal unit to obtain a corrected interference segment; if the start timestamp of the interference segment falls within a non-singing time period, then the start timestamp of the interference segment is corrected to the start timestamp of the next vocal unit adjacent to the non-singing time period to obtain a corrected interference segment; if the end timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit, then the end timestamp of the interference segment is corrected to the end timestamp of any vocal unit to obtain a corrected interference segment; if the end timestamp of the interference segment falls within a non-singing time period, then the end timestamp of the interference segment is corrected to the end timestamp of the previous vocal unit adjacent to the non-singing time period to obtain a corrected interference segment.
[0125] Step S1006: Based on the corrected interference fragments, obtain the vocal fragment to be repaired from the original vocals of the target song.
[0126] The vocal segment to be repaired is a sung vocal segment containing reverb and / or harmony.
[0127] In one embodiment, before step S902, the method further includes: separating the accompaniment and vocals of the target song to obtain the separated original vocals; identifying the non-singing time periods between lines in the lyrics based on the position information of each pronunciation unit in the lyrics of the target song; the position information includes the endpoint timestamp of the pronunciation unit and the line identifier information of the pronunciation unit in the lyrics; and replacing the audio corresponding to the non-singing time periods in the separated original vocals with blank audio to obtain the original vocals of the target song.
[0128] Furthermore, in one embodiment, the audio corresponding to the non-singing time period in the separated original vocals is replaced with blank audio to obtain the original vocals of the target song. This includes: determining the beginning and end reserved time periods in the non-singing time periods between any two lines in each row; taking the other time periods between any two lines, excluding the beginning and end reserved time periods, as the trimming time periods; and replacing the audio corresponding to the trimming time periods in the separated original vocals with blank audio to obtain the original vocals of the target song.
[0129] Step S1008: Determine whether there is a clean segment in the original vocals that matches the vocal segment to be repaired. If it exists, proceed to step S910; otherwise, proceed to step S912.
[0130] Among them, the pure segment is a vocal segment without reverb and harmony.
[0131] Step S1010: The clean segment is identified as the target vocal segment that matches the vocal segment to be repaired, and step S1016 is executed.
[0132] Step S1012: Obtain the candidate voice set.
[0133] The candidate voice set includes multiple candidate voices.
[0134] As an exemplary embodiment, obtaining a candidate voice set includes: obtaining multiple dry voices corresponding to the target song; the dry voices are non-original vocals that do not have reverberation or harmony and whose audio length is the same as the audio length of the original vocals; and determining multiple candidate voices from the multiple dry voices to obtain a candidate voice set based on the sound quality index information of the dry voices; the sound quality index information includes at least one of loudness, clarity, and pitch.
[0135] Step S1014: Based on at least one of the singer's gender, pitch rise / fall, and loudness of the candidate voice, determine the target vocal segment that matches the vocal segment to be repaired from multiple candidate voices, and then execute step S1016.
[0136] Step S1016: Replace the vocal segment to be repaired in the original vocal segment with the target vocal segment to obtain the repaired original vocal segment, and extract the singing characteristics of the repaired original vocal segment to obtain the guiding vocal characteristics.
[0137] Step S1018: Extract the timbre features from the audio material provided by the user, and generate a cover version of the target song based on the timbre features and the vocal features.
[0138] Cover songs are generated by simulating the user's voice when singing the target song.
[0139] It should be noted that the specific limitations of the above steps can be found in the specific limitations of a method for generating cover songs mentioned above, and will not be repeated here.
[0140] For example, the method for generating cover songs can be divided into two processes: the first process is to generate vocal lead features, and the second process is to generate the cover song. The first process involves: firstly, obtaining the original vocals of the target song using vocal separation technology; then, using de-harmony and de-reverb algorithms to remove and separate the harmony and reverb of the original vocals, improving the quality of the vocal lead features generated from the original vocals; next, selecting high-quality vocal segments from non-original singers based on multi-dimensional vocal evaluation scores and audio detection technology; then, using a binary classification model to detect the time regions corresponding to severely residual harmony and reverb in the original vocals, and replacing these time regions with high-quality vocal segments from non-original singers; finally, extracting the phoneme posterior probability map (PPG) and fundamental frequency (F0) to generate vocal lead features for subsequent cover song production. The second process is as follows: First, the user's voice material is collected and its features are extracted and trained to obtain a timbre extraction model. The timbre features of the user are extracted through the timbre extraction model. Then, the cover vocals are generated based on the timbre features and the lead vocal features. Then, the cover vocals and the accompaniment of the target song are professionally mixed to generate professional and high-quality cover songs.
[0141] In summary, existing cover song generation methods rely on vocal characteristics, and the available song library is relatively limited. This application's cover song generation method utilizes multi-dimensional vocal evaluation scores, audio detection technology for selecting vocal segments, original vocal extraction based on vocal separation, audio noise reduction algorithms, harmony and reverberation detection algorithms, and sound effect mixing technology based on the cover vocals and accompaniment. This allows users to personalize their song choices and generate high-quality cover songs tailored to their preferences. Therefore, this method only requires users to input their own vocal materials and the song they want to cover to create a high-quality cover song using their timbre, performed according to the original song's professional singing style, and refined through mixing. It solves the problems of damaged audio points and volume imbalances in existing cover song generation methods, detects and repairs damaged audio points (such as vocal segments with reverberation and harmony), and, combined with professional mixing, significantly improves the performance and listening quality of cover songs, meeting users' demands for more personalized cover song production.
[0142] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0143] Based on the same inventive concept, this application also provides a cover song generation device for implementing the cover song generation method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more cover song generation device embodiments provided below can be found in the limitations of the cover song generation method described above, and will not be repeated here.
[0144] In an exemplary embodiment, Figure 11 As shown, a cover song generation device is provided, including: a recognition module 1110, an acquisition module 1120, a replacement module 1130, and a generation model 1140, wherein:
[0145] The identification module 1110 is used to identify the vocal segment to be repaired in the original vocals of the target song selected by the user for a cover song; the vocal segment to be repaired is a vocal segment with reverb and / or harmony.
[0146] The acquisition module 1120 is used to acquire a target vocal segment that matches the vocal segment to be repaired; the target vocal segment is a vocal segment without reverb and harmony, and the target vocal segment and the vocal segment to be repaired correspond to the same singing content in the target song.
[0147] The replacement module 1130 is used to replace the vocal segment to be repaired in the original vocal segment with the target vocal segment to obtain the repaired original vocal segment, and to extract the singing features of the repaired original vocal segment to obtain the guiding vocal features.
[0148] The generation module 1140 is used to extract the timbre features from the sound material provided by the user, and generate a cover song of the target song based on the timbre features and the vocal features; the cover song is a song generated by simulating the user's timbre in singing the target song.
[0149] In an exemplary embodiment, the identification module 1110 is used to detect interference segments in the original vocals of the target song; the interference segments are vocal segments with reverberation and / or harmony; for any interference segment, if the interference segment partially falls within the non-singing time period of the original vocals, or if the interference segment completely falls within the singing time period of the original vocals and the endpoint timestamp of the interference segment is not aligned with the endpoint timestamp of the sound unit, the two endpoint timestamps of the interference segment are adjusted to obtain a corrected interference segment; the two endpoint timestamps of the corrected interference segment are both aligned with the endpoint timestamp of a sound unit and the corrected interference segment does not include the non-singing time period; based on the corrected interference segment, the vocal segment to be repaired in the original vocals of the target song is obtained.
[0150] In an exemplary embodiment, the identification module 1110 is configured to: if the start timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit in the original vocal performance, then correct the start timestamp of the interference segment to the start timestamp of any vocal unit to obtain the corrected interference segment; if the start timestamp of the interference segment falls within the non-singing time period, then correct the start timestamp of the interference segment to the start timestamp of the next vocal unit adjacent to the non-singing time period to obtain the corrected interference segment; if the end timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit, then correct the end timestamp of the interference segment to the end timestamp of any vocal unit to obtain the corrected interference segment; if the end timestamp of the interference segment falls within the non-singing time period, then correct the end timestamp of the interference segment to the end timestamp of the previous vocal unit adjacent to the non-singing time period to obtain the corrected interference segment.
[0151] In an exemplary embodiment, the acquisition module 1120 is configured to, if there is a clean segment in the original vocal recording that matches the vocal segment to be repaired, determine the clean segment as the target vocal segment that matches the vocal segment to be repaired; the clean segment is a vocal segment without reverb and harmony; if there is no clean segment in the original vocal recording that matches the vocal segment to be repaired, determine the target vocal segment that matches the vocal segment to be repaired from the candidate vocal set.
[0152] In an exemplary embodiment, the acquisition module 1120 is used to acquire a candidate voice set; the candidate voice set includes multiple candidate voices; and based on at least one of the singer's gender, pitch rise / fall, and loudness of the candidate voices, a target vocal segment that matches the vocal segment to be repaired is determined from the multiple candidate voices.
[0153] In an exemplary embodiment, the acquisition module 1120 is used to acquire multiple dry voices corresponding to the target song; the dry voices are non-original vocals that do not have reverb or harmony and whose audio length is the same as the audio length of the original vocals; based on the sound quality index information of the dry voices, the multiple candidate voices are determined from the multiple dry voices to obtain the candidate voice set; the sound quality index information includes at least one of loudness, clarity, and pitch.
[0154] In an exemplary embodiment, the device further includes a separation module, which is used to separate the accompaniment and vocals of the target song to obtain the separated original vocals; identify the non-singing time periods between lines in the lyrics based on the position information of each pronunciation unit in the lyrics of the target song; the position information includes the endpoint timestamp of the pronunciation unit and the line identifier information of the pronunciation unit in the lyrics; and replace the audio corresponding to the non-singing time periods in the separated original vocals with blank audio to obtain the original vocals of the target song.
[0155] In an exemplary embodiment, the separation module is configured to determine the beginning and end reserved time periods in the non-singing time periods between any two rows in each row; to use the other time periods in the non-singing time periods between any two rows, excluding the beginning and end reserved time periods, as trimming time periods; and to replace the audio corresponding to the trimming time periods in the separated original vocals with blank audio to obtain the original vocals of the target song.
[0156] Each module in the aforementioned cover song generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0157] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 12As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for generating cover songs. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0158] Those skilled in the art will understand that Figure 12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0159] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the cover song generation method described above. The steps of the cover song generation method described here can be steps from one of the cover song generation methods in the various embodiments described above.
[0160] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the steps of the cover song generation method described above. The steps of the cover song generation method described here may be steps from one of the cover song generation methods in the various embodiments described above.
[0161] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, causes the processor to perform the steps of the cover song generation method described above. The steps of the cover song generation method described here may be steps from one of the cover song generation methods in the various embodiments described above.
[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0163] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0164] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0165] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for generating cover songs, characterized in that, The method includes: For the target song selected by the user for a cover, identify the vocal segment to be repaired in the original vocals of the target song; the vocal segment to be repaired is a vocal segment with reverb and / or harmony; Obtain a target vocal segment that matches the vocal segment to be repaired; the target vocal segment is a vocal segment without reverb and harmony, and the target vocal segment and the vocal segment to be repaired correspond to the same singing content in the target song; The target vocal segment is used to replace the vocal segment to be repaired in the original vocal segment to obtain the repaired original vocal segment. The singing characteristics of the repaired original vocal segment are extracted to obtain the singing characteristics. Extract the timbre features from the user-provided audio material, and generate a cover version of the target song based on the timbre features and the vocal guide features; the cover song is a song generated by simulating the user's timbre when singing the target song.
2. The method according to claim 1, characterized in that, The process of identifying the vocal segment to be repaired in the original vocals of the target song includes: Interference segments in the original vocals of the target song were detected; the interference segments were vocal segments containing reverb and / or harmony. For any of the aforementioned interference segments, if the interference segment partially falls within the non-singing time period of the original vocals, or if the interference segment completely falls within the singing time period of the original vocals and the endpoint timestamp of the interference segment is not aligned with the endpoint timestamp of the sound unit, the two endpoint timestamps of the interference segment are adjusted to obtain a corrected interference segment; the endpoint timestamp of the corrected interference segment is aligned with the endpoint timestamp of the sound unit and the corrected interference segment does not include the non-singing time period; Based on the corrected interference fragment, the vocal fragment to be repaired in the original vocals of the target song is obtained.
3. The method according to claim 2, characterized in that, The adjustment of the two endpoint timestamps of the interference fragment includes: If the start timestamp of the interference segment falls between the two endpoint timestamps of any vocal unit in the original vocal recording, the start timestamp of the interference segment is corrected to the start timestamp of any vocal unit. If the start timestamp of the interference segment falls within the non-singing time period, the start timestamp of the interference segment is corrected to the start timestamp of the next vocal unit adjacent to the non-singing time period. If the end timestamp of the interference segment falls between the two endpoint timestamps of any of the pronunciation units, the end timestamp of the interference segment is corrected to the end timestamp of any of the pronunciation units; If the end timestamp of the interference segment falls within the non-singing time period, the end timestamp of the interference segment is corrected to the end timestamp of the previous vocal unit adjacent to the non-singing time period.
4. The method according to claim 1, characterized in that, Obtaining the target voice segment that matches the voice segment to be repaired includes: If there is a clean segment in the original vocals that matches the vocal segment to be repaired, then the clean segment is determined as the target vocal segment that matches the vocal segment to be repaired; the clean segment is a vocal segment without reverb and harmony. If there is no pure segment in the original vocals that matches the vocal segment to be repaired, then a target vocal segment that matches the vocal segment to be repaired is determined from the candidate vocal set.
5. The method according to claim 4, characterized in that, The step of determining the target human voice segment that matches the human voice segment to be repaired from the candidate voice set includes: Obtain a candidate voice set; the candidate voice set includes multiple candidate voices; Based on at least one of the singer's gender, pitch rise / fall, and loudness of the candidate voices, a target vocal segment matching the vocal segment to be repaired is determined from the plurality of candidate voices.
6. The method according to claim 5, characterized in that, The acquisition of the candidate voice set includes: Obtain multiple dry vocals corresponding to the target song; the dry vocals are non-original vocals that do not have reverb or harmony and whose audio length is the same as the audio length of the original vocals; Based on the sound quality index information of the dry sound, the multiple candidate sounds are determined from the multiple dry sounds to obtain the candidate sound set; the sound quality index information includes at least one of loudness, clarity, and pitch.
7. The method according to claim 1, characterized in that, Before identifying the vocal segment to be repaired in the original vocals of the target song, the method further includes: The accompaniment and vocals of the target song are separated to obtain the original vocals after separation; Based on the positional information of each pronunciation unit in the lyrics of the target song, the non-singing time periods between each line in the lyrics are identified; the positional information includes the endpoint timestamp of the pronunciation unit and the line identifier information of the pronunciation unit in the lyrics; The audio corresponding to the non-singing time period in the separated original vocals is replaced with blank audio to obtain the original vocals of the target song.
8. The method according to claim 7, characterized in that, The step of replacing the audio corresponding to the non-singing time period in the separated original vocals with blank audio to obtain the original vocals of the target song includes: Determine the beginning and end reserved time periods in the non-singing time periods between any two rows in each of the rows; The non-singing time intervals between any two lines, excluding the reserved time intervals at the beginning and the reserved time intervals at the end, are used as the trimming time intervals; The audio corresponding to the trimmed time segment in the separated original vocals is replaced with blank audio to obtain the original vocals of the target song.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Song identification method, computer equipment and storage medium
CN116189707A
Recording fragment identification method, computer equipment and storage medium
CN116386667A