Song screening method, song production method, song screening apparatus, and electronic device
By using machine learning models to filter songs from both the music and lyrics dimensions, the problem of low efficiency in traditional manual screening is solved, enabling fast and accurate song selection and production.
Patent Information
- Application Number
- PCT/CN2024/127094
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2026-04-30
AI Technical Summary
Traditional manual song selection is inefficient and cannot meet the needs of song production.
Machine learning models are used to filter songs based on their musical and lyrical features and their matching degree with the features of the target songs. This includes using large language models or natural language processing models to determine the matching degree between candidate songs and target features.
It enables the rapid and accurate filtering of songs that meet the song selection criteria, improving the efficiency and accuracy of song selection.
Smart Images

Figure CN2024127094_30042026_PF_FP_ABST
Abstract
Description
Song selection methods, song production methods, song selection devices and electronic equipment Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a song selection method, a song production method, a song selection device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] In the traditional song production process, songs need to be screened manually. After listening to each song, people evaluate whether the songs meet the requirements.
[0003] However, as the number of candidate songs that need to be screened continues to increase, the low efficiency of manual song screening leads to longer and longer production times, which cannot meet the production needs.
[0004] Summary of the Invention
[0005] In view of this, embodiments of this disclosure propose a song screening method, a song production method, a song screening device, an electronic device, a computer-readable storage medium, and a computer program product. By utilizing a machine learning model, songs are screened from two dimensions: music and lyrics, based on the degree of matching between the music features and lyric features of the song and the features of the song screening target, thereby improving screening efficiency.
[0006] According to a first aspect of some embodiments of this disclosure, a song filtering method is provided, comprising: acquiring features of a song filtering target, the features of the song filtering target including target music features and target lyrics features; using a first machine learning model to determine a first matching degree between the music features of candidate songs and the target music features; using a second machine learning model to determine a second matching degree between the lyrics features of the candidate songs and the target lyrics features; filtering songs from the candidate songs that match the target features based on the first matching degree and the second matching degree; and determining a target song based on the filtered songs.
[0007] According to a second aspect of some embodiments of this disclosure, a song production method is provided, comprising: a song screening method as described above, screening a target song from candidate songs; obtaining a first human voice audio of the target song; replacing the sound features of the first human voice audio with the target sound features to obtain a second human voice audio; generating an accompaniment matching the second human voice audio based on the second human voice audio; and mixing the second human voice audio and the accompaniment to obtain a mixed song.
[0008] According to a third aspect of some embodiments of this disclosure, a song filtering apparatus is provided, comprising: an acquisition module configured to acquire features of a song filtering target, the features of the song filtering target including target music features and target lyrics features; a first matching module configured to use a first machine learning model to determine a first matching degree between the music features of candidate songs and the target music features; a second matching module configured to use a second machine learning model to determine a second matching degree between the lyrics features of the candidate songs and the target lyrics features; a filtering module configured to filter songs from the candidate songs that match the target features based on the first matching degree and the second matching degree; and a determination module configured to determine a target song based on the filtered songs.
[0009] According to a fourth aspect of some embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform a song selection method as described above or a song production method as described above based on instructions stored in the memory.
[0010] According to a fifth aspect of some embodiments of the present disclosure, a computer-readable storage medium is provided having computer instructions stored thereon that, when executed by a processor, implement the song selection method or the song production method as described above.
[0011] According to a sixth aspect of some embodiments of the present disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement the song selection method or the song production method as described above.
[0012] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0013] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:
[0014] Figure 1 shows a schematic flowchart of a song selection method according to some embodiments of the present disclosure;
[0015] Figure 2 illustrates a flowchart of obtaining features of song selection targets according to some embodiments of the present disclosure;
[0016] Figure 3 illustrates a flowchart of determining a first matching degree according to some embodiments of the present disclosure;
[0017] Figure 4 illustrates a flowchart of determining a second matching degree according to some embodiments of the present disclosure;
[0018] Figure 5 illustrates a flowchart of the song selection process according to some embodiments of the present disclosure;
[0019] Figure 6 illustrates a flowchart of a song selection process according to other embodiments of the present disclosure;
[0020] Figure 7 illustrates a flowchart of determining a target song according to some embodiments of the present disclosure;
[0021] Figure 8 shows a schematic flowchart of a song selection method according to other embodiments of the present disclosure;
[0022] Figure 9 shows a flowchart of a song production method according to some embodiments of the present disclosure;
[0023] Figure 10 shows a block diagram of a song selection apparatus according to some embodiments of the present disclosure;
[0024] Figure 11 shows a block diagram of an electronic device according to some embodiments of the present disclosure;
[0025] Figure 12 shows a block diagram of an electronic device according to other embodiments of the present disclosure. Detailed Implementation
[0026] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0027] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. Unless otherwise specifically stated, the relative arrangement and numerical values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0028] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".
[0029] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0030] The terms “a” and “a plurality” used in this disclosure are illustrative and not restrictive, and those skilled in the art should understand that they should be understood as “one or more” unless otherwise expressly indicated in the context.
[0031] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0032] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0033] In the current song production process, the low efficiency of manual song selection cannot meet the needs of song production. To improve the efficiency of song selection and thus increase production efficiency, this disclosure provides a new song selection method. This method utilizes a machine learning model to select songs based on the matching degree between the song's musical and lyrical features and the features of the target song selection criteria, considering both the music and lyrics.
[0034] Figure 1 shows a schematic flowchart of a song selection method according to some embodiments of the present disclosure.
[0035] As shown in Figure 1, the song selection method includes: Step S11, obtaining the features of the song selection target, wherein the features of the song selection target include target music features and target lyrics features; Step S12, using a first machine learning model, determining a first matching degree between the music features of the candidate songs and the target music features; Step S13, using a second machine learning model, determining a second matching degree between the lyrics features of the candidate songs and the target lyrics features; Step S14, selecting songs from the candidate songs that match the target features based on the first matching degree and the second matching degree; Step S15, determining the target song based on the selected songs.
[0036] The characteristics of the target songs refer to the features of the songs that are to be produced, which are used to filter out target songs. Specifically, they include target music features and target lyric features.
[0037] Target musical characteristics include, for example, sad love songs and pop music. Target lyrical characteristics include, for example, father themes and high readability of the lyrics.
[0038] Machine learning models include, for example, Large Language Models (LLMs) or other natural language processing (NLP) models. For instance, machine learning models such as Variational Autoencoders (VAEs), Convolutional Neural Networks (CNNs), and Transformers can be used to achieve fast and accurate song selection.
[0039] The first and second machine learning models mentioned above can be trained by inputting music and lyrics from the same song as training data, thereby achieving the above functions.
[0040] For example, for the theme of wandering, positive samples in the training corpus could include "The lights are dim, where is home?", while negative samples could include "He has broad shoulders and hardworking hands" (father theme). As another example, for more colloquial song lyrics, positive samples could include "Does he really love you?", while negative samples could include "The days and nights alternate, I hesitate and waver."
[0041] Similar to training lyrics, different positive and negative samples can be selected for training based on different features, enabling the machine learning model to achieve the above functions.
[0042] By using machine learning models, songs can be filtered based on their musical and lyrical features and their matching degree with the features of the target songs. This allows for quick and accurate song filtering.
[0043] The song filtering method according to embodiments of the present disclosure will be further described below with reference to FIG2. FIG2 shows a schematic flowchart of obtaining features of song filtering targets according to some embodiments of the present disclosure.
[0044] In some embodiments, musical features may include descriptive and attribute information of the music. The descriptive information of the music includes at least one of theme, imagery, emotional type, and style type. The attribute information of the music includes at least one of rhythm, melody, timbre, and harmony. Lyric features include descriptive and attribute information of the lyrics. The descriptive information of the lyrics includes at least one of theme, imagery, emotional type, and style type. The attribute information of the lyrics includes at least one of clarity, consistency, and readability.
[0045] In other words, descriptive information refers to the subjective characteristics of music or lyrics, including subjective evaluations of the music or lyrics. Attribute information, on the other hand, refers to the basic attributes of music or lyrics, including the inherent properties of the music or lyrics themselves.
[0046] The music can be described in various ways, including themes such as fatherhood, wandering, homesickness, remembrance of friends, and celebration of success. Musical imagery can include sunsets, waves, rocks, streets, and architecture. Emotional types can range from sadness and joy to pain and emotion. Genres can include rock, jazz, pop, folk, and rap.
[0047] Regarding the lyrics, the themes include fatherhood, wandering, longing for home, remembrance of friends, and celebration of success. The imagery includes sunsets, waves, rocks, streets, and buildings. The emotional types include sadness, joy, pain, and emotion. The stylistic types include narrative, lyrical, inspirational, satirical, and colloquial.
[0048] Regarding the attributes of music, rhythm includes, for example, the percentage of eighth notes, sixteenth notes, etc. Melody includes, for example, pitch, tonic, scale, etc. Timbre includes, for example, the percentage of high-frequency components, low-frequency components, noise components, etc. Harmony includes, for example, the number, density, and frequency of chords, etc.
[0049] Regarding the attributes of lyrics, clarity includes aspects such as word overlap and grammatical accuracy. Consistency includes aspects such as contextual relevance and coherence. Readability includes aspects such as average sentence length and the number of complex words.
[0050] Specifically, target music features can include descriptive and attribute information of the target music, and target lyrics features can include descriptive and attribute information of the target lyrics. The following section focuses on describing the process of obtaining the features of the song selection target, including the aforementioned target music features and target lyrics features.
[0051] As shown in Figure 2, obtaining the features of the song selection target in step S11 includes: step S111, obtaining the description information of the target music and the attribute information matching the description information of the target music; step S112, obtaining the description information of the target lyrics and the attribute information matching the description information of the target lyrics.
[0052] In other words, in the above embodiments, the attribute information of the target music and target lyrics matches the descriptive information. For example, if the emotional type of the obtained target music is sadness, the rhythm of the target music is a slow rhythm that matches sadness. If the theme of the obtained target lyrics is daily life, the readability of the target lyrics is high readability that matches daily life.
[0053] Through the above processing, the descriptive and attribute information of the target music features and target lyrics features used in the song selection process can be matched with each other, avoiding the selection of songs that do not conform to basic logic, thereby improving the accuracy of song selection.
[0054] The above description, with reference to Figure 2, illustrates how the features of the song selection target are obtained in step S11. The following description, with reference to Figure 3, describes how the first matching degree between the musical features of the candidate songs and the target musical features is determined in step S12 of the above embodiment. Figure 3 shows a schematic flowchart illustrating the determination of the first matching degree according to some embodiments of this disclosure.
[0055] Similar to the target music features, the music features of candidate songs can also include descriptive and attribute information. The following section focuses on the process of determining the first degree of match between music features including descriptive and attribute information and the target music features.
[0056] Step S12, using a first machine learning model, determines the first matching degree between the musical features of the candidate song and the target musical features, including: step S121, determining the matching degree of the candidate music with the target music's description information and the matching degree of its attribute information; step S122, determining the weights of the matching degree of the description information and the matching degree of the attribute information based on the description information of the target music; and step S123, determining the first matching degree based on the matching degree of the description information, the matching degree of the attribute information, and the weights of the matching degree of the description information and the matching degree of the attribute information.
[0057] In step S121, the matching degree of description information and the matching degree of attribute information between the candidate music and the target music are determined respectively, including the matching degree of theme and rhythm between the candidate music and the target music.
[0058] In some embodiments, determining the matching degree of the description information between the candidate music and the target music includes: determining the matching degree of the description information between the candidate music and the target music based on the similarity of the description information between the candidate music and the target music.
[0059] In contrast, determining the matching degree of attribute information between the candidate music and the target music may include: determining the matching degree of attribute information between the candidate music and the target music based on the difference in attribute information between the candidate music and the target music.
[0060] In the above embodiments, the attribute information of the music can be quantified using numerical values, such as the number of harmonics. The matching degree between the attribute information of the candidate music and the target music can be accurately determined based on the difference between their numerical values. In this process, the attribute information of the target music can be understood as the target threshold for filtering.
[0061] Conversely, descriptive information about music can be difficult to quantify, such as emotional type. Descriptive vectors can be generated separately based on the descriptive information of the candidate music and the target music. For example, the emotional type labels of the music can be converted into an encoding that can be processed by a machine learning model, and a descriptive vector can be generated based on this encoding.
[0062] Based on the similarity between the description vectors of candidate music and target music, the matching degree of description information between the candidate and target music can be accurately determined. In the above process, the description information of the target music can be understood as the target type to be filtered.
[0063] In step S122, during the process of determining the first matching degree, different weights can be assigned to the matching degree of description information and the matching degree of attribute information based on the description information of the target music.
[0064] By assigning different weights, important items in the music features can be given higher priority during the screening process, thereby improving the matching degree of important features of the screened music and further improving the accuracy of the screening.
[0065] For example, the descriptive information of target music could include the genre being pop. Since pop songs typically have simple, uncomplicated emotions, fast tempos, and simple, repetitive melodies, they easily attract listeners' attention. Therefore, for pop songs, attribute information such as rhythm is more important than descriptive information such as emotion; thus, the weight of musical attribute information is greater than the weight of musical descriptive information.
[0066] In some embodiments, when the music's attribute information or description information includes more than one item, different weights can be assigned to different items of the attribute information or description information.
[0067] For example, if the description information of the target music includes theme, emotion type, and style type, assuming that the theme of the target music is daily life, the emotion type is happiness, and the style type is rap, since the rap type is significantly different from other types, and the daily life theme is less different from other themes, the weight of each item in the description information of the target music can be style type weight greater than emotion type weight greater than theme weight.
[0068] In step S123, the matching degree between the candidate music and the target music, i.e. the first matching degree, is determined by the matching degree of the description information, the matching degree of the attribute information, and their respective weights.
[0069] Through the above processing steps, the matching degree between candidate music and target music can be accurately determined.
[0070] The first matching degree between the musical features of the candidate song and the target musical features was determined in step S12, as described above with reference to Figure 3. The second matching degree between the lyric features of the candidate song and the target lyric features is then determined in step S13, with reference to Figure 4. Figure 4 shows a schematic flowchart illustrating the determination of the second matching degree according to some embodiments of this disclosure.
[0071] Similar to the target lyric features, the lyric features of candidate songs can also include descriptive and attribute information about the candidate music. The following section focuses on the process for determining the second degree of matching between the lyric features, which include descriptive and attribute information, and the target lyric features.
[0072] As shown in Figure 4, in step S13, the second matching degree between the lyrics features of the candidate song and the target lyrics features is determined using the second machine learning model, including: step S131, determining the matching degree of the description information and the matching degree of the attribute information between the candidate lyrics and the target lyrics; step S132, determining the weights of the matching degree of the description information and the matching degree of the attribute information based on the description information of the target lyrics; and step S133, determining the second matching degree based on the matching degree of the description information, the matching degree of the attribute information, and the weights of the matching degree of the description information and the matching degree of the attribute information.
[0073] In step S131, the matching degree of descriptive information and attribute information between the candidate lyrics and the target lyrics are determined respectively, including the emotional matching degree and readability matching degree between the candidate lyrics and the target lyrics.
[0074] Similar to music, in some embodiments, determining the matching degree of the descriptive information between the candidate lyrics and the target lyrics includes: determining the matching degree of the descriptive information between the candidate lyrics and the target lyrics based on the similarity of the descriptive information between the candidate lyrics and the target lyrics.
[0075] Accordingly, determining the matching degree of attribute information between the candidate lyrics and the target lyrics may include: determining the matching degree of attribute information between the candidate lyrics and the target lyrics based on the difference in attribute information between the candidate lyrics and the target lyrics.
[0076] In the above embodiments, the attribute information of the lyrics can be quantified using numerical values, such as word overlap rate. The matching degree between the attribute information of the candidate lyrics and the target lyrics can be accurately determined based on the difference between their numerical values. In this process, the attribute information of the target lyrics can be understood as the target threshold for filtering.
[0077] Conversely, descriptive information in lyrics can be difficult to quantify, such as emotion type. Descriptive vectors can be generated separately based on the descriptive information of candidate lyrics and target lyrics. For example, the emotion type labels of the lyrics can be converted into an encoding that can be processed by a machine learning model, and a descriptive vector can be generated based on this encoding.
[0078] Based on the similarity between the description vectors of candidate lyrics and target lyrics, the matching degree of descriptive information between the candidate lyrics and target lyrics can be accurately determined. In the above process, the descriptive information of the target lyrics can be understood as the target type to be filtered.
[0079] In step S132, during the process of determining the second matching degree, different weights can be assigned to the matching degree of description information and the matching degree of attribute information based on the description information of the target lyrics.
[0080] By assigning different weights, important items in the lyrics features can be given higher priority during the screening process, thereby improving the matching degree of important features of the selected lyrics and further improving the accuracy of the screening.
[0081] For example, the descriptive information in the target lyrics could include the theme of fatherhood and the emotion of bitterness. Because songs about fathers generally have more relaxed linguistic requirements, while expressing more complex emotions, they are more likely to resonate with listeners. Therefore, for songs about fathers, the importance of descriptive information such as emotion outweighs the importance of attribute information such as consistency. Thus, the weight of descriptive information in the lyrics is greater than the weight of attribute information in the lyrics.
[0082] In some embodiments, when the attribute information or description information of the lyrics includes more than one item, different weights can be assigned to different items of the attribute information or description information.
[0083] For example, if the descriptive information of the target lyrics includes theme, emotion type, and style type, assuming the theme of the target lyrics is father, the emotion type is bitterness, and the style type is colloquial, since the father theme is relatively specific and differs greatly from other themes, and the colloquial style differs less from other styles, the weight of each item in the descriptive information of the target lyrics can be that the weight of theme is greater than the weight of emotion type, which is greater than the weight of style type.
[0084] In step S133, the matching degree between the candidate lyrics and the target lyrics, i.e., the second matching degree, is determined by the matching degree of the description information, the matching degree of the attribute information, and their respective weights.
[0085] Similarly, when matching music, the above processing steps can accurately determine the degree of matching between candidate lyrics and target lyrics.
[0086] The above description, with reference to Figures 3 and 4, outlines some specific processes for determining the matching degree in this disclosure. The following description, with reference to Figure 5, describes, according to some embodiments of this disclosure, how, in step S14, songs matching the target features are selected from candidate songs based on a first matching degree and a second matching degree. Figure 5 shows a schematic flowchart of the song selection process according to some embodiments of this disclosure.
[0087] As mentioned earlier, the target music features include descriptive and attribute information of the target music, and the target lyrics features include descriptive and attribute information of the target lyrics.
[0088] As shown in Figure 5, in step S14, filtering songs that match the target feature from the candidate songs based on the first matching degree and the second matching degree includes: step S141, determining the weights of the first matching degree and the second matching degree based on at least one of the description information of the target music and the language of the song selection target; step S142, filtering songs that match the target feature from the candidate songs based on the first matching degree, the second matching degree, and the weights of the first matching degree and the second matching degree.
[0089] In step S141, the target language can be filtered by the description information of the target music and / or the song, and different weights can be assigned to the first matching degree and the second matching degree respectively.
[0090] For example, the description of the target music might include the genre being rap. For rap songs, the core is usually the lyrics, which convey the song's theme and message. In this case, the weight of the second match can be greater than the weight of the first match, thus focusing the song selection on the lyrics.
[0091] For example, when the target language for song selection is Chinese, due to the vastness of Chinese culture, Chinese songs typically place greater emphasis on the beauty of the lyrics and the vivid imagery they create. Therefore, for Chinese songs, the weight of the second match degree can be greater than that of the first match degree. Conversely, when the target language for song selection is English, English songs typically place greater emphasis on the beauty of the melody and the rhythm of the music. Therefore, for English songs, the weight of the first match degree can be greater than that of the second match degree.
[0092] In step S142, the matching degrees are weighted and summed according to the weights of the first matching degree and the second matching degree, and the songs are filtered according to the overall matching degree between the songs and the target obtained by the weighted summation.
[0093] By using the above process, different focuses can be assigned during the screening process, thereby selecting higher-quality songs that are more closely matched to the target and improving the accuracy of song selection.
[0094] The above description, with reference to Figure 5, illustrates how, in step S14, songs matching the target features are selected from candidate songs based on a first matching degree and a second matching degree, according to some embodiments of the present disclosure. The following description, with reference to Figures 6 to 8, further describes how, according to some embodiments of the present disclosure, a matching degree between music and lyrics is additionally introduced during the song selection process.
[0095] Figure 6 illustrates a flowchart of song selection according to other embodiments of the present disclosure.
[0096] As shown in Figure 6, in step S14, filtering songs that match the target feature from the candidate songs based on the first matching degree and the second matching degree includes: step S141', using a third machine learning model to determine the third matching degree between the music and lyrics of the candidate songs based on the music features and lyrics features of the candidate songs; step S142', filtering songs that match the target feature from the candidate songs based on the first matching degree, the second matching degree, and the third matching degree.
[0097] The third machine learning model is similar to the first and second machine learning models, and may include large language models or other natural language processing models. For example, machine learning models such as variational autoencoders, convolutional neural networks, and transformers can be used to achieve fast and accurate song selection.
[0098] In step S141', a third machine learning model is used to determine the third degree of matching between the music and lyrics of the candidate song. This third degree of matching refers to the degree of matching between the music and lyrics of the candidate song.
[0099] The matching degree of lyrics and music mainly reflects the coordination between the lyrics and the music, such as emotional consistency, thematic consistency, rhythmic synchronization, harmonic support, and linguistic rhythm.
[0100] As an example, thematic consistency is determined by the similarity between the theme of the music and the theme of the lyrics. When both the music and the lyrics describe the theme of missing family, the song has high thematic consistency. Conversely, when the music describes the theme of missing family while the lyrics describe the theme of a happy life, the lyrics have low thematic consistency.
[0101] In some embodiments, the music can be divided into different segments, and the lyric-music matching degree between each music segment and its corresponding lyric segment can be calculated separately, ultimately integrating them into a third matching degree. For example, music segments can be divided based on at least one of the following: the rhythm of the music and repetitive parts in the music.
[0102] Dividing music according to its rhythm refers to, for example, dividing the music into different segments based on rhythmic changes, so that the rhythm is relatively consistent within each segment.
[0103] Dividing music based on repetition refers to identifying musical phrases or sections that appear repeatedly. Musical segments created in this way follow the same phrase or section, resulting in multiple segments arranged in a parallel structure. Typically, the lyrics corresponding to the same phrase or section should be identical or similar.
[0104] Similar to music, in other embodiments, lyrics can be divided into different segments, and the word-music matching degree between each lyric segment and its corresponding musical segment can be calculated separately, ultimately integrating them into a third matching degree. For example, lyric segments can be divided based on at least one of the following: sentence segments or repeated parts of the lyrics.
[0105] Dividing lyrics into phrases and segments means, for example, treating each complete line of lyrics as a lyric segment, ensuring that the meaning of each lyric segment is coherent.
[0106] Dividing lyrics based on repetition means, for example, dividing the lyrics into different segments at the same words or phrases that appear repeatedly. For instance, lyrics describing a father may contain the word "father" in multiple places. By dividing the lyrics into different segments at these words, multiple lyric segments with a similar parallel structure can be obtained.
[0107] Using the methods described above, the third degree of match between the music and lyrics of candidate songs can be accurately determined.
[0108] Step S142': Based on the first matching degree, the second matching degree, and the third matching degree, select matching songs from the candidate songs by combining the three dimensions.
[0109] In the embodiment shown in Figure 6, the third matching degree, namely the lyric-music matching degree of the candidate song, is used together with the first matching degree and the second matching degree for song selection. For example, the overall matching degree between the candidate song and the target can be obtained by assigning weights to the first matching degree, the second matching degree, and the third matching degree respectively, and then selecting songs that match the target features based on the overall matching degree.
[0110] By following the steps above, an additional dimension is introduced when filtering songs, which avoids mismatches between the music and lyrics of the filtered songs and further improves the accuracy of the matching.
[0111] The above, based on Figure 6, introduces one application method of lyric-music matching, which involves using the lyric-music matching score along with the first and second matching scores to filter candidate songs. The following, in conjunction with Figures 7 and 8, introduces other applications of lyric-music matching, using multiple rounds of filtering with the lyric-music matching score, the first matching score, and the second matching score respectively to determine the target song.
[0112] Figure 7 illustrates a flowchart of determining a target song according to some embodiments of the present disclosure. In the embodiment shown in Figure 7, candidate songs are screened in a first round using a first matching degree and a second matching degree, and the songs selected in the first round are screened in a second round using a lyric-music matching degree.
[0113] As shown in Figure 7, in step S15, determining the target song based on the selected songs includes: step S151, using a third machine learning model to determine the fourth matching degree between the music and lyrics of the selected songs based on the music features and lyrics features of the selected songs; step S152, determining the target song based on the fourth matching degree.
[0114] Both steps S151 and S141' determine the matching degree between music and lyrics. The difference is that step S141' determines the matching degree between the music and lyrics of all candidate songs, while step S151 determines the matching degree between the music and lyrics of the songs selected based on the first matching degree and the second matching degree.
[0115] By only calculating the lyric-music match rate of songs filtered by the first and second matching rates, the amount of calculation required for lyric-music match rate can be reduced, thus improving the efficiency of song filtering.
[0116] Furthermore, similar to the embodiment shown in Figure 6, the embodiment shown in Figure 7 introduces an additional dimension when filtering songs, avoiding mismatches between the music and lyrics of the filtered songs, and further improving the accuracy of the matching.
[0117] The above description, with reference to Figure 7, illustrates one application of the lyric-music matching degree. The following description, with reference to Figure 8, illustrates another application of the lyric-music matching degree. Figure 8 shows a flowchart illustrating a song selection method according to other embodiments of this disclosure. In the embodiment shown in Figure 8, a first round of selection is performed using a fifth matching degree to determine candidate songs, and a second round of selection is performed using a first matching degree and a second matching degree.
[0118] As shown in Figure 8, based on the embodiment shown in Figure 1, the song selection method may further include the following steps before step S11: Step S01, using a third machine learning model to determine the fifth matching degree between the music and lyrics of the songs in the song library based on the music features and lyrics features of the songs in the song library; Step S02, selecting candidate songs from the song library based on the fifth matching degree.
[0119] In steps S01 to S02, candidate songs are determined by the matching degree of lyrics and music in the song library. This song library includes, for example, all songs provided by suppliers. Step S01 is similar to step S141' in Figure 6, which also determines the matching degree between music and lyrics, and determines the matching degree between music and lyrics of all songs in the song library.
[0120] In the embodiment shown in Figure 8, by using the word-song matching degree for initial screening, the number of candidate songs identified can be reduced, thereby reducing the computational load when calculating the first matching degree and the second matching degree, and improving the efficiency of song screening.
[0121] In addition, similar to the embodiments shown in Figures 6 and 7, the embodiment shown in Figure 8 introduces an additional dimension when filtering songs, which avoids mismatches between the music and lyrics of the filtered songs and further improves the accuracy of the matching.
[0122] The song selection method proposed in this disclosure has been described above with reference to Figures 1 to 8, which can quickly and accurately complete song selection. The song production method based on the above selection method, proposed in this disclosure, is described below with reference to Figure 9. Figure 9 shows a schematic flowchart of a song production method according to some embodiments of this disclosure.
[0123] As shown in Figure 9, the song production method includes: step S91, as described in any of the above embodiments, selecting a target song from candidate songs; step S92, obtaining a first human voice audio of the target song; step S93, replacing the sound features of the first human voice audio with the target sound features to obtain a second human voice audio; step S94, generating an accompaniment that matches the second human voice audio based on the second human voice audio; and step S95, mixing the second human voice audio and the accompaniment to obtain a mixed song.
[0124] In step S91, the target songs that match the filtering target are filtered out using the song filtering method described above.
[0125] In some embodiments, the song selection targets in the song selection method can be determined based on the target audience of the song production, and then the selection can be carried out. For example, since middle-aged and elderly people usually like to listen to melancholic love songs, when producing songs for middle-aged and elderly people, it is necessary to select corresponding songs for production, that is, the characteristics of the song selection targets include melancholic love songs.
[0126] In step S92, the first vocal audio of the target song is obtained. The vocal audio refers to the audio of a human singing the target song. For example, the first vocal audio may include a dry audio file of a singer performing the target song.
[0127] By having a singer perform the target song first and then replacing the target song's vocal characteristics, compared to directly generating a target song with the target vocal characteristics, the song can be made more natural and the quality of the song production can be improved.
[0128] In step S93, the target sound feature can be a sound feature that matches the feature of the song selection target. For example, when creating a song for middle-aged and elderly people, if the feature of the song selection target is a somber and melancholic love song, the sound feature of the first vocal audio can be replaced with a weathered timbre that matches a somber and melancholic love song.
[0129] In step S94, a machine learning model can be used to automatically generate an accompaniment that matches the second human voice audio, such as generating an accompaniment whose rhythm and emotion match the second human voice audio, thereby improving the efficiency of song production.
[0130] In step S95, a machine learning model can be used to merge different sounds together for mixing. Using a machine learning model for mixing can effectively improve the efficiency of song production.
[0131] For example, to prevent accompaniment from interfering with vocals, machine learning models can be used to automatically adjust the weights of various sounds at different points in a song, increasing the weight of the accompaniment when there are no vocals and decreasing its weight during key vocal sections. Similarly, since drum and piano sounds conflict, one can be given a higher weight and the other a lower weight, thus avoiding mutual interference between them. This process eliminates the need to differentiate sound features during the initial vocal audio production and allows for subsequent replacement with target sound features that match the vocals, improving song production efficiency. For instance, the same singer can perform multiple songs, and the sound features of each song can be replaced with different sound features later.
[0132] Furthermore, by replacing only the vocal features, other features in the first vocal audio, such as emotional features, can be preserved, thereby improving the quality of song production.
[0133] In summary, the song production method proposed in this disclosure can improve the efficiency of song selection and production, thereby improving the overall efficiency of the song production process.
[0134] The above describes the song selection method and song production method provided by embodiments of this disclosure. The following description, with reference to FIG10, describes a song selection apparatus according to embodiments of this disclosure, used to perform any of the above-described song selection methods.
[0135] Figure 10 shows a block diagram of a song selection apparatus according to some embodiments of the present disclosure.
[0136] As shown in Figure 10, the song filtering device 10 includes: an acquisition module 101 configured to acquire features of a song filtering target, the features of which include target music features and target lyrics features; a first matching module 102 configured to use a first machine learning model to determine a first matching degree between the music features of a candidate song and the target music features; a second matching module 103 configured to use a second machine learning model to determine a second matching degree between the lyrics features of the candidate song and the target lyrics features; a filtering module 104 configured to filter songs from the candidate songs that match the target features based on the first matching degree and the second matching degree; and a determination module 105 configured to determine a target song based on the filtered songs.
[0137] The acquisition module 101 of the song filtering device 10 can be used to execute step S11 of FIG1. The first matching module 102 of the song filtering device 10 can be used to execute step S12 of FIG1. The second matching module 103 of the song filtering device 10 can be used to execute step S13 of FIG1. The filtering module 104 of the song filtering device 10 can be used to execute step S14 of FIG1. The determining module 105 of the song filtering device 10 can be used to execute step S15 of FIG1.
[0138] The song filtering device disclosed herein can quickly and accurately filter songs, thereby improving the efficiency of song filtering.
[0139] The song selection apparatus according to an embodiment of the present disclosure has been described above with reference to FIG. 10. Next, with reference to FIGS. 11 and 12, an electronic device according to an embodiment of the present disclosure will be described for performing any embodiment of the aforementioned song selection method or any embodiment of the aforementioned song production method.
[0140] Figure 11 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0141] As shown in FIG11, the electronic device 11 includes: a memory 111; and a processor 112 coupled to the memory 111, the processor 112 being configured to execute the song selection method or the song production method described in any of the foregoing embodiments based on instructions stored in the memory 111.
[0142] Memory 111 is used to store one or more computer-readable instructions. Memory 111 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 111 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.
[0143] The processor 112 is configured to execute computer-readable instructions to implement the song selection method or the song production method described in any of the foregoing embodiments. For specific implementation details of each step of the song selection method or song production method, please refer to the above embodiments; repeated details will not be elaborated here.
[0144] Processor 112 can be configured to execute the steps shown in Figures 1 through 9. Processor 112 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.
[0145] The processor 112 and the memory 111 can communicate with each other directly or indirectly. For example, the processor 112 and the memory 111 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 112 and the memory 111 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0146] It should be noted that the components of the electronic device 11 shown in Figure 11 are exemplary and not limiting. The electronic device 11 may have other components depending on the actual application requirements. The processor 112 can control other components in the electronic device 11 to perform the desired functions.
[0147] Electronic device 11 can be implemented by software, firmware and / or hardware, and can be integrated into a device with relevant applications installed.
[0148] Figure 12 shows a block diagram of an electronic device according to other embodiments of the present disclosure.
[0149] The electronic device 12 shown in Figure 12 can be a computer system with a dedicated hardware structure, which can perform corresponding functions when relevant applications are installed.
[0150] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0151] As shown in Figure 12, the Central Processing Unit (CPU) 121 performs various processes based on programs stored in the Read-Only Memory (ROM) 122 or programs loaded from the Storage Section 128 into the Random Access Memory (RAM) 123. The RAM 123 stores data required as needed when the CPU 121 performs various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 122, RAM 123, and Storage Section 128 can be various forms of computer-readable storage media. It should be noted that although the ROM 122, RAM 123, and Storage Section 128 are shown separately in Figure 12, one or more of them can be combined or located in the same or different memories or storage modules.
[0152] CPU 121, ROM 122 and RAM 123 are interconnected via bus 124. Input / output interface 125 is also connected to bus 124.
[0153] The following components are connected to the input / output interface 125: input section 126, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 127, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 128, including hard disks, magnetic tapes, etc.; and communication section 129, including network interface cards such as LAN cards, modems, etc. The communication section 129 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various parts of the electronic device 12 shown in Figure 12 communicate via bus 124, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0154] As needed, drive 1210 is also connected to input / output interface 125. Removable media 1211, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1210 as needed, so that computer programs read from them can be installed into storage section 128 as needed.
[0155] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 1211.
[0156] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the song selection method or the song production method described in any of the foregoing embodiments. The computer program product includes a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 129, or installed from storage section 128, or installed from ROM 122. When the computer program is executed by CPU 121, the song selection method or song production method of the embodiments of this disclosure is performed.
[0157] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0158] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0159] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the song selection method or the song production method described in any of the foregoing embodiments.
[0160] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0161] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0162] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the song selection method or the song production method described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.
[0163] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0165] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0166] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A song selection method, comprising: The features of the target songs to be filtered are obtained, including target music features and target lyrics features. Using a first machine learning model, a first degree of matching is determined between the musical features of the candidate song and the target musical features; Using a second machine learning model, a second matching degree is determined between the lyric features of the candidate song and the target lyric features; Based on the first matching degree and the second matching degree, select songs from the candidate songs that match the target features; Based on the selected songs, determine the target song.
2. The song selection method according to claim 1, wherein, The target music features include descriptive and attribute information of the target music, and the target lyrics features include descriptive and attribute information of the target lyrics. The features for obtaining the song selection target include: Obtain the description information of the target music and the attribute information that matches the description information of the target music; Obtain the description information of the target lyrics and the attribute information that matches the description information of the target lyrics.
3. The song selection method according to claim 1, wherein, The target music features include descriptive and attribute information of the target music, and the target lyrics features include descriptive and attribute information of the target lyrics. Based on the first and second matching degrees of the candidate songs, filtering songs from the candidate songs that match the target features includes: The weights of the first matching degree and the second matching degree are determined based on at least one of the description information of the target music and the language of the song selection target; Based on the first matching degree, the second matching degree, and the weights of the first matching degree and the second matching degree, songs that match the target features are selected from the candidate songs.
4. The song selection method according to claim 2, wherein, The musical features of the candidate songs include descriptive information and attribute information of the candidate songs. Determining the first degree of matching between the musical features of the candidate songs and the target musical features includes: Determine the matching degree of the candidate music with the target music in terms of description information and attribute information; Based on the description information of the target music, determine the weights of the matching degree of the description information and the matching degree of the attribute information; The first matching degree is determined based on the matching degree of the description information, the matching degree of the attribute information, and the weights of the matching degree of the description information and the matching degree of the attribute information.
5. The song selection method according to claim 4, wherein, Determining the matching degree of the candidate music with the target music in terms of description information and attribute information includes: The matching degree of the description information of the candidate music and the target music is determined based on the similarity between their description information. The matching degree of attribute information between the candidate music and the target music is determined based on the difference in attribute information between the candidate music and the target music.
6. The song selection method according to claim 2, wherein, The lyric features of the candidate songs include descriptive information and attribute information of the candidate lyrics. Determining the second matching degree between the lyric features of the candidate songs and the target lyric features includes: Determine the matching degree of the candidate lyrics with the target lyrics in terms of description information and attribute information; Based on the description information of the target lyrics, determine the weights of the matching degree of the description information and the matching degree of the attribute information; The second matching degree is determined based on the matching degree of the description information, the matching degree of the attribute information, and the weights of the matching degree of the description information and the matching degree of the attribute information.
7. The song description method according to claim 6, wherein, Determining the matching degree of the candidate lyrics with the target lyrics in terms of description information and attribute information includes: The matching degree between the candidate lyrics and the target lyrics is determined based on the similarity between their descriptive information. The matching degree of attribute information between the candidate lyrics and the target lyrics is determined based on the difference in attribute information between the candidate lyrics and the target lyrics.
8. The song selection method according to claim 1, wherein, Based on the first matching degree and the second matching degree, filtering songs from the candidate songs that match the target features includes: Using a third machine learning model, the third degree of matching between the music and lyrics of the candidate songs is determined based on the music features and lyrics features of the candidate songs; Based on the first matching degree, the second matching degree, and the third matching degree, songs that match the target features are selected from the candidate songs.
9. The song selection method according to claim 1, wherein, Based on the selected songs, the target songs include: Using a third machine learning model, a fourth degree of matching between the music and lyrics of the selected songs is determined based on the music features and lyrics features of the selected songs; The target song is determined based on the fourth matching degree.
10. The song selection method according to claim 1, further comprising: Using a third machine learning model, the fifth degree of matching between the music and lyrics of the songs in the song library is determined based on the music features and lyrics features of the songs in the song library; Based on the fifth matching degree, candidate songs are selected from the song library.
11. The song selection method according to claim 1, wherein: The musical features include descriptive information and attribute information of the music. The descriptive information of the music includes at least one of theme, imagery, emotional type, and style type. The attribute information of the music includes at least one of rhythm, melody, timbre, and harmony. The lyric features include descriptive information and attribute information of the lyrics. The descriptive information of the lyrics includes at least one of theme, imagery, emotional type, and style type. The attribute information of the lyrics includes at least one of clarity, consistency, and readability of the lyrics.
12. A method for producing a song, comprising: The song filtering method as described in any one of claims 1 to 11 filters out target songs from candidate songs; Obtain the first vocal audio of the target song; The sound features of the first human voice audio are replaced with the target sound features to obtain the second human voice audio; Generate an accompaniment that matches the second human voice audio; The second vocal audio and the accompaniment are mixed to obtain the mixed song.
13. The song production method according to claim 11, wherein, The target sound features are sound features that match the features of the song selection target.
14. A song selection device, comprising: The acquisition module is configured to acquire features of the song selection target, the features of the song selection target including target music features and target lyrics features; The first matching module is configured to use a first machine learning model to determine a first degree of matching between the musical features of the candidate song and the target musical features. The second matching module is configured to use a second machine learning model to determine a second degree of matching between the lyric features of the candidate song and the target lyric features; The filtering module is configured to filter songs that match the target feature from the candidate songs based on the first matching degree and the second matching degree; The "determine" module is configured to determine the target song based on the filtered songs.
15. An electronic device comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the song selection method as claimed in any one of claims 1 to 11 or the song production method as claimed in any one of claims 12 to 13, based on instructions stored in the memory.
16. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the song selection method of any one of claims 1 to 11 or the song production method of any one of claims 12 to 13.
17. A computer program product, when run on a computer, causes the computer to implement the song selection method of any one of claims 1 to 11 or the song production method of any one of claims 12 to 13.
Citation Information
Patent Citations
Music searching and vehicle-mounted music playing method and device, equipment and storage medium
CN110717062A
Song tone conversion method, system and device and storage medium
CN112331222A
Song matching method and device, equipment, medium and product
CN114840707A
Composition using correlation between melody and lyrics
US20140174279A1