A music generation method, system, device, medium, and computer program product

By displaying the matching relationship of music feature sets on the terminal device, the server generates music after the user selects the target feature set, which solves the problem that users have difficulty accurately describing the music they want, and improves the accuracy of generated music and user experience.

CN122135674APending Publication Date: 2026-06-02PETAL CLOUD TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PETAL CLOUD TECH CO LTD
Filing Date
2024-12-02
Publication Date
2026-06-02

Smart Images

  • Figure CN122135674A_ABST
    Figure CN122135674A_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing, and more particularly to a music generation method, system, device, medium, and computer program product. In this method, when a user uses a music generation application to generate target music, the terminal device acquires multiple sets of music features selected by the user, determines the matching degree between each pair of music feature sets, and displays it to the user, allowing the user to determine the music features to select based on the matching degree. The user-selected music feature sets are then sent to the corresponding server, which generates the music. In this way, the server can generate music based on multiple music feature sets with high matching degrees selected by the user, making the music generated by the server more closely match the user's desired target music, effectively improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more particularly to a music generation method, system, device, medium, and computer program product. Background Technology

[0002] Currently, users are no longer limited to listening to music on terminal devices (such as mobile phones), but are creating their own music based on usage scenarios. Music generation applications installed on terminal devices can automatically generate music according to the user's settings. For example, if a user needs opening music for a wedding, they can set corresponding tags in the music generation application installed on their terminal device, such as wedding, soothing, violin, etc. The server corresponding to the music generation application will then generate the opening music according to the user's set tags using algorithms such as deep learning.

[0003] In some cases, the music generated by the server of a music generation application based on user-defined tags may not meet the user's needs. In such situations, the user needs to reset the tags, and the server needs to regenerate the music until it satisfies their requirements, thus impacting the user experience. Summary of the Invention

[0004] To address the problem that users cannot accurately describe the music they want, and have to repeatedly change the selected tags and perform the music generation operation until their needs are met, resulting in a poor user experience, this application provides a music generation method, system, device, medium, and computer program product.

[0005] In a first aspect, embodiments of this application provide a music generation method, the method comprising: a terminal device acquiring multiple music feature sets selected by a user and sending the multiple music feature sets to a server, the music feature sets including one or more music features; the server determining a first matching relationship of the multiple music feature sets based on the multiple music feature sets, and sending the first matching relationship to the terminal device, the first matching relationship including the matching degree between every two music feature sets in the multiple music feature sets; the terminal device displaying the first matching relationship; the terminal device acquiring a target music feature set selected by the user from the multiple music feature sets based on the first matching relationship, and sending the target music feature set to the server; and the server generating target music based on the target music feature set.

[0006] In some alternative embodiments, the terminal device can display multiple music features to the user by showing a selection interface, such as the first selection interface and the second selection interface mentioned below, so that the user can select multiple sets of music features from the multiple music features. Then, the terminal device can obtain the multiple sets of music features selected by the user.

[0007] In some optional embodiments, corresponding to the terminal device acquiring a first music feature set selected by the user from a first selection interface and a second music feature set selected from a second selection interface, the terminal device can send the acquired first and second music feature sets to the server. The server then determines the matching degree between the first and second music feature sets based on the first and second music feature sets and sends the matching degree to the terminal device. Thus, the terminal device can display a music feature graph to the user, including the first music feature set, the second music feature set, and a first matching relationship. The first matching relationship includes the matching degree between the first and second music feature sets. Then, the terminal device can acquire a target music feature set selected by the user based on the music feature graph and send the target music feature set to the server. The server can generate target music based on the target music feature set.

[0008] Therefore, by displaying multiple music feature sets, the terminal device can clearly describe the various parts related to the target music and their matching relationships. Furthermore, by displaying the matching relationships between each music feature set, the user's preferences can be shown more accurately. Moreover, using a music feature graph to display multiple music feature sets and their matching relationships makes it easier to show users low-quality music feature sets, which in turn helps the server generate target music that meets the user's needs based on the music feature sets selected by the user.

[0009] In this way, the server can generate the target music that the user needs based on the set of target music features selected by the user, thereby improving the user experience.

[0010] In one possible implementation of the first aspect described above, the multiple music feature sets include a first music feature set and a second music feature set. The server determines a first matching relationship between the multiple music feature sets based on these multiple music feature sets, including: the server determining a first degree of matching between the first music feature set and the second music feature set based on the first music set corresponding to the first music feature set, the second music set corresponding to the second music feature set, and the stored music library; the server determining a second degree of matching between the first music feature set and the second music feature set based on the first music set corresponding to the first music feature set, the second music set corresponding to the second music feature set, and the stored user playlists; and the server determining the first matching relationship between the multiple music feature sets based on the first degree of matching and the second degree of matching between the first music feature set and the second music feature set.

[0011] In some optional embodiments, the first matching degree may represent the proportion of the number of music items corresponding to the music feature set in the music library, and the second matching degree may represent the proportion of the number of music items corresponding to the music feature set in the user's playlist.

[0012] In some optional embodiments, a higher first matching degree indicates that the server is more likely to generate music based on the user-selected music feature set, and the generated music does not contain semantic contradictions or lack aesthetic appeal. For example, a sad melody paired with festive lyrics. A higher second matching degree indicates that the music generated by the server based on the user-selected music feature set is more in line with the user's preferences, and the music generated based on the user's playlist does not contain semantic contradictions or lack aesthetic appeal.

[0013] In one possible implementation of the first aspect above, the server determines a first matching degree between the first music feature set and the second music feature set based on the first music set corresponding to the first music feature set, the second music set corresponding to the second music feature set, and the stored music library. This includes: the server determining the first matching degree based on the ratio of the sum of the number of music pieces corresponding to the first music set and the number of music pieces corresponding to the second music set to the number of music pieces in the music library.

[0014] In some optional embodiments, the matching degree between the first music feature set and the second music feature set includes a first matching degree and a second matching degree. The server can determine a first matching relationship between the first music feature set and the second music feature set based on the first matching degree and the second matching degree. This first matching relationship can be represented in any implementable manner, such as a music feature graph.

[0015] In some optional instances, the server can determine the first matching degree based on the ratio of the sum of the number of music items corresponding to the first music set and the second music set to the number of music items in the music library. For example, if the number of music items corresponding to the first music set is q, the number of music items corresponding to the second music set is p, and the number of music items in the music library is m, then the first matching degree can be (q+p) / m.

[0016] In one possible implementation of the first aspect above, the server determines a second matching degree between the first music feature set and the second music feature set based on the first music set corresponding to the first music feature set, the second music set corresponding to the second music feature set, and the stored user playlist. This includes: the server determining the second matching degree based on the ratio of the sum of the number of music pieces corresponding to the first music set and the number of music pieces corresponding to the second music set to the number of music pieces in the user playlist.

[0017] In some optional instances, the server can determine the second matching degree based on the ratio of the sum of the number of music items corresponding to the first music set and the second music set to the number of music items in the user's playlist. For example, if the number of music items corresponding to the first music set is q, the number of music items corresponding to the second music set is p, and the number of music items in the user's playlist is n, then the first matching degree can be (q+p) / n.

[0018] In one possible implementation of the first aspect above, the terminal device acquires a set of multiple music features corresponding to the user's selection, including: the terminal device displaying a first selection interface, the first selection interface including multiple attribute features, each attribute feature corresponding to one or more music features; the terminal device acquiring a first set of music features selected by the user from the music features corresponding to the multiple attribute features, and using the first set of music features as a first music feature set; the terminal device displaying a second selection interface, the second selection interface including multiple attribute features; and the terminal device acquiring a second set of music features selected by the user from the second selection interface, and using the second set of music features as a second music feature set.

[0019] In some optional instances, the first selection interface may correspond to selection interface 301 mentioned below, and the second selection interface may correspond to selection interface 501 mentioned below.

[0020] In some optional instances, for example, a user can sequentially select vocals (children's voice), subject, scene, back-to-school season, and song title 1 in the first selection interface, i.e., the first set of musical features. Based on the user's selection of the first set of musical features, the terminal device determines that the first set of musical features can be {children's voice_subject_scene_back-to-school season_song title 1}. The terminal device can also determine the second set of musical features as {male voice_subject_year_1980s} based on the user's selection of the second set of musical features, sequentially selecting vocals (male voice), subject, era, and 1980s.

[0021] In one possible implementation of the first aspect above, corresponding to a user selecting a second music feature set after selecting a first music feature set, and the server generating target music based on the first music feature set and the second music feature set, the first weight corresponding to the first music feature set is greater than the second weight corresponding to the second music feature set.

[0022] In some optional instances, to ensure the server can generate high-quality music, during the user's selection of a first music feature set, the terminal device can display a second music feature associated with the first music feature in the second attribute features, based on the first music feature selected by the user in the first attribute features. Therefore, when the server generates target music based on the first and second music feature sets, the first weight corresponding to the first music feature set is greater than the second weight corresponding to the second music feature set.

[0023] In one possible implementation of the first aspect described above, the music features corresponding to the multiple attribute features displayed on the second selection interface are associated with the first set of music features.

[0024] In one possible implementation of the first aspect above, the multiple attribute features include a first attribute feature and a second attribute feature. The terminal device acquires a first set of music features selected by the user from the music features corresponding to the multiple attribute features, including: the terminal device acquiring a first music sub-feature selected by the user from the first music feature of the first attribute feature in a first selection interface; the terminal device displaying a second music feature corresponding to the second attribute feature in the first selection interface, wherein the first music sub-feature is associated with the second music feature; and the terminal device acquiring a second music sub-feature selected by the user from the second music feature corresponding to the second attribute feature.

[0025] In some optional instances, during the process of the user selecting a second set of music features, the terminal device can prioritize music features associated with multiple music features in the first set of music features based on the user's selected first set of music features and a preset feature distribution method, so that the user can choose from them.

[0026] For example, the terminal device can prioritize music features associated with the back-to-school season based on the user's selection of the first music feature set, according to a preset feature distribution method, so that the user can choose.

[0027] In this way, when selecting the first set of music features, it can be ensured that each music feature in the first set of music features is related to the others; when selecting the second set of music features, it can be ensured that the music features in the second set of music features are related to the music features in the first set of music features, thereby effectively improving the quality of the generated music.

[0028] In some optional embodiments, associating the musical features of each musical feature set with the user's selection of multiple musical feature sets can effectively avoid situations where lyrics contradict the melody, lyrics contradict each other, or the generated music lacks aesthetic appeal. For example, this could result in generated music that combines celebratory lyrics with a sad melody, two consecutive lines of lyrics expressing opposite meanings, or generated music that is unpleasant to listen to.

[0029] In one possible implementation of the first aspect above, the attribute features include at least one of the following: category attribute, structure attribute, tag attribute, tag sub-attribute, and song name attribute.

[0030] In one possible implementation of the first aspect above, the musical features corresponding to the category attribute include at least one of the following: lyrics, instruments, vocals; the musical features corresponding to the structure attribute include at least one of the following: introduction, chorus, bridging, main body, coda; and the musical features corresponding to the tag attribute include at least one of the following: back to school, fitness, traveling, office, party, atmosphere, mood.

[0031] In one possible implementation of the first aspect above, the server determines a first matching relationship of multiple music feature sets based on a first matching degree and a second matching degree between the first music feature set and the second music feature set, including: the server determines a third matching degree based on the first matching degree, the second matching degree and a preset matching degree algorithm, and sends the third matching degree to the terminal device; the terminal device displays a first music feature map based on the third matching degree; the first music feature map includes the first music feature set, the second music feature set and the third matching degree.

[0032] In some optional embodiments, the server can determine a third matching degree corresponding to the first matching relationship between the first music feature set and the second music feature set based on the first matching degree, the second matching degree, and a preset matching degree algorithm (e.g., a weighted average algorithm).

[0033] For example, the weight corresponding to the first matching degree is 1, the weight corresponding to the second matching degree is 0, and the third matching degree is determined based on the weight corresponding to the first matching degree, the weight corresponding to the second matching degree, and the weighted average algorithm; wherein, the weight corresponding to the first matching degree can be 0, the weight corresponding to the second matching degree can be 1; or the weight corresponding to the first matching degree can be 0.4, the weight corresponding to the second matching degree can be 0.6, etc.

[0034] Therefore, the server sends the third matching degree to the terminal device, which can then display the first music feature graph to the user based on the third matching degree. In this way, the server can determine whether it is easier to generate music from the first and second music feature sets based on the first matching degree; and it can determine whether the music generated from the first and second music feature sets matches the user's preferences based on the second matching degree.

[0035] In some optional instances, the server can send multiple sets of music features selected by the user and the matching relationships between every two sets of music features to the end device, which can then display the music feature graph to the user.

[0036] In some alternative instances, the music feature map can be generated by the server based on multiple music feature sets and the matching relationship between every two music feature sets. The server then sends the generated music feature map to the terminal device for display.

[0037] In one possible implementation of the first aspect above, the method further includes: the terminal device responding to a user's modification, deletion, or addition operation on music features in the first music feature set, obtaining the modified first music feature set, and sending the modified first music feature set to the server; the server determining a second matching relationship based on the modified first music feature set and the second music feature set, and sending the second matching relationship to the terminal device; and the terminal device displaying the second matching relationship.

[0038] In some optional instances, if the music generated by the server based on the user-selected first and second music feature sets does not match the target music, the user can modify, delete, or add music features in the first and / or second music feature sets via a terminal device. The terminal device then sends the modified music feature map to the server. The server determines a second matching relationship based on the modified music feature map and sends the second matching relationship to the terminal device.

[0039] Furthermore, the terminal device can determine the second music feature map based on the modified music feature map and the second matching relationship. Thus, the user can use the modified first and second music feature sets from the second music feature map as the target music feature set, based on the second music feature map and the second matching relationship. The terminal device then sends the target music feature set to the server. The server can generate the target music based on the user-selected target music feature set.

[0040] In one possible implementation of the first aspect above, the server determines a first matching relationship of multiple music feature sets based on multiple music feature sets, including: if the server determines that the first music feature set in the multiple music feature sets is less than the matching degree threshold of each music feature set, the server deletes the first music feature set and determines the first matching relationship of the multiple music feature sets other than the first music feature set in the multiple music feature sets.

[0041] In some optional instances, the terminal device can be configured for automatic deletion. When determining the music feature map, the server can delete music feature sets from multiple music feature sets where the matching degree with each music feature set is less than a matching degree threshold. The terminal device will then prompt the user with the music feature sets to be deleted, allowing the user to decide whether to proceed.

[0042] In one possible implementation of the first aspect above, the method further includes: if the server determines that the first music feature set in the plurality of music feature sets is less than the matching degree threshold of each music feature set, the server sends a first prompt message to the terminal device, the first prompt message indicating that the first music feature set is less than the matching degree threshold of each music feature set; the display interface of the terminal device displays the first prompt message.

[0043] In some optional instances, when a terminal device displays a music feature map to a user, the terminal device can display a prompt message to the user based on the fact that the matching degree between the music feature set in the multiple music feature sets in the music feature map and each music feature set is less than the matching degree threshold.

[0044] For example, if the matching degree between the second music feature set and the third and fourth music feature sets is low, such as below the matching degree threshold, the terminal device can prompt the user that "the matching degree between the second music feature set and other music feature sets is low, which affects the quality of the generated music."

[0045] In one possible implementation of the first aspect above, the method includes: a terminal device sending a recommendation set instruction to a server in response to a recommendation set operation performed by a user; the server determining a third music feature set based on the recommendation set instruction and the selected first music feature set and second music feature set; the server sending the third music feature set to the terminal device, wherein the matching degree between the third music feature set and the first music feature set and the second music feature set is greater than or equal to a matching degree threshold; and the terminal device displaying the third music feature set to the user.

[0046] In some optional instances, the terminal device, in response to a user's recommendation set operation, sends a recommendation set instruction to the server. The server receives the instruction, generates a recommended music feature set based on the user's selected music feature set, and sends it to the terminal device. The terminal device can then combine the recommended music feature set with the music feature graph to determine a new music feature graph.

[0047] In one possible implementation of the first aspect above, the method further includes: when the user has selected the first music feature set and the second music feature set, the server sends the fourth music feature set to the terminal device based on preset recommendation settings, wherein the matching degree between the fourth music feature set and the first music feature set and the second music feature set is greater than or equal to the matching degree threshold; the terminal device displays the fourth music feature set to the user.

[0048] In some optional instances, the terminal device can be configured with automatic recommendations. When determining the music feature map, the server can send music feature sets to the terminal device that have a matching degree greater than or equal to a matching degree threshold among multiple music feature sets. The terminal device can prompt the user to add recommended music feature sets, and the user can decide whether to add them based on the prompts.

[0049] In this way, during the music generation process, the terminal device can display to the user the matching degree between at least two music feature sets selected by the user. Furthermore, the user can perform operations such as adding, modifying, and deleting multiple music features in the music feature sets based on the matching degree, making the music generated by the server more closely match the user's target music and effectively improving the user experience.

[0050] Secondly, embodiments of this application provide a music generation method applied to a terminal device. The method includes: acquiring multiple music feature sets selected by a user, each music feature set including one or more music features; sending the music feature sets to a server to receive a first matching relationship from the server, the first matching relationship being determined by the multiple music feature sets and including the matching degree between every two music feature sets; displaying the first matching relationship; acquiring a target music feature set selected by the user from the multiple music feature sets based on the first matching relationship, and sending the target music feature set to the server; and receiving target music sent by the server, the target music being generated by the server based on the target music feature set.

[0051] In some optional instances, the terminal device may acquire a set of multiple music features selected by the user through a selection interface, such as the terminal device mentioned below displaying multiple music features to the user through a first selection interface and a second selection interface.

[0052] In some optional embodiments, corresponding to the terminal device acquiring a first set of music features selected by the user from a first selection interface and a second set of music features selected from a second selection interface, the terminal device can send the acquired first and second music feature sets to the server. The server then determines the matching degree between the first and second music feature sets based on the first and second music feature sets, and sends the matching degree to the terminal device.

[0053] Therefore, the terminal device can display a music feature map to the user, including a first music feature set, a second music feature set, and a first matching relationship. The first matching relationship includes the matching degree between the first and second music feature sets. Then, the terminal device can obtain the target music feature set selected by the user based on the music feature map and send the target music feature set to the server. The server can then generate target music based on the target music feature set.

[0054] In this way, by displaying multiple music feature sets, the terminal device can clearly describe the various parts related to the target music and their matching relationships. Furthermore, by displaying the matching relationships between each music feature set, the user's preferences can be displayed more accurately. Moreover, using a music feature graph to display multiple music feature sets and their matching relationships makes it easier to show users low-quality music feature sets, which in turn helps the server generate target music that meets the user's needs based on the music feature sets selected by the user.

[0055] Thirdly, embodiments of this application provide a music generation method applied to a server. The method includes: receiving a set of multiple music features selected by a user from a terminal device, wherein the music feature sets include one or more music features; determining a first matching relationship between the multiple music feature sets based on the multiple music feature sets, wherein the first matching relationship includes the matching degree between every two music feature sets; receiving a target music feature set selected by the user from the multiple music feature sets based on the first matching relationship from the multiple music feature sets via the terminal device; generating target music based on the target music feature set; and sending the target music to the terminal device.

[0056] In some optional instances, the matching degree between any two music feature sets in the multiple music feature sets includes a first matching degree and a second matching degree.

[0057] In some optional instances, a higher first degree of matching indicates that the server is more likely to generate music based on the user's selected music feature set, and that the generated music does not contain semantic contradictions or lack aesthetic appeal. For example, a sad melody paired with celebratory lyrics. A higher second degree of matching indicates that the music generated by the server based on the user's selected music feature set is more in line with the user's preferences, and that the music generated based on the user's playlist does not contain semantic contradictions or lack aesthetic appeal.

[0058] Therefore, the server sends the first and second matching degrees between the determined first and second music feature sets to the terminal device, which then displays them to the user through a display interface. The user can then determine whether further operations are needed on the music feature sets based on the first and second matching degrees displayed on the terminal device, such as adding, modifying, or deleting multiple music features within the sets.

[0059] In this way, the matching degree between each pair of music feature sets can be effectively improved, which helps the generated music to match the target music better and meet the user's preferences.

[0060] Fourthly, this application provides a system comprising a terminal device and a server. The terminal device is used to execute the steps performed by the terminal device in the music generation method mentioned in the first aspect and any one of the first aspects of this application, and the server is used to execute the steps performed by the server in the music generation method mentioned in the first aspect and any one of the first aspects of this application.

[0061] Fifthly, embodiments of this application provide a terminal device, the terminal device comprising: a memory for storing instructions executed by one or more processors of the terminal device; and a processor, one of the processors of the terminal device, for executing the instructions stored in the memory to implement the music generation method mentioned in the second aspect above.

[0062] In a sixth aspect, embodiments of this application provide a server, the server comprising: a memory for storing instructions executed by one or more processors of the server; and a processor, one of the processors of the server, for executing the instructions stored in the memory to implement the music generation method mentioned in the third aspect above.

[0063] In a seventh aspect, this application provides a readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the music generation methods mentioned in any one of the first, second, and third aspects of this application.

[0064] Eighthly, embodiments of this application provide a computer program product, including: a non-volatile computer-readable storage medium containing computer program code for executing the music generation methods mentioned in any one of the first, second, and third aspects of this application. Attached Figure Description

[0065] Figure 1 According to some embodiments of this application, a schematic diagram of a system architecture is shown;

[0066] Figure 2 According to some embodiments of this application, a schematic diagram of a music generation process is shown;

[0067] Figure 3A According to some embodiments of this application, a schematic diagram of a selection interface 301 for a terminal device is shown;

[0068] Figure 3B According to some embodiments of this application, schematic diagrams of selection interfaces 321 to 324 of a terminal device are shown.

[0069] Figure 4 According to some embodiments of this application, a schematic diagram showing the unfolded music features in a selection interface 301 of a terminal device is provided.

[0070] Figure 5 According to some embodiments of this application, a schematic diagram of a selection interface 501 for a terminal device is shown;

[0071] Figure 6A According to some embodiments of this application, a schematic diagram of a selection interface 601 of a terminal device is shown;

[0072] Figure 6B According to some embodiments of this application, a schematic diagram of a selection interface 611 of a terminal device is shown;

[0073] Figure 6C According to some embodiments of this application, a schematic diagram of a semantic association graph is shown;

[0074] Figure 7 According to some embodiments of this application, a schematic diagram of a display interface 701 of a terminal device is shown;

[0075] Figure 8 According to some embodiments of this application, a schematic diagram of a display interface 801 of a terminal device is shown;

[0076] Figure 9A According to some embodiments of this application, a schematic diagram of a prompt window 901 of a display interface 801 of a terminal device is shown;

[0077] Figure 9B According to some embodiments of this application, a schematic diagram of a music feature spectrum is shown;

[0078] Figure 9C According to some embodiments of this application, a schematic diagram of a music feature spectrum is shown;

[0079] Figure 9D According to some embodiments of this application, a schematic diagram of selecting a set of music features based on a display interface 801 is shown;

[0080] Figure 10 According to some embodiments of this application, a schematic diagram of a music generation process is shown;

[0081] Figure 11 According to some embodiments of this application, a schematic diagram of the hardware structure of a terminal device 1000 is shown;

[0082] Figure 12 According to some embodiments of this application, a block diagram of a server 001 is shown. Detailed Implementation

[0083] The illustrative embodiments of this application include, but are not limited to, a music generation method, system, device, medium, and computer program product.

[0084] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0085] The terminal devices mentioned in this application are described below. It is understood that the music generation method provided in this application can be applied to any terminal device, including but not limited to mobile stations (MS), mobile terminals (MT), etc. For example, the terminal device can be a mobile phone, tablet computer, desktop computer, laptop computer, etc. This application does not limit the specific form of the terminal device.

[0086] The following describes the server mentioned in this application. It is understood that the music generation method provided in this application can be applied to any server, including but not limited to application servers, database servers, server clusters, etc. This application does not limit the specific form of the server.

[0087] In some optional instances, Figure 1 An embodiment of this application illustrates a system architecture. This system architecture includes a terminal device and a server. The terminal device can provide users with selectable tags through a selection interface, such as jazz, piano, or party tags. The terminal device can then obtain the user-selected tags and send them to the server. The server then generates music based on the user-selected tags.

[0088] In some embodiments, a music generation application may be installed on the terminal device, and the server may automatically generate music according to the music generation method set by the user in the music generation application.

[0089] The music generation method can be one or more combinations of the following: generating music based on user-selected tags, generating music based on user-input lyrics, generating music based on user-selected singer's timbre, and generating music based on the user's own timbre. This can reduce the cost and difficulty of music production for non-professional producers.

[0090] For example, the server can use deep learning and other algorithms to select and generate at least one piece of male jazz music for an annual meeting from the stored music library based on tags such as male vocals, jazz, and annual meeting selected by the user in the music generation application, so that the user can choose from it.

[0091] For example, the server can select multiple songs from the stored music library that match the timbre of singer A, based on the timbre of singer A selected by the user in the music generation application, and then use algorithms such as deep learning to generate at least one song sung by singer A.

[0092] Similarly, in other instances, servers can generate music based on lyrics entered by users in music generation applications or based on users' own voices, enabling users to create music in a low-cost manner and reducing the difficulty of music production.

[0093] It's understandable that in some cases, the music a user expects may not match the generated music. Users may be unable to accurately describe their desired music, and in such situations, they may have to delete and / or add tags, then use the music generation app again until the generated music matches their expectations. For example, when generating music based on user-selected tags, the user can only repeatedly change the selected tags and perform the music generation process. Furthermore, because the tags are simple in description, it's difficult to precisely control the generated music to move closer to the desired result. Additionally, when users change selected tags, they may blindly modify them due to a lack of understanding of their impact on the generated music, potentially causing the generated music to deviate further from the expected result.

[0094] Therefore, to solve the above problems, this application provides a music generation method. In this method, a terminal device can display multiple music features, allowing the user to select multiple sets of music features from these features. The terminal device can acquire the multiple music feature sets selected by the user, such as two music feature sets, and determine the matching degree between each pair of music feature sets, displaying it to the user so that the user can determine the music features to be selected based on the matching degree of the music feature sets. For example, the user can select two music feature sets with a high matching degree as the music feature sets of the music to be generated. The terminal device can send the music feature sets selected by the user to the corresponding server, and the server generates the music. In this way, the server generates music based on the music feature sets selected by the user, making the music generated by the server more closely match the target music required by the user, effectively improving the user experience. It can be understood that a high matching degree between two music feature sets indicates a high probability that the music generated based on those two music feature sets is the target music.

[0095] In some embodiments, if the music generated by the server based on the user-selected set of music features satisfies the target music, then the user-selected set of music features is the target music feature set corresponding to the target music. If the music generated by the server does not satisfy the target music, the user can add, modify, and delete music features in the selected set of music features through the terminal device. The user can then select the target music feature set from the modified set, allowing the server to generate the target music based on the user-selected target music feature set. Thus, during the music generation process, the terminal device can display the matching degree between at least two music feature sets selected by the user. Furthermore, the user can perform operations such as adding, modifying, and deleting multiple music features in the music feature set based on the matching degree, making the music generated by the server more closely match the user's desired target music and effectively improving the user experience.

[0096] In some embodiments, the terminal device displays multiple attribute features to the user, each attribute feature containing multiple musical features. These attribute features may include category attributes, structural attributes, tag attributes, tag sub-attributes, song title attributes, etc. Musical features corresponding to category attributes may include lyrics, instruments, vocals, or overall (i.e., lyrics + instruments + vocals). Musical features corresponding to structural attributes may include introduction, chorus, bridging, main body, coda, or overall (i.e., introduction + chorus + bridging + main body + coda). Musical features corresponding to tag attributes may include artist name, scene, group, era, function, atmosphere, mood, etc. Tag sub-attributes are further subdivisions of tag attributes; for example, the scene tag attribute can be further subdivided into tag sub-attributes such as back-to-school season, fitness, travel, office, party, celebration, sadness, etc. Therefore, the musical features corresponding to the further subdivisions of the musical features corresponding to tag attributes can be called the musical features corresponding to the tag sub-attributes. The terminal device can obtain a set of multiple musical features selected by the user from the multiple musical features, and then send the selected set of multiple musical features to the server. The server can then generate music based on the set of musical features selected by the user.

[0097] For example, if a user selects the following music features from category attributes, structure attributes, tag attributes, tag sub-attributes, and song name attributes: instrument (erhu), coda, scene, travel, and first song, the terminal device can determine the first music feature set as {erhu_coda_scene_travel_first song}. Similarly, the terminal device can determine the second music feature set as {male voice_main_year_1980s} based on the user's selected second music feature set: instrument (piano), chorus, scene, and wedding. Furthermore, the terminal device can send the user's selected first, second, and third music feature sets to the server, where the server determines the matching degree between each pair of music feature sets and displays it on the terminal device. For example, the matching degree from high to low can be as follows: 90% matching degree between the first and second music feature sets, 70% matching degree between the second and third music feature sets, and 50% matching degree between the first and third music feature sets. Therefore, the user can select the first and second music feature sets as the target music feature sets to generate the target music based on the higher matching degree displayed on the terminal device. Alternatively, the user can perform operations such as adding, modifying, or deleting multiple music features in the first, second, and / or third music feature sets. The terminal device will then send the modified music feature sets as the target music feature sets to the server, which will then generate the target music.

[0098] In some instances, a user sequentially selecting at least one music feature from attributes such as category, structure, tag, tag sub-attributes, and song name can be considered as performing a first selection operation. Each first selection operation corresponds to the terminal device acquiring a music feature set. To ensure the generated music better matches the user's desired target music, the following example illustrates how the terminal device acquires four music feature sets in response to four first selection operations by the user: a first music feature set, a second music feature set, a third music feature set, and a fourth music feature set. It can be understood that each music feature set comprises at least one of the following attributes: category, structure, tag, tag sub-attributes, and song name. Each attribute can include multiple music features.

[0099] In some optional instances, when the server generates music based on a first music feature set, a second music feature set, a third music feature set, and a fourth music feature set, the first weight corresponding to the first music feature set is greater than the second weight corresponding to the second music feature set, the second weight corresponding to the second music feature set is greater than the third weight corresponding to the third music feature set, and the third weight corresponding to the third music feature set is greater than the fourth weight corresponding to the fourth music feature set. Furthermore, the order of these weights corresponds to the order in which the user selects the music feature sets; the user's first execution of the first selection operation corresponds to the terminal device acquiring the first music feature set.

[0100] In some optional instances, the matching degree between each two sets of musical features may include a first matching degree (e.g., public matching degree) and a second matching degree (e.g., private matching degree), the first matching degree being different from the second matching degree.

[0101] The first matching degree can represent the proportion of music items corresponding to a specific music feature set in the music library. For example, the server can determine the first matching degree between the first and second music sets based on the ratio of the sum of the music items corresponding to the first and second music sets to the total number of music items in the music library. The second matching degree can represent the proportion of music items corresponding to a specific music feature set in the user's playlist. Again, the server can determine the second matching degree based on the ratio of the sum of the music items corresponding to the first and second music sets to the total number of music items in the user's playlist.

[0102] In some optional instances, a higher first matching degree indicates that the server is more likely to generate music based on the user-selected music feature set, and the generated music is free from semantic contradictions and aesthetic flaws. For example, a sad melody paired with celebratory lyrics. A higher second matching degree indicates that the music generated by the server based on the user-selected music feature set is more in line with the user's preferences, and the music generated based on the user's playlist is free from semantic contradictions and aesthetic flaws. Therefore, by adding, modifying, and deleting multiple music features in each music feature set based on the first and second matching degrees displayed on the terminal device, users can effectively improve the matching degree between each pair of music feature sets, thus making the generated music more closely match the target music and more in line with the user's preferences.

[0103] In some instances, to ensure a higher degree of matching between any two music feature sets and to improve the quality of the generated music, the category attribute, structural attribute, tag attribute, tag sub-attribute, and song name attribute are correlated, and the music features within each music feature set are also correlated. For example, during the process of a user selecting a first music feature set, the terminal device obtains the user's selection of a first music sub-feature from the first music features of the first attribute set. When the user selects a second attribute feature, the terminal device can display the second music feature that is related to the first music sub-feature. After the user selects the first music feature set, during the process of the user selecting the second music feature set, music features related to multiple music features in the first music feature set can be displayed in the attribute features. For example, during the process of the user selecting the first music feature set, if the user selects "children's voice" in the category attribute and "scene" in the tag attribute, the terminal device can, since the user has already selected "children's voice" in the category attribute, not display "office" in the tag sub-attribute, and prioritize music features related to "children's voice," such as "back to school season," for the user to choose from. Therefore, during the process of the user selecting the second music feature set, the terminal device can select the tag sub-attribute "back to school" in the first music feature set, and not show the user tag sub-attributes such as "fitness" which are different from the music described by "back to school".

[0104] It is understood that the above-mentioned terminal device does not display music features such as "office" to the user only as an example. In some specific implementations, the terminal device may display music features that are not relevant, but the user cannot select them, or the terminal device may prioritize music features that are relevant and hide music features that are not relevant, but the user can select them, etc. This application does not impose any specific restrictions on this.

[0105] It is understood that a terminal device can obtain the set of music features selected by the user through user clicks on the terminal device's display interface, selection interface, etc., or respond to user operations such as deleting, adding, modifying, or generating music features. The terminal device can also respond to user voice input. Alternatively, it can obtain the set of music features selected by the user and respond to user operations in other ways. This application does not impose specific limitations. To avoid repetition, the following describes the terminal device responding to user clicks, obtaining the set of music features selected by the user, and responding to user operations.

[0106] The following is combined with Figure 1 The system architecture shown provides a detailed description of a music generation method mentioned in this application.

[0107] In some optional instances, such as Figure 2 The diagram illustrates a process for generating music, which may include:

[0108] S201: The terminal device obtains the multiple music feature sets selected by the user and sends the multiple music feature sets to the server.

[0109] In some alternative embodiments, such as Figure 3A The diagram illustrates a selection interface 301 for a music generation application on a terminal device (which may correspond to the first selection interface mentioned above). The terminal device can display attribute features to the user, such as category attributes, structural attributes, tag attributes, tag sub-attributes, and song name attributes. Each attribute feature contains multiple musical features. For example, the musical features corresponding to the category attribute may include lyrics, instruments, vocals, and overall structure. The musical features corresponding to the structural attribute may include introduction, chorus, bridging, main body, coda, and overall structure. The musical features corresponding to the tag attribute may include artist name, scene, group, era, function, atmosphere, and mood. Tag sub-attributes can be further subdivided from tag attributes; for example, the scene tag attribute can be further subdivided into tag sub-attributes such as back-to-school season, fitness, travel, office, party, celebration, and sadness. The musical features corresponding to the song name attribute can be song names generated by the server based on the musical features selected by the user, such as song name 1, song name 2, song name 3, song name 4, song name 5, etc.

[0110] In some optional embodiments, the server can generate a song title based on multiple music features corresponding to user-selected category attributes, structural attributes, tag attributes, tag sub-attributes, etc., and combine this with algorithms such as deep learning, and then send the song title to the terminal device. The terminal device can then display the generated song title in the song title attribute.

[0111] In some alternative embodiments, the selection interface displayed by the terminal device to the user may be, for example... Figure 3A The selection interface 301 shown displays multiple attribute features and their corresponding musical features. In some optional embodiments, the terminal device can sequentially display the attribute features and their corresponding musical features to the user according to a set arrangement. For example... Figure 3B The diagram illustrates selection interfaces 321 to 324 for a music generation application on a terminal device. It is understood that the terminal device can progressively display attribute features to the user; for example, based on a pre-defined arrangement, the terminal device may first display… Figure 3B The selection interface 321 shown in (a) can display music features corresponding to category attributes and music features corresponding to structure attributes. Corresponding to the user's selection of music features in the structure attributes, the terminal device can convert the selection interface 321 into... Figure 3BThe selection interface 322 shown in (b) further displays the music features corresponding to the tag attributes to the user for selection.

[0112] Similarly, the terminal device can sequentially display to the user Figure 3B The selection interface 323 shown in (c) is as follows. Figure 3B The selection interface 324 is shown in (d) above. In selection interface 323, the user can further select the music feature corresponding to the tag sub-attribute, and then in selection interface 324, select the music feature corresponding to the song name attribute. Therefore, the selection interface displayed to the user by the terminal device can be as follows: Figure 3A and 3B As shown. To avoid repetition, below... Figure 3A The selection interface shown below will be used as an example for explanation.

[0113] In some specific implementations, the terminal device can obtain a set of multiple music features selected by the user from multiple music features, and send the set of multiple music feature sets to the server. The server can then generate target music based on the target music feature set selected by the user. It can be understood that the music feature set includes one or more music features.

[0114] In some specific implementations, when a user clicks on the component corresponding to a music feature, the terminal device can display the music sub-feature corresponding to the clicked music feature to the user in the selection interface 301. For example, such as... Figure 4 As shown, corresponding to the user clicking on the instrument component 311, the terminal device can further display musical sub-features of instruments such as piano, violin, and erhu in the selection interface 301. For example, human voice can include musical sub-features such as male voice, female voice, and children's voice; female voice can be further divided into soprano, mezzo-soprano, and contralto (not shown). In this way, further subdividing musical features can effectively improve the quality of music generated by the server.

[0115] In some specific implementations, for example, users can select human voice (children's voice), subject, scene, back-to-school season, and song name 1 in sequence on the selection interface 301, which is the first set of music features. The terminal device determines that the first music feature set can be {children's voice_subject_scene_back-to-school season_song name 1} based on the first set of music features selected by the user.

[0116] In some optional embodiments, corresponding to the user selecting a first set of music features, the user can click the "Enter Next Page" component 302 in the selection interface 301 to select a second set of music features. For example, Figure 5The selection interface 501 shown (which can correspond to the second selection interface mentioned above) allows users to sequentially select lyrics, theme, scene, and schoolboy (not shown), i.e., the second set of musical features. Based on the second set of musical features selected by the user, the terminal device determines that the second musical feature set can be {lyrics_theme_scene_schoolboy}. Then, the user can click the "Enter Next Page" component 502 in the selection interface 501 to select the third set of musical features. Alternatively, the user can click the "Complete" component 503 in the selection interface 501 without selecting the third set of musical features. In response to the user clicking the "Complete" component 503, the terminal device can display the musical feature graphs corresponding to the first and second musical feature sets selected by the user (e.g., as shown below). Figure 7 The musical feature diagram shown is 702. For details, please refer to S203 below. To avoid repetition, it will not be elaborated further here.

[0117] In some alternative embodiments, such as Figure 6A The selection interface 601 shown corresponds to the user selecting the third set of music features. The user can sequentially select lyrics, chorus, function, and encouragement (not shown) in selection interface 601, i.e., the third set of music features. Based on the user's selection of the third set of music features, the terminal device determines that the third set of music features can be {lyrics_chorus_function_encouragement}. Then, the user can click the "Enter Next Page" component 602 in selection interface 601 to select the fourth set of music features.

[0118] In some alternative embodiments, the user can click the completion component 603 in the selection interface 601 without selecting the fourth music feature set. The terminal device can respond to the user clicking the completion component 603 by displaying the music feature graphs corresponding to the first, second, and third music feature sets selected by the user. For details, please refer to S203 below; to avoid repetition, further elaboration will not be provided here.

[0119] Similarly, users in Figure 6B When selecting the fourth music feature set in the selection interface 611 shown, the terminal device can determine the fourth music feature set as {female voice_main body_celebrity name_female celebrity A} based on the user's selection of the fourth music feature set as female voice, main body, celebrity name, female celebrity A (not shown).

[0120] In some optional embodiments, the user can click the completion component 612 in the selection interface 611. In response to the user clicking the completion component 612, the terminal device can display to the user the music feature spectra corresponding to the first music feature set, the second music feature set, the third music feature set, and multiple fourth music feature sets selected by the user (e.g., as described below). Figure 8The musical feature diagram shown is 802. For details, please refer to S203 below. To avoid repetition, it will not be elaborated further here.

[0121] Therefore, by using the set of music features selected by the user multiple times, the music generated by the server can be more closely matched with the target music.

[0122] In some specific implementations, to ensure that the server can generate high-quality music, during the process of the user selecting the first music feature set, the terminal device can display the second music feature associated with the first music feature in the second attribute feature, based on the first music feature selected by the user in the first attribute feature.

[0123] For example, in Figure 3A In the selection interface 301 shown, the terminal device can recommend tag attributes associated with the selected category attribute to the user. For example, if the user selects human voice (children's voice) as a music feature in the category attribute, and the user selects a scene in the tag attribute, the terminal device can recommend music features associated with children's voices in the tag sub-attributes, such as "back to school season." Alternatively, the terminal device can prioritize music features associated with children's voices according to a preset feature distribution method, for example, placing "back to school season" first for the user to choose from.

[0124] In other specific implementations, during the user's selection of a second music feature set, the terminal device can prioritize music features associated with multiple music features in the first music feature set based on the user's selected first music feature set and a preset feature distribution method, allowing the user to choose accordingly. For example, in Figure 5 In the selection interface 501 shown, the terminal device can prioritize music features associated with the "back to school season" tag sub-attributes according to a preset feature distribution method, based on the user's selection of the "back to school season" in the first music feature set, for the user to choose from. This ensures that each music feature in the first music feature set is associated with the others, and that the music features in the second music feature set are associated with the music features in the first music feature set, thereby effectively improving the quality of the generated music.

[0125] It is understandable that when a user selects a third or fourth set of music features, the terminal device can display music features associated with the selected set. The terminal device can then send the user-selected first, second, third, and fourth sets of music features to the server.

[0126] In some optional embodiments, associating the musical features of each musical feature set with the user's selection of multiple musical feature sets can effectively avoid situations where lyrics contradict the melody, lyrics contradict each other, or the generated music lacks aesthetic appeal. For example, this could result in generated music that combines celebratory lyrics with a sad melody, two consecutive lines of lyrics expressing opposite meanings, or generated music that is unpleasant to listen to.

[0127] It is understandable that during the process of a user selecting a set of music features, the terminal device can obtain the attribute features that the user has not selected as a whole according to the preset selection method. For example, corresponding to the type attribute that the user has not selected, the terminal device can obtain the type feature that the user has selected as a whole, namely lyrics, instruments and vocals.

[0128] S202: The server determines the first matching relationship of multiple music feature sets based on multiple music feature sets, and sends the first matching relationship to the terminal device.

[0129] In some optional embodiments, the server can determine a first matching relationship among multiple music feature sets selected by the user. This first matching relationship may include the matching degree between every two music feature sets. For example, the user can sequentially select a first music feature set, a second music feature set, a third music feature set, and a fourth music feature set. The determination of the first matching relationship between the first and second music feature sets will be illustrated below.

[0130] In some optional embodiments, the matching degree between the first music feature set and the second music feature set includes a first matching degree and a second matching degree. The server can determine a first matching relationship between the first music feature set and the second music feature set based on the first matching degree and the second matching degree. This first matching relationship can be represented in any implementable manner, such as a music feature graph.

[0131] In some optional instances, the server can determine the first matching degree between the first music feature set and the second music feature set based on the first music set corresponding to the first music set, the second music set corresponding to the second music feature set, and the music library stored in the server.

[0132] In some optional instances, the server can determine the first matching degree based on the ratio of the sum of the number of music items corresponding to the first music set and the second music set to the number of music items in the music library. For example, if the number of music items corresponding to the first music set is q, the number of music items corresponding to the second music set is p, and the number of music items in the music library is m, then the first matching degree can be (q+p) / m.

[0133] In some optional instances, the server can determine a second matching degree between the first music feature set and the second music feature set based on the first music set corresponding to the first music set, the second music set corresponding to the second music feature set, and the user playlists stored in the server.

[0134] In some optional instances, the server can determine the second matching degree based on the ratio of the sum of the number of music items corresponding to the first music set and the second music set to the number of music items in the user's playlist. For example, if the number of music items corresponding to the first music set is q, the number of music items corresponding to the second music set is p, and the number of music items in the user's playlist is n, then the first matching degree can be (q+p) / n.

[0135] Furthermore, the server can determine the third matching degree corresponding to the first matching relationship between the first music feature set and the second music feature set based on the first matching degree, the second matching degree, and a preset matching degree algorithm (such as a weighted average algorithm). For example, the weight corresponding to the first matching degree is 1, the weight corresponding to the second matching degree is 0, and the third matching degree is determined based on the weight corresponding to the first matching degree, the weight corresponding to the second matching degree, and the weighted average algorithm; wherein the weight corresponding to the first matching degree can be 0, the weight corresponding to the second matching degree can be 1; or the weight corresponding to the first matching degree can be 0.4, the weight corresponding to the second matching degree can be 0.6, etc.

[0136] Therefore, the server sends the third matching score to the terminal device, which can then display a music feature graph (such as the first music feature graph mentioned above) to the user based on the third matching score. For example, corresponding to the user's... Figure 6A When the user clicks the "OK" component 603 on the selection interface 601, the terminal device can respond by sending a "Confirm Matching Degree" command to the server. The server can receive the command, determine a third matching degree based on the first and second matching degrees between the first and second music feature sets, and send the third matching degree to the terminal device. Then, the terminal device can display the corresponding music feature graph to the user based on the first, second, and third matching degrees.

[0137] Thus, the server can determine whether it is easier to generate music from the first set of music features and the second set of music features based on the first matching degree; and can determine whether the music generated from the first set of music features and the second set of music features matches the user's preferences based on the second matching degree.

[0138] In some optional instances, a higher first degree of matching indicates that, based on the user's selected set of music features, the server is more likely to generate music that is semantically incompatible or aesthetically unappealing. A higher second degree of matching indicates that, based on the user's selected set of music features, the music generated by the server better matches the user's preferences; and, based on the user's playlist, the music generated is semantically incompatible or aesthetically unappealing. For example... Figure 6C The semantic association graph shown corresponds to multiple music features in the first and second music feature sets, including festive music features. Therefore, when generating music, the server can... Figure 6C The generated music incorporates elements that match the musical characteristics of weddings, travel, and celebrations, as shown in (a). Furthermore, it avoids... Figure 6C As shown in (a), themes such as weddings and sorrow lead to a combination of celebratory lyrics and melancholic melodies in the generated music. This effectively improves the quality of the generated music.

[0139] In some optional instances, the server can send multiple sets of music features selected by the user and the matching relationships between every two sets of music features to the end device, which can then display the music feature graph to the user.

[0140] In some alternative instances, the music feature map can be generated by the server based on multiple music feature sets and the matching relationship between every two music feature sets. The server then sends the generated music feature map to the terminal device for display.

[0141] S203: The terminal device displays the first matching relationship.

[0142] In some optional instances, the music feature maps corresponding to the first music feature set, the second music feature set, and the third matching degree can be as follows: Figure 7 As shown, in Figure 7 In the display interface 701, a music feature map 702 and a music generation component 713 may be included. The music feature map 702 may include a first music feature set 711, a second music feature set 712, and a third matching degree (e.g., 90%). The user can then click the music generation component 713. In response to this action, the terminal device sends a music generation command to the server. The server can then execute step S205, generating the target music based on the target music feature set.

[0143] In some optional instances, corresponding to the user selecting a first music feature set, a second music feature set, a third music feature set, and a fourth music feature set, the server can determine the matching relationship between each pair of music feature sets and send it to the terminal device, which then displays the matching relationship.

[0144] For example, Figure 8 A schematic diagram of a music feature map is shown, used to represent the matching relationship between every two music feature sets in multiple music feature sets. Figure 8 In the display interface 801, a music feature graph 802 and a music generation component 813 may be included. The music feature graph 802 may include a first music feature set 821, a second music feature set 822, a third music feature set 823, and a fourth music feature set 824. In some implementations, the degree of matching between two music feature sets can be represented by the thickness of the lines between them.

[0145] In some optional instances, in the music feature map 801, the highest matching degree can be determined based on the thickness of association lines 811, 814, and 813, indicating the highest matching degree between the first and second music feature sets, the highest matching degree between the first and third music feature sets, and the highest matching degree between the third and fourth music feature sets. A relatively thick association line 815 indicates a high matching degree between the first and fourth music feature sets. A relatively thin association line 812 indicates a low matching degree between the second and fourth music feature sets. And the thinnest association line 816 indicates the lowest matching degree between the second and fourth music feature sets.

[0146] In some optional instances, when a terminal device displays a music feature map to a user, it can display a prompt message if the matching degree between a music feature set and each of the multiple music feature sets in the music feature map is less than a matching degree threshold. For example, if the matching degree between the second music feature set and the third and fourth music feature sets is low, such as less than the matching degree threshold, the terminal device can prompt the user that "the matching degree between the second music feature set and other music feature sets is low, affecting the quality of the generated music." Figure 9A As shown in the prompt window 901, the user can click the return component 911 to return to the display interface 801 and perform modification, addition, or deletion operations on the music features corresponding to the second music feature set. The user can also click the delete music feature set component 912, and the terminal device will delete the second music feature set in response to the user's click to delete the music feature set component 912.

[0147] In some alternative instances, if the terminal device acquires multiple music feature sets selected by the user, including music feature set T, music feature set W, music feature set X, music feature set Y, and music feature set Z; then the music feature map determined based on music feature set T, music feature set W, music feature set X, music feature set Y, and music feature set Z can be as follows: Figure 9B As shown, when the matching degree between music feature set X and music feature sets T, W, Y, and Z is all less than the matching degree threshold, the terminal device can display a prompt message to the user. This prompt message indicates that music feature set X is less than the matching degree threshold with each of the other music feature sets. Therefore, the user can delete music feature set X based on the prompt.

[0148] In some optional instances, in Figure 8 In the display interface 801, a recommendation component 826 is included. Users can click on the recommendation component 826. In response to this click, the terminal device sends a recommendation set instruction to the server. The server receives the recommendation set instruction and, based on the user's selected music feature set, generates a fifth music feature set, which is then sent to the terminal device. The terminal device can combine the fifth music feature set with the music feature graph 802 to determine a new music feature graph.

[0149] In other optional instances, the terminal device can be configured to automatically delete music feature sets from multiple music feature sets when determining the music feature map, where the matching degree with each music feature set is less than a matching degree threshold. A first prompt message will then be displayed, prompting the user to confirm whether the music feature sets to be deleted. The user can then decide whether to proceed with the deletion based on the prompt.

[0150] In some optional instances, the terminal device can be configured with automatic recommendations. When determining the music feature map, the server can send music feature sets to the terminal device that have a matching degree greater than or equal to a matching degree threshold among multiple music feature sets. The terminal device can prompt the user to add recommended music feature sets, and the user can decide whether to add them based on the prompts.

[0151] In some specific implementations, the matching relationship between multiple music feature sets selected by the user and every two music feature sets can also be displayed to the user in the form of a list, and this application does not impose specific restrictions.

[0152] It is understood that the matching relationship between two sets of musical features can also be displayed to the user through different colors, specific numerical values ​​corresponding to the matching degree, combinations of different colors and specific numerical values ​​corresponding to the matching degree, or combinations of line thickness and specific numerical values ​​corresponding to the matching degree, etc., and this application does not impose specific limitations. For example, the terminal device can display the matching relationship between two sets of musical features in the order of red, yellow, blue, and black, corresponding to a sequential decrease from high to low. For another example... Figure 9C The diagram illustrates a musical feature graph. A terminal device can display the matching relationship between two musical feature sets to the user by combining line thickness with numerical values ​​corresponding to specific matching degrees. For example, in... Figure 9C In the example, the terminal device displays the thickest connection line between the first and second music feature sets, indicating a 90% match rate. Alternatively, the terminal device can display the thinnest connection line between the first and third music feature sets, indicating a 30% match rate, and so on.

[0153] The matching degree, which reflects the matching relationship between two music feature sets, can be determined by the terminal device according to a preset matching degree algorithm. For example, the matching degree can be determined based on a first matching degree, a second matching degree, and a weighted average algorithm.

[0154] In some optional instances, users can click on a music feature set to view music related to that set. For example, a user can click on the first music feature set 821 to view music related to that set.

[0155] Thus, by displaying multiple sets of musical features, the various parts related to the target music and their matching relationships can be described more clearly. Furthermore, by displaying the matching relationships between each set of musical features, the user's preferences can be shown more accurately. Moreover, using a musical feature graph to display multiple sets of musical features and their matching relationships makes it easier to show users low-quality sets of musical features, thereby facilitating the user's modification, deletion, and addition of features from these low-quality sets.

[0156] S204: The terminal device obtains the target music feature set selected by the user from multiple music feature sets based on the first matching relationship, and sends the target music feature set to the server.

[0157] In some optional instances, the terminal device acquires the first, second, third, and fourth music feature sets selected by the user, and displays the matching relationship between each pair of music feature sets through a music feature graph 802. The user can select at least two music feature sets as target music feature sets based on the matching relationship. Then, the terminal device acquires the target music feature sets selected by the user and sends the target music feature sets to the server.

[0158] For example, a user can use the first music feature set and the second music feature set as the target music feature set based on the matching degree between them being greater than a matching degree threshold. For instance, as shown... Figure 9D As shown, Figure 8 The first and second music feature sets are highlighted (marked as gray). For example, a user can select the first and third music feature sets as the target music feature sets based on the matching degree between them exceeding a matching degree threshold. Alternatively, the user can choose to use the first, second, third, and fourth music feature sets as the target music feature sets.

[0159] S205: The server generates target music based on the target music feature set.

[0160] In some optional instances, corresponding to the user selecting a first music feature set and a second music feature set as the target music feature set, the terminal device can send the target music feature set corresponding to the first music feature set and the second music feature set to the server. The server can then generate the target music based on the target music feature set.

[0161] For example, see Figure 9D The display interface 801 shown allows users to sequentially click on the first music feature set 821, the second music feature set 822, and the music generation component 825. The terminal device responds to the user's clicks by sending the target music feature sets corresponding to the first and second music feature sets to the server. The server can then generate the target music based on the first and second music feature sets.

[0162] In some optional instances, if the music generated by the server based on the user-selected first and second music feature sets does not match the target music, the user can modify, delete, or add music features in the first and / or second music feature sets via a terminal device. The terminal device then sends the modified music feature graph to the server. The server determines a second matching relationship based on the modified music feature graph and sends this relationship to the terminal device. Furthermore, the terminal device can determine a second music feature graph based on the modified music feature graph and the second matching relationship. Thus, the user can use the modified first and second music feature sets in the second music feature graph as the target music feature set based on the second music feature graph and the second matching relationship. The terminal device then sends the target music feature set to the server. The server can then generate the target music based on the user-selected target music feature set.

[0163] For example, if the music generated by the server based on the user's selected first and second music feature sets does not match the target music, the user can return to display interface 801 and click the recommendation component 826. In response to the user clicking the recommendation component 826, the terminal device sends a recommendation set instruction to the server. The server receives the recommendation set instruction and generates a recommended music feature set based on the user's selected music feature set, then sends it to the terminal device. The terminal device can combine the recommended music feature set and the music feature graph 802 to determine a new music feature graph. The user can then select the target music feature set again, and the terminal device will send the target music feature set to the server. The server can then generate the target music based on the user's selected target music feature set.

[0164] In some optional instances, if the music generated by the server based on the user's selected first and second music feature sets does not match the target music, the user is returned to display interface 801. The user can then select the second music feature set 822 based on the lowest matching degree between the second and fourth music feature sets in the music feature graph, and perform operations such as adding, modifying, or deleting multiple music features in the second music feature set. The terminal device then sends the modified second music feature set to the server. The server determines the matching degree between each pair of music feature sets based on the modified second music feature set and sends this result back to the terminal device. The terminal device can then determine a new music feature graph. The user can then select the target music feature set again, and the terminal device will send the target music feature set to the server. The server can then generate the target music based on the user's selected target music feature set.

[0165] In other optional instances, the server can also generate target music based on a user-selected set of target music features, combined with information such as sound, text, and video. For example... Figure 10 The flowchart shown illustrates that the terminal device can acquire user-inputted audio and text, as well as user-uploaded videos or videos stored on the terminal device, and then send this information to the server. The server can then generate target music based on the target music feature set selected by the user in the music feature map, combined with the text, audio, and video information.

[0166] Thus, during the music generation process, the terminal device can display to the user the matching degree between at least two music feature sets selected by the user, that is, to show the user the matching degree between the generated music and the target music. Furthermore, the user can perform operations such as adding, modifying, and deleting multiple music features in the music feature sets based on the matching degree, making the music generated by the server more closely match the user's desired target music, effectively improving the user experience. In addition, compared to manually creating music, the music generation method provided in this application embodiment can effectively reduce the cost and time required for music production.

[0167] The hardware structure of terminal device 1000 will be introduced below.

[0168] Figure 11 A schematic diagram of the hardware structure of terminal device 1000 is shown. It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on terminal device 1000. In other embodiments of this application, terminal device 1000 may include more or fewer components than illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0169] Terminal device 1000 may include a processor 1101, an external memory interface 1102, an internal memory 1103, a universal serial bus (USB) interface 1104, a charging management module 1105, a power management module 1106, a battery 1107, antenna 1, antenna 2, a mobile communication module 1108, a wireless communication module 1109, an audio module 1110, a sensor module 1111, buttons 1112, a motor 1113, an indicator 1114, a camera 1115, a display screen 1116, and a subscriber identification module (SIM) card interface 1117, etc. The sensor module 1111 may include a pressure sensor 1111A, a gyroscope sensor 1111B, a barometric pressure sensor 1111C, a magnetic sensor 1111D, an accelerometer sensor 1111E, a distance sensor 1111F, a proximity sensor 1111G, a fingerprint sensor 1111H, a temperature sensor 1111J, a touch sensor 1111K, an ambient light sensor 1111L, a bone conduction sensor 1111M, etc.

[0170] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0171] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. The processor can be used to execute the music generation method on the terminal device side mentioned in the embodiments of this application.

[0172] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0173] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal device 1000. While charging the battery 142, the charging management module 140 can also supply power to the terminal device 1000 via the power management module 141.

[0174] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0175] The wireless communication function of the terminal device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.

[0176] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 1000 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0177] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the terminal device 1000. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0178] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 1000, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0179] The terminal device 1000 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0180] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, terminal device 1000 may include one or N displays 194, where N is a positive integer greater than 1.

[0181] In some alternative instances, display screen 194 may be used as mentioned in the embodiments of this application. Figure 3A , Figure 3B , Figure 4 , Figure 5 , Figure 6A , Figure 6B , Figure 7 , Figure 8 , Figure 9A , Figure 9B and Figure 9D The interface shown.

[0182] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device 1000. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0183] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of terminal device 1000 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of terminal device 1000 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0184] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to establish contact with and disconnect from the terminal device 1000. The terminal device 1000 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The terminal device 1000 interacts with the network through the SIM card to achieve functions such as voice calls and data communication. In some embodiments, the terminal device 1000 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal device 1000 and cannot be separated from the terminal device 1000.

[0185] Figure 12 The structure of the server mentioned in the embodiments of this application will be introduced using server 001 as an example.

[0186] Figure 12 This is a block diagram of server 001 provided in an embodiment of this application. In some embodiments, server 001 may include one or more processors 1204, system control logic 1208 connected to at least one of the processors 1204, system memory 1212 connected to system control logic 1208, non-volatile memory (NVM) 1216 connected to system control logic 1208, and network interface 1220 connected to system control logic 1208.

[0187] In some embodiments, processor 1204 may include one or more single-core or multi-core processors. In some embodiments, processor 1204 may include any combination of general-purpose processors and special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In embodiments where server 001 employs an Evolved Node B (eNB) or Radio Access Network (RAN) controller, processor 1204 may be configured to execute the music generation method performed on the server side as mentioned in the embodiments of this application.

[0188] In some embodiments, system control logic 1208 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1204 and / or any suitable device or component communicating with system control logic 1208.

[0189] In some embodiments, system control logic 1208 may include one or more memory controllers to provide an interface to system memory 1212. System memory 1212 may be used to load and store data and / or instructions. In some embodiments, system memory 1212 of server 001 may include any suitable volatile memory, such as suitable dynamic random access memory (DRAM).

[0190] The non-volatile memory (NVM) 1216 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory (NVM) 1216 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.

[0191] The non-volatile memory (NVM) 1216 may include a portion of the storage resources on the device on which the server 001 is mounted, or it may be accessible by the device, but is not necessarily part of the device. For example, the non-volatile memory (NVM) 1216 may be accessed over a network via network interface 1220.

[0192] System memory 1212 and non-volatile memory (NVM) 1216 may each include a temporary copy and a permanent copy of instructions 1224. Instructions 1224 may include instructions that, when executed by at least one of processors 1204, cause server 001 to implement the music generation method mentioned in the embodiments of this application. In some embodiments, instructions 1224, hardware, firmware, and / or their software components may additionally / alternatively be located in system control logic 1208, network interface 1220, and / or processor 1204.

[0193] Network interface 1220 may include a transceiver for providing a radio interface to server 001, thereby enabling communication with any other suitable device (such as a front-end module, antenna, etc.) via one or more networks. In some embodiments, network interface 1220 may be integrated into other components of server 001. For example, network interface 1220 may be integrated into at least one of processor 1204, system memory 1212, non-volatile memory (NVM) 1216, and firmware device (not shown) with instructions that, when at least one of processor 1204 executes the instructions, server 001 implements the music generation method mentioned in the embodiments of this application.

[0194] The network interface 1220 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1220 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.

[0195] In some embodiments, at least one of the processors 1204 may be packaged together with the logic of one or more controllers for system control logic 1208 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1204 may be integrated on the same die with the logic of one or more controllers for system control logic 1208 to form a system-on-a-chip (SoC).

[0196] Server 001 may further include an input / output (I / O) device 1232. The I / O device 1232 may include a user interface enabling a user to interact with server 001; the peripheral component interface is designed to allow peripheral components to also interact with server 001. In some embodiments, server 001 further includes sensors for determining at least one of environmental conditions and location information related to server 001.

[0197] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.

[0198] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.

[0199] In some embodiments, the sensor may include, but is not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of or interact with the network interface 1220 to communicate with components of the positioning network (e.g., Global Positioning System (GPS) satellites).

[0200] This application provides a system including a terminal device and a server. The terminal device executes the steps of the music generation method mentioned in this application, and the server executes the steps of the music generation method mentioned in this application.

[0201] This application provides a computer-readable medium storing instructions that, when executed on a terminal device or server, cause the terminal device or server to perform the steps on the terminal device or server side of the music generation method mentioned in this application.

[0202] This application provides a computer program product, including a non-volatile computer-readable storage medium containing computer program code for performing the music generation method mentioned in this application.

[0203] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0204] It should be noted that, in the examples and description of this patent, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0205] Although this application has been illustrated and described with reference to certain embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.

Claims

1. A method for generating music, characterized in that, The method includes: The terminal device acquires multiple music feature sets selected by the user and sends the multiple music feature sets to the server, wherein the music feature sets include one or more music features; The server determines a first matching relationship between the multiple music feature sets based on the multiple music feature sets, and sends the first matching relationship to the terminal device. The first matching relationship includes the matching degree between every two music feature sets in the multiple music feature sets. The terminal device displays the first matching relationship; The terminal device obtains the target music feature set selected by the user from the plurality of music feature sets based on the first matching relationship, and sends the target music feature set to the server; The server generates target music based on the target music feature set.

2. The method according to claim 1, characterized in that, The plurality of music feature sets include a first music feature set and a second music feature set. The server determines a first matching relationship among the plurality of music feature sets based on the plurality of music feature sets, including: The server determines a first matching degree between the first music feature set and the second music feature set based on the first music feature set corresponding to the first music set, the second music feature set corresponding to the second music set, and the stored music library; The server determines a second matching degree between the first music feature set and the second music feature set based on the first music feature set corresponding to the first music set, the second music feature set corresponding to the second music set, and the stored user playlists. The server determines the first matching relationship of the plurality of music feature sets based on the first matching degree and the second matching degree between the first music feature set and the second music feature set.

3. The method according to claim 2, characterized in that, The server determines a first matching degree between the first music feature set and the second music feature set based on the first music feature set corresponding to the first music set, the second music feature set corresponding to the second music set, and the stored music library, including: The server determines the first matching degree based on the ratio of the sum of the number of music items corresponding to the first music set and the number of music items corresponding to the second music set to the number of music items in the music library.

4. The method according to claim 3, characterized in that, The server determines a second matching degree between the first music feature set and the second music feature set based on the first music feature set corresponding to the first music set, the second music feature set corresponding to the second music set, and the stored user playlists, including: The server determines the second matching degree based on the ratio of the sum of the number of music items corresponding to the first music set and the number of music items corresponding to the second music set to the number of music items in the user's playlist.

5. The method according to claim 2, characterized in that, The terminal device acquires multiple music feature sets corresponding to the user's selection, including: The terminal device displays a first selection interface, which includes multiple attribute features, each of which corresponds to one or more music features. The terminal device obtains a first set of music features selected by the user from the music features corresponding to the multiple attribute features, and uses the first set of music features as the first music feature set. The terminal device displays a second selection interface, which includes multiple attribute features. The terminal device acquires the second set of music features selected by the user from the second selection interface, and uses the second set of music features as the second music feature set.

6. The method according to claim 5, characterized in that, When a user selects a first music feature set and then selects a second music feature set, and the server generates the target music based on the first music feature set and the second music feature set, the first weight corresponding to the first music feature set is greater than the second weight corresponding to the second music feature set.

7. The method according to claim 5, characterized in that, The music features corresponding to the multiple attribute features displayed in the second selection interface are associated with the first group of music features.

8. The method according to claim 5, characterized in that, The multiple attribute features include a first attribute feature and a second attribute feature. The terminal device acquires the first set of music features selected by the user from the music features corresponding to the multiple attribute features, including: The terminal device acquires the first music sub-feature selected by the user from the first music feature of the first attribute feature in the first selection interface; The terminal device displays the second music feature corresponding to the second attribute feature in the first selection interface, and the first music sub-feature is associated with the second music feature; The terminal device obtains the second music sub-feature selected by the user from the second music feature corresponding to the second attribute feature.

9. The method according to any one of claims 5-8, characterized in that, The attribute features include at least one of the following: Category attributes, structure attributes, tag attributes, tag sub-attributes, and song name attributes.

10. The method according to claim 9, characterized in that, The musical features corresponding to the category attributes include at least one of the following: lyrics, instruments, and vocals; The musical features corresponding to the structural attributes include at least one of the following: introduction, chorus, bridging, main body, and coda; The music features corresponding to the tag attributes include at least one of the following: back to school season, fitness, traveling, office, party, atmosphere, mood.

11. The method according to claim 2, characterized in that, The server determines the first matching relationship of the plurality of music feature sets based on the first matching degree and the second matching degree between the first music feature set and the second music feature set, including: The server determines a third matching degree based on the first matching degree, the second matching degree, and a preset matching degree algorithm, and sends the third matching degree to the terminal device; The terminal device displays a first music feature spectrum based on the third matching degree; The first music feature map includes the first music feature set, the second music feature set, and the third matching degree.

12. The method according to claim 11, characterized in that, The method further includes: The terminal device responds to the user's modification, deletion, or addition operations on the music features in the first music feature set by obtaining the modified first music feature set and sending the modified first music feature set to the server; The server determines a second matching relationship based on the modified first music feature set and the second music feature set, and sends the second matching relationship to the terminal device; The terminal device displays the second matching relationship.

13. The method according to claim 1, characterized in that, The server determines a first matching relationship for the multiple music feature sets based on the multiple music feature sets, including: If the server determines that the first music feature set in the plurality of music feature sets is less than the matching degree threshold for each of the music feature sets, the server deletes the first music feature set and determines the first matching relationship corresponding to the plurality of music feature sets other than the first music feature set in the plurality of music feature sets.

14. The method according to claim 13, characterized in that, The method further includes: If the server determines that the first music feature set in the plurality of music feature sets is less than the matching degree threshold for each of the music feature sets, the server sends a first prompt message to the terminal device. The first prompt message is used to indicate that the first music feature set is less than the matching degree threshold for each of the music feature sets. The terminal device displays the first prompt message on its display interface.

15. The method according to claim 14, characterized in that, The method includes: In response to the user's recommended set operation, the terminal device sends a recommended set instruction to the server; Based on the recommendation set instruction, the server determines the third music feature set according to the selected first and second music feature sets. The server sends the third music feature set to the terminal device, wherein the matching degree between the third music feature set and the first music feature set and the second music feature set is greater than or equal to the matching degree threshold. The terminal device displays the third music feature set to the user.

16. The method according to claim 15, characterized in that, The method further includes: When the user has selected the first music feature set and the second music feature set, the server sends the fourth music feature set to the terminal device based on the preset recommendation settings. The matching degree of the fourth music feature set with the first music feature set and the second music feature set is greater than or equal to the matching degree threshold. The terminal device displays the fourth music feature set to the user.

17. A music generation method, applied to a terminal device, characterized in that, The method includes: Obtain multiple music feature sets selected by the user, wherein the music feature sets include one or more music features; Send a set of music features to the server to receive a first matching relationship from the server. The first matching relationship is determined by the plurality of music feature sets and includes the matching degree between every two music feature sets in the plurality of music feature sets. Display the first matching relationship; Obtain the target music feature set selected by the user from the plurality of music feature sets based on the first matching relationship, and send the target music feature set to the server; The server receives target music sent by the server, which is generated by the server based on the target music feature set.

18. A music generation method, applied to a server, characterized in that, The method includes: The receiving terminal device sends a set of multiple music features selected by the user, wherein the set of music features includes one or more music features; Based on the multiple music feature sets, a first matching relationship of the multiple music feature sets is determined, and the first matching relationship includes the matching degree between every two music feature sets in the multiple music feature sets; Receive the target music feature set sent by the terminal device, which is selected by the user from the plurality of music feature sets based on the first matching relationship; Based on the target music feature set, generate the target music; Send the target music to the terminal device.

19. A system, characterized in that, The system includes a terminal device and a server. The terminal device is used to execute the steps performed by the terminal device in any one of the music generation methods of claims 1 to 16, and the server is used to execute the steps performed by the server in any one of the music generation methods of claims 1 to 16.

20. A terminal device, characterized in that, include: A memory for storing instructions executed by one or more processors of the terminal device, wherein the processor is one of the one or more processors of the terminal device for performing the music generation method of claim 17.

21. A server, characterized in that, include: A memory for storing instructions executed by one or more processors of the terminal device, wherein the processor is one of the one or more processors of the terminal device for executing the music generation method of claim 18.

22. A computer-readable medium, characterized in that, The readable medium stores instructions that, when executed on a computer, cause the computer to perform the music generation method according to any one of claims 1 to 18.

23. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the music generation method according to any one of claims 1 to 18.