Music generation method, music generation device, electronic equipment and storage medium

By obtaining the emotional characteristics of classical literature and real-time hot texts, the music description text is generated, and the problem of single style of music generation model is solved, and high-quality and diverse music generation is achieved.

CN120544523APending Publication Date: 2025-08-26SHENZHEN CONSYS SCI&TECH CO LTD

Patent Information

Application Number
CN202510413905.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the music sample size in the open source data set is limited and low-quality, resulting in a single music style generated by the music generation model, which is difficult to meet the user's needs for diverse music types.

Method used

By obtaining classical literary texts and real-time hot texts, emotional characteristics are extracted and integrated respectively, the emotional characteristics of the literary text of classical literary texts are used to enrich the emotional characteristics of real-time hot texts, and the music description text is generated based on scene prompts, and the target music is finally generated.

Benefits of technology

It improves the quality and diversity of music generation, ensures that the music is highly consistent with user needs, and overcomes the shortcomings of the existing technology in music complexity and emotional richness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544523A_ABST
    Figure CN120544523A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a music generation method, a music generation device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring a classical literature text and a real-time hotspot text; then extracting sentiment features of the classical literature text to obtain sentiment features of the literature text, and extracting sentiment features of the real-time hot text to obtain sentiment features of the hot text; and performing feature fusion on the literature text emotion features and the hot text emotion features to obtain comprehensive emotion features. And then obtaining a target music element from a preset music material according to the comprehensive emotion feature and the classical literature text, and generating a music description text according to the target music element. And then acquiring a scene prompt instruction and generating target music according to the music description text and the scene prompt instruction. According to the embodiment of the invention, high-quality music meeting user requirements can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a music generation method, a music generation device, an electronic device, and a storage medium. Background Art

[0002] Related technologies typically rely on open-source datasets to train music generation models. However, these datasets often contain limited music samples and often contain low-quality music samples. This results in the music generated by these models being too monotonous in style and lacking in complexity in terms of melody, rhythm, and harmony, making it difficult to meet users' diverse needs for various music genres. Therefore, improving the quality of music generation has become an urgent issue. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose a music generation method, a music generation device, an electronic device and a storage medium, aiming to generate high-quality music that meets user needs.

[0004] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application provides a music generation method, the method comprising:

[0005] Obtain classical literature texts and real-time hot texts;

[0006] Extracting sentiment features from the classical literature text to obtain sentiment features of the literature text;

[0007] Extracting emotional features from the real-time hot text to obtain emotional features of the hot text;

[0008] Performing feature fusion on the emotional features of the literary text and the emotional features of the hot text to obtain a comprehensive emotional feature;

[0009] Acquire target music elements from preset music materials according to the comprehensive emotional characteristics and the classical literature text, and generate music description text according to the target music elements;

[0010] Get scene prompt instructions;

[0011] Target music is generated according to the music description text and the scene prompt instruction.

[0012] In some embodiments, the comprehensive emotional feature includes multiple emotional categories and the emotional intensity of each emotional category, the target music element includes a target rhythm sub-element, a target scale sub-element, and a target instrument sub-element, and obtaining the target music element from a preset music material based on the comprehensive emotional feature and the classical literature text, and generating a music description text based on the target music element includes:

[0013] Selecting a target emotion category from the plurality of emotion categories according to each emotion intensity;

[0014] Querying the preset music material according to the target emotion category to obtain the target rhythm sub-element, target scale sub-element and candidate instrument sub-element;

[0015] Get the number of candidate instrument sub-elements to get the number of instruments;

[0016] Screening the candidate instrument sub-elements according to the number of instruments and the classical literature text to obtain a target instrument sub-element;

[0017] The music description text is generated according to the target rhythm sub-element, the target scale sub-element and the target instrument sub-element.

[0018] In some embodiments, screening the candidate instrument sub-elements according to the number of instruments and the classical literature text to obtain the target instrument sub-element includes:

[0019] If the number of musical instruments is greater than one, performing artistic conception feature extraction on the classical literature text to obtain text artistic conception features;

[0020] The candidate instrument sub-elements are screened according to the text artistic conception feature to obtain the target instrument sub-element.

[0021] In some embodiments, obtaining the classical literature text and the real-time hot text includes:

[0022] Obtain original classical texts and real-time hot texts;

[0023] Perform keyword extraction on the original classical text to obtain a first keyword;

[0024] Perform keyword extraction on the real-time hotspot text to obtain a second keyword;

[0025] Calculating the semantic similarity between the first keyword and the second keyword;

[0026] The classical literature text is obtained according to the semantic similarity.

[0027] In some embodiments, calculating the semantic similarity between the first keyword and the second keyword includes:

[0028] Calculating the similarity between the first keyword and the second keyword to obtain a first similarity;

[0029] If the first similarity is greater than or equal to a preset similarity threshold, determining the first similarity as the semantic similarity;

[0030] If the first similarity is less than the preset similarity threshold, performing attention transformation on the first keyword to obtain a first attention feature;

[0031] Performing attention transformation on the second keyword to obtain a second attention feature;

[0032] Calculating a similarity between the first attention feature and the second attention feature to obtain a second similarity;

[0033] The semantic similarity is determined according to the first similarity and the second similarity.

[0034] In some embodiments, generating target music according to the music description text and the scene prompt instruction includes:

[0035] Performing semantic analysis on the scene prompt instruction to obtain scene semantic text;

[0036] updating the music description text according to the scene semantic text to obtain a supplementary description text;

[0037] The target music is generated according to the supplementary description text.

[0038] In some embodiments, the feature fusion of the literary text sentiment feature and the hot text sentiment feature to obtain a comprehensive sentiment feature includes:

[0039] Performing weighted calculation on the emotional features of the literary text and the emotional features of the hot text to obtain an initial emotional feature;

[0040] The initial emotional features are normalized to obtain the comprehensive emotional features.

[0041] To achieve the above-mentioned purpose, a second aspect of the embodiments of the present application provides a music generation device, comprising:

[0042] The first acquisition module is used to acquire classical literature texts and real-time hot texts;

[0043] A first emotion feature extraction module is used to extract emotion features from the classical literature text to obtain emotion features of the literature text;

[0044] A second emotion feature extraction module is used to extract emotion features from the real-time hot text to obtain emotion features of the hot text;

[0045] A feature fusion module is used to fuse the emotional features of the literary text and the emotional features of the hot text to obtain a comprehensive emotional feature;

[0046] a text generation module, configured to obtain target music elements from preset music materials based on the comprehensive emotional characteristics and the classical literature text, and generate a music description text based on the target music elements;

[0047] The second acquisition module is used to obtain scene prompt instructions;

[0048] The music generation module is used to generate target music according to the music description text and the scene prompt instruction.

[0049] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0050] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0051] The music generation method, music generation device, electronic device, and storage medium proposed in this application obtain classical literary text and real-time hot text, extract emotional features from each, and then fuse the emotional features of the classical literary text and the real-time hot text to obtain a comprehensive emotional feature. Compared to traditional music generation models that rely on limited open source data sets for training, the method of this embodiment utilizes the emotional features of classical literary texts to enrich the emotional features and cultural heritage of real-time hot texts. During the generation process, the method of this embodiment extracts target musical elements from the musical material based on the comprehensive emotional features and generates music description text based on the target musical elements. This ensures a deep match between the generated music and the text's artistic conception, further enhancing the artistry and diversity of the target music. Furthermore, by combining specific scene prompts and music description text for music generation, the target music can be highly aligned with the user's needs. Therefore, the method of this embodiment can better meet users' demand for diverse musical styles and overcome the shortcomings of existing technologies in terms of musical complexity and emotional richness, thereby improving the quality of music generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of the music generation method provided in an embodiment of the present application;

[0053] Figure 2 yes Figure 1 Flowchart of step S101 in FIG.

[0054] Figure 3 yes Figure 2 Flowchart of step S204 in FIG.

[0055] Figure 4 yes Figure 1 Flowchart of step S104 in FIG.

[0056] Figure 5 yes Figure 1 Flowchart of step S105 in FIG.

[0057] Figure 6 is a structural diagram of a music generating device provided in an embodiment of the present application;

[0058] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0062] First, let’s analyze some of the terms used in this application:

[0063] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0064] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages ​​(such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.

[0065] In today's digital age and information explosion, real-time hot topics, spread through social media, news platforms, and other information dissemination media, can quickly spark widespread discussion and strong emotional resonance. While people are paying attention to these hot topics, they also crave a corresponding immersive experience. Music, as a powerful tool for expressing emotions and creating atmosphere, can greatly enhance people's feelings about these hot topics.

[0066] Related technologies typically rely on open-source datasets to train music generation models. However, these datasets often contain limited music samples and often contain low-quality music samples. This results in the music generated by these models being too monotonous in style and lacking in complexity in terms of melody, rhythm, and harmony, making it difficult to meet users' diverse needs for various music genres. Therefore, improving the quality of music generation has become an urgent issue.

[0067] Based on this, the embodiments of the present application provide a music generation method, a music generation device, an electronic device and a storage medium, which aim to generate high-quality music that meets user needs.

[0068] The music generation method and device, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the music generation method in the embodiments of the present application is described.

[0069] Figure 1 This is an optional flowchart of the music generation method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S107.

[0070] Step S101: Acquire classical literature texts and real-time hot texts.

[0071] Step S102: extracting sentiment features from the classical literature text to obtain sentiment features of the literature text.

[0072] Step S103: extract emotional features from the real-time hot text to obtain emotional features of the hot text.

[0073] Step S104: Fusing the emotional features of the literary text and the emotional features of the hot text to obtain a comprehensive emotional feature.

[0074] Step S105 , obtaining target music elements from preset music materials according to the comprehensive emotional characteristics and classical literature texts, and generating music description text according to the target music elements.

[0075] Step S106: Obtain scene prompt instructions.

[0076] Step S107: Generate target music according to the music description text and the scene prompt instruction.

[0077] In the steps S101 to S107 shown in the embodiment of the present application, by obtaining classical literature texts and real-time hot texts, and extracting emotional features respectively, the emotional features of the classical literature texts and the real-time hot texts are then integrated to obtain comprehensive emotional features. Compared with the traditional music generation model that relies on limited open source data sets for training, the method of this embodiment utilizes the emotional features of the literary texts of classical literature texts to enrich the emotional features and cultural heritage of the real-time hot texts. During the generation process, the method of this embodiment obtains target music elements from the music material based on the comprehensive emotional features, and generates a music description text based on the target music elements, which can ensure that the music generation is deeply matched with the textual artistic conception, further enhancing the artistry and diversity of the target music. In addition, music generation is performed by combining specific scene prompts and music description texts, so that the target music can be highly consistent with the needs of the user. Therefore, the method of this embodiment can better meet the user's demand for diverse music styles, and overcome the shortcomings of the existing technology in music complexity and emotional richness, thereby improving the quality of music generation.

[0078] In step S101 of some embodiments, the resulting music must match each real-time hot topic text. Current hot topic text refers to events that are currently attracting widespread social attention and discussion, as well as topic-related textual content. In this embodiment, the real-time hot topic text is the text object to which the resulting music must match. Classical literature texts refer to literary works covering different historical periods and genres that share similar or identical themes with the real-time hot topic texts. A pre-set classical literature library can contain a large number of literary works from different historical periods and genres. It is understandable that when composing music for real-time hot topic texts, it is necessary to ensure that the music emotionally matches the real-time hot topic text. Therefore, when subsequently generating music based on classical literature texts and real-time hot topic literature, the themes and emotional characteristics of the selected classical literature texts must also be highly consistent with the hot topic texts. For example, for hot topics related to holiday celebrations, it would be unreasonable to select a classical literature text on a war-related theme. This ensures that the music not only reflects the emotional atmosphere of the current hot topic content, but also leverages the emotional depth and artistic value of classical literature, ultimately enhancing the music's expressiveness and appeal.

[0079] For a selection of classical literary texts, see Figure 2 In some embodiments, step S101 may include but is not limited to steps S201 to S205:

[0080] Step S201: Acquire original classical text and real-time hot text.

[0081] Step S202: extract keywords from the original classical text to obtain the first keyword.

[0082] Step S203: extract keywords from the real-time hotspot text to obtain second keywords.

[0083] Step S204: Calculate the semantic similarity between the first keyword and the second keyword.

[0084] Step S205: Acquire classical literature texts based on semantic similarity.

[0085] In step S201 of some embodiments, real-time hot texts can be captured based on real-time page views on news platforms and social media. Original classical texts are literary texts covering different historical periods, genres, themes, and expressing different emotions.

[0086] In step S202 of some embodiments, preprocess the original classical text. It can be understood that the original classical text is not a modern language and may also be a non-locally set language. The locally set language can be set by the user. For example, if the user sets the locally set language to Chinese, then other languages (such as English, German, and Japanese) are non-locally set languages. First, translate the original classical text into a text with the locally set language and an expression form that conforms to the contemporary expression. Use the word segmentation tool in natural language processing technology to segment the translated text into individual words. Then, remove the stop words. Stop words refer to words that frequently appear in the text but contribute little to the core meaning of the text. For example, when the original classical text is classical Chinese, the corresponding stop words can be "zhi", "hu", "zhe", and "ye", etc. A predefined stop word list can be used to filter out the stop words in the text. Finally, use a keyword extraction algorithm to extract the first keyword. The first keyword refers to the keyword of the text after translating the original classical text, reflecting important features such as the theme, emotion, and imagery of the original classical text. The keyword extraction algorithm can be the Term Frequency-Inverse Document Frequency algorithm (TF-IDF), the graph-based keyword extraction algorithm (TextRank), and the deep learning-based keyword extraction algorithm (BERT-based Embedding), etc., without limitation.

[0087] It can be understood that if the original classical text is a text material in the preset classical literature material library, the mapping relationship of its first keyword can be preset for each original classical text at the beginning of constructing the material library. When the real-time hot text is updated, the first keyword can be directly called.

[0088] In step S203 of some embodiments, the second keyword refers to the keyword of the real-time hot text, reflecting the main topics, key figures, emotions, and important plots of the hot event. The preprocessing operation of the real-time hot text data is almost the same as the preprocessing of the original classical text in step S202 except for the translation step, and will not be elaborated here.

[0089] In step S204 of some embodiments, please refer to Figure 3 , step S204 may include but is not limited to steps S301 to S306:

[0090] Step S301, calculate the similarity between the first keyword and the second keyword to obtain the first similarity.

[0091] Step S302: If the first similarity is greater than or equal to a preset similarity threshold, the first similarity is determined as semantic similarity.

[0092] Step S303: If the first similarity is less than a preset similarity threshold, perform attention transformation on the first keyword to obtain a first attention feature.

[0093] Step S304: Perform attention transformation on the second keyword to obtain a second attention feature.

[0094] Step S305: Calculate the similarity between the first attention feature and the second attention feature to obtain a second similarity.

[0095] Step S306: Determine semantic similarity based on the first similarity and the second similarity.

[0096] In step S301 of some embodiments, the first similarity refers to an indicator that measures the preliminary similarity between the first keyword and the second keyword at the semantic level. A word vector-based method can be used to calculate the first similarity. First, a word vector model that has been pre-trained based on a large-scale corpus is used to convert the first keyword and the second keyword into corresponding vector representations. Then, the cosine similarity between the two vectors is calculated. Cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them. The closer its value is to 1, the more similar the semantics of the two keywords are. If its value is closer to -1, the greater the semantic difference is.

[0097] In step S302 of some embodiments, a preset similarity threshold is used to determine whether the semantic similarity between the first keyword and the second keyword is sufficient. If the cosine similarity is calculated, the preset similarity threshold may be 0.5. It is understandable that in some classical literary works, themes and emotions are directly expressed, such as "sadness", "joy" and "home and country", so the first keyword may have clearly expressed the theme and emotion of the corresponding original classical text in the literal sense. At this time, directly calculating the cosine similarity between the first keyword and the second keyword can determine whether the original classical text and the real-time hot text match at the literal level. If the first similarity is greater than the preset similarity threshold, the first similarity can be directly used as the semantic similarity.

[0098] In step S303 of some embodiments, when the first similarity is less than the preset similarity threshold, it does not mean that the first keyword and the second keyword are unrelated. It is understandable that in some classical literary works, metaphorical writing techniques are often used, and some texts often have deeper meanings. For example, the second keyword of the real-time hot text includes "reunion", while the first keyword of the original classical text may be "full moon". The full moon literally has a completely different meaning from reunion, but the full moon is an image commonly used to represent family reunions. Therefore, even if it is determined that the first similarity is low, it is necessary to explore the deep semantics of the first keyword and the second keyword.

[0099] In this embodiment, attention transformation can be performed using the attention mechanism based on the Transformer architecture. First, the word vector sequence of the first keyword is input into a multi-head attention layer, and multiple different attention heads are used for parallel calculation, and each attention head has its own independent weight matrix. Then, the outputs of each attention head are spliced ​​and linearly transformed to obtain the first attention feature. The obtained first attention feature can better capture the important semantic information of the first keyword and provide a more accurate feature representation for subsequent similarity calculations.

[0100] In step S304 of some embodiments, the process of performing attention transformation on the second keyword is consistent with the principle of the attention transformation process on the first keyword in step S303, and will not be described in detail here.

[0101] In step S305 of some embodiments, the second similarity may be cosine similarity, Euclidean distance, Jaccard similarity, etc. The embodiment of the present application does not strictly limit the specific calculation method of the first similarity and the second similarity.

[0102] In step S306 of some embodiments, when the calculation dimensions and value ranges of the first and second similarities are the same, the first and second similarities can be directly added together to obtain the semantic similarity. The higher the value, the more closely the first and second keywords are aligned in terms of subject matter. Alternatively, weights can be pre-set for the two similarity data, and the first and second similarities can be weighted together to obtain the semantic similarity.

[0103] Steps S301 to S306 shown in the embodiment of the present application significantly improve the accuracy and reliability of keyword semantic similarity calculation through a phased and hierarchical similarity calculation method. The preliminary calculation of the first similarity between the first keyword and the second keyword can quickly screen out situations with high semantic similarity at the literal level, thereby improving processing efficiency. When the first similarity meets the preset threshold, the semantic similarity is directly determined, avoiding unnecessary subsequent complex calculations. When the first similarity does not meet the threshold, the attention transformation is used to mine deeper semantic information of the keyword and calculate the second similarity. This deep mining can capture semantic associations that are difficult to discover with traditional methods. It can effectively improve the matching degree between classical literature texts and real-time hot texts in subsequent music generation.

[0104] In some embodiments, in step S205, a semantic similarity threshold may be set, which may be 0.7, but is not limited thereto. When the semantic similarity score exceeds the semantic similarity threshold, the original classical text containing the corresponding first keyword is considered to have a high correlation with the real-time hot text. All original classical texts that meet the criteria are screened and identified as classical literature texts.

[0105] Steps S201 to S205 shown in the embodiment of this application, by extracting keywords from the original classical text and the real-time hot text and calculating semantic similarity, accurately establish the connection between classical culture and current hot topics, so that traditional classical literature can fit the hot text of the times, enrich the cultural connotation of hot events, and inject them with a deep historical heritage. Secondly, in the field of music generation, it breaks the limitations of traditional reliance on limited open source data sets and introduces a variety of classical literary materials closely related to hot topics, greatly expanding the sources of inspiration for music creation.

[0106] In step S102 of some embodiments, an open source sentiment feature analysis model is used to extract sentiment features from classical literary texts. The sentiment features of literary texts are the sentiment features of classical literary texts. Before the sentiment features are extracted, the classical literary texts can be translated into a user-specified language, and the expression method conforms to modern expression habits. In some embodiments, the translated text can also be supplemented with relevant interpretation texts, and the supplemented text can be cleaned, for example, to remove unnecessary punctuation, annotations, and noise data. The supplemented text is input into the sentiment feature analysis model, and then the sentiment features of the literary text are output.

[0107] In this embodiment, the emotional feature of a literary text is a vector that includes emotional intensity data of multiple emotional categories, including happiness, sadness, anger, fear, surprise, disgust, trust, and expectation. Each emotional intensity data corresponds to an emotional category, indicating the relative intensity of the corresponding emotion in the classical literary text. The value is [0, 1], where 0 indicates that the corresponding emotion does not exist and 1 indicates that the corresponding emotion is the strongest. For example, the emotional feature of a literary text can be specifically expressed as Among them, E cj Represents the emotional characteristics of literary texts, The emotional intensity data representing the happy emotional category, The emotional intensity data representing the sad emotional category, The emotional intensity data representing the expected emotional category.

[0108] In step S103 of some embodiments, an open-source sentiment feature analysis model is used to extract sentiment features from the real-time hot text. The sentiment features of the hot text are the emotional features of the real-time hot text. The real-time hot text can be pre-cleaned to remove colloquial expressions, internet slang, and other noise. The cleaned text is then input into the sentiment feature analysis model to obtain the sentiment features of the hot text.

[0109] Consistent with the emotional features of literary texts, the emotional features of hot texts also include emotional intensity data of multiple emotional categories. The vector dimensions of the emotional features of hot texts are the same as those of the emotional features of literary texts, as well as the value range of the emotional intensity data. Ensuring the consistency of the dimensions of the two emotional feature vectors can facilitate subsequent feature fusion.

[0110] For example, the sentiment feature of hot text can be specifically expressed as Among them, E ti Represents the emotional characteristics of literary texts, The emotional intensity data representing the happy emotional category, The emotional intensity data representing the sad emotional category, The emotional intensity data representing the expected emotional category.

[0111] In step S104 of some embodiments, the comprehensive emotional feature is a comprehensive feature representation obtained by fusing the emotional features of literary texts extracted from classical literary texts with the emotional features of hot texts extracted from real-time hot texts.

[0112] Specifically, see Figure 4 Step S104 may include but is not limited to steps S401 to S402:

[0113] Step S401 : performing weighted calculation on the emotional features of literary texts and the emotional features of hot texts to obtain initial emotional features.

[0114] Step S402: normalize the initial emotion features to obtain comprehensive emotion features.

[0115] In step S401 of some embodiments, weights are determined for the two vectors of literary text emotional features and hot text emotional features respectively. For example, it is assumed that the weight corresponding to the literary text emotional features is α, and the weight of the hot text emotional features is (1-α). The determination of the weights can be adjusted according to the specific application scenarios and needs. For example, if more attention is paid to the impact of the emotional background of classical literature on the final result, α can be set to a larger value. If more attention is paid to the current emotional influence of real-time hot spots, the value of α can be reduced. The initial emotional feature is a fusion feature representation after the weighted fusion of the literary text emotional features and the hot text emotional features, and its calculation process satisfies the following analytical formula:

[0116] E fusion =αE cj +(1-α)E ti (1),

[0117] Among them, E fusion represents the initial emotional feature, E cj Indicates the emotional characteristics of literary texts, E ti represents the sentiment feature of the hot text, α represents the weight coefficient of the sentiment feature of the literary text, and (1-α) represents the weight coefficient of the sentiment feature of the hot text.

[0118] In step S402 of some embodiments, in order to ensure that the values ​​of each dimension of the fused emotion vector are between [0, 1], the initial emotion feature needs to be normalized to obtain the final comprehensive emotion feature. Specifically, the comprehensive emotion feature satisfies the following analytical formula:

[0119]

[0120] Among them, E fusion ' represents the comprehensive emotional features, max{E fusion} represents the initial emotional feature E fusion The emotional intensity data with the largest value.

[0121] Steps S401 to S402 shown in the embodiment of the present application effectively integrate the emotional characteristics of literary texts and hot texts through weighted calculation and normalization processing, which brings significant advantages to music generation. Through weighted calculation, the generated initial emotional characteristics not only retain the profound emotional background of classical literature, but also highlight the current emotional influence of real-time hot spots, and accurately capture the emotional information of texts from different sources. The subsequent normalization processing eliminates the dimensional differences between different features, so that the comprehensive emotional characteristics can be compared and processed under a unified scale. This standardized operation allows the subsequent music generation system to interpret emotional information more accurately and accurately generate appropriate music elements based on emotional characteristics.

[0122] In step S105 of some embodiments, the comprehensive emotional feature includes multiple emotional categories and the emotional intensity of each emotional category, and the target music element includes a target rhythm sub-element, a target scale sub-element and a target instrument sub-element. The target music element is a music attribute element selected from a preset music material library and used to generate music. The preset music material can be stored in the form of a database, which stores the mapping relationship between different emotional categories and various types of music materials. The categories of music material can include rhythm, scale and musical instrument, and each element corresponding to the rhythm category is a rhythm sub-element, each element corresponding to the scale category is a scale sub-element, and each element corresponding to the instrument category is an instrument sub-element. Exemplarily, the mapping relationship between the preset music material and the emotional category is shown in Table 1.

[0123]

[0124] Table 1

[0125] See also Figure 5 In some embodiments, step S105 may also include but is not limited to steps S501 to S505:

[0126] Step S501 : selecting a target emotion category from a plurality of emotion categories according to each emotion intensity.

[0127] Step S502: query the preset music material according to the target emotion category to obtain the target rhythm sub-element, the target scale sub-element and the candidate instrument sub-element.

[0128] Step S503: Obtain the number of candidate instrument sub-elements to obtain the number of instruments.

[0129] Step S504 , screening candidate instrument sub-elements according to the number of instruments and classical literature texts to obtain target instrument sub-elements.

[0130] Step S505 : Generate a music description text according to the target rhythm sub-element, the target scale sub-element, and the target instrument sub-element.

[0131] In step S501 of some embodiments, the target emotion category refers to the emotion category that is dominant or has a significant influence in the current comprehensive emotion feature. Specifically, an intensity threshold is set, such as 0.5 (which can be adjusted according to actual conditions). Traverse all emotion categories and their corresponding emotion intensities, and select emotion categories with emotion intensities greater than or equal to the threshold. Or sort according to the size of the intensity value, and select the first few emotion categories with larger intensity values ​​as the target emotion categories. Since multiple emotions can be expressed simultaneously in the same text, the embodiment of the present application does not strictly limit the number of target emotion categories.

[0132] For example, the comprehensive emotional feature E is known fusion ' = [0.6, 0.3, 0.1, 0.0, 0.2, 0.0, 0.5, 0.4], where the vector structure is the same as that in step S103 regarding the emotional features of literary texts and the emotional features of hot texts. The preset strength threshold is 0.5, and it is further determined that the emotion types corresponding to 0.6 and 0.5 are happiness and trust, respectively. Therefore, happiness and trust are the target emotion categories.

[0133] In step S502 of some embodiments, taking Table 1 as an example, based on happiness and trust, it can be determined that the target rhythm sub-elements are light and stable, the target scale sub-elements are major and harmonious scales, and the candidate instrument sub-elements are piano, guitar, flute, harp and string ensemble.

[0134] In step S503 of some embodiments, according to the example of step S502, the candidate instrument sub-elements are piano, guitar, flute, harp and string ensemble, and the number of instruments is five. It is understandable that if the instrument sub-elements are selected directly by the emotion category, this method may cause the selection of instruments to be too broad. For example, when expressing happy emotions, you can choose a variety of instruments with different timbres such as piano, flute, suona, etc., but instruments with different timbres can express different styles and artistic conceptions. Taking regionality as an example, if the real-time hot text and classical literature text have regional characteristics, then in the selection of instruments, priority is given to the traditional instruments of the region, so that the generated music will have more regional cultural characteristics and emotional resonance. For another example, if there are descriptions of the artistic conception of "mountains and rivers" and "oceans" in the text, you can choose the instrument timbre with a brighter timbre to generate music.

[0135] Therefore, further screening of the initially selected instruments is required.

[0136] In step S504 of some embodiments, if the number of musical instruments is greater than one, artistic conception features are extracted from the classical literary text based on a preset artistic conception feature analysis model to obtain text artistic conception features. It should be noted that artistic conception refers to an artistic realm formed by the interplay of the scenes and atmosphere depicted in a literary work with the emotions expressed by the author. A pre-trained neural network model can be used to analyze the imagery, emotional tone, and language style of the classical literary text to extract features that represent the artistic conception, i.e., the text artistic conception features.

[0137] Then, the candidate instrument sub-elements are screened according to the text's artistic conception features to obtain the target instrument sub-elements. An instrument-artistic conception matching rule table can be established, which records the type of artistic conception that each instrument is suitable for expressing.

[0138] For example, continuing with the example of step S503, if the classical literature text includes the words "Liushui Renjia" and "Jiangnan", the method of this embodiment can select the flute as the target instrument sub-element from the candidate instrument sub-elements (piano, guitar, flute, harp, and string ensemble). It should be noted that there can be multiple target instrument sub-elements, and this embodiment does not strictly limit the number of target instrument sub-elements.

[0139] In other embodiments, if the number of instruments is equal to one, the candidate instrument sub-element is directly used as the target instrument sub-element.

[0140] In step S505 of some embodiments, a music description text is generated according to a certain language template based on the characteristics of the target rhythm sub-element, the target scale sub-element and the target instrument sub-element. For example, the language template may be "use [target instrument sub-element] to play with the rhythm of [target rhythm sub-element], adopt [target scale sub-element] to create the atmosphere of [target emotion category]". Substitute the specific target rhythm sub-element, target scale sub-element and target instrument sub-element into this template, and input the text expression substituted into the template into the large language model for polishing, and the corresponding music description text can be generated. For example, if the target rhythm sub-element is "light and stable", the target scale sub-element is "harmonious major", the target instrument sub-element is "flute", and the target emotion category is "joy", then the generated music description text may be "use the flute to play with a light and stable rhythm, adopt a major key with higher harmony, create an atmosphere of joy and expectation for the future, and have a strong traditional flavor."

[0141] Steps S501 to S505 shown in the embodiment of the present application accurately select the target emotion category to ensure that the generated music can accurately convey the required emotion. Then, based on the target emotion category, the preset music material is queried to obtain the target rhythm sub-element, the target scale sub-element and the candidate instrument sub-element. Subsequently, the candidate instrument sub-element is screened based on the number of instruments and the classical literature text to obtain the target instrument sub-element, so that the selection of instruments takes into account both emotional expression and the artistic conception of classical literature, thereby enhancing the cultural connotation and artistic appeal of the music. Finally, a music description text is generated based on these elements, which provides a specific and accurate creative guidance for the music generation model, so that the generated music can achieve an organic unity in emotional expression, style, artistic conception and cultural heritage, and better meet the user's demand for diversified and personalized music.

[0142] In step S106 of some embodiments, the scene prompt instruction is scene feature information obtained through user input, sensor data, or contextual information, and can include application scenarios such as family gatherings, relaxation meditation, television programs, and advertising marketing. The scene prompt instruction can be obtained by the user manually entering the scene on the interface, or by the smart device's sensors collecting environmental data and analyzing the scene information, but is not limited to this.

[0143] In step S107 of some embodiments, the scene prompt instruction is first semantically parsed to obtain a scene semantic text, and then the music description text is updated according to the scene semantic text to obtain a supplementary description text. Specifically, since the scene prompt instruction may be a voice, image or text input, these inputs are semantically parsed, the user's current application scenario is analyzed and expressed in text form, and then the scene semantic text is obtained. The scene semantic text is used to supplement the music description text generated in the previous step. For example, the scene semantic text includes the text of "family gathering". The supplementary description text can be "using a flute with a light and steady rhythm, using a relatively high harmony major key to play, creating an atmosphere of joy and anticipation for the future, suitable for family gatherings, and with a strong traditional flavor."

[0144] Finally, the target music is generated based on the supplementary description text. An appropriate music generation model is selected, such as a deep learning-based generative adversarial network (GAN) or variational autoencoder (VAE). The supplementary description text is converted into an input format acceptable to the model, for example, by encoding information such as rhythm, scale, and instrumentation in the text into a vector form. The encoded input data is input into the music generation model, which processes and generates the input based on its own learning mechanism and parameters. During the generation process, the model combines the various requirements in the supplementary description text to generate a music sequence that meets the characteristics of rhythm, scale, instrument playing style, and scene atmosphere. Finally, the generated music sequence is post-processed, such as adjusting the volume balance and adding appropriate sound effects, to make it a complete, high-quality musical work, namely the target music. The generated target music can be saved in common music file formats (such as MP3 and WAV) for user use.

[0145] See also Figure 6 The present application also provides a music generation device that can implement the above-mentioned music generation method. The device includes:

[0146] The first acquisition module 601 is used to acquire classical literature texts and real-time hot texts;

[0147] The first emotion feature extraction module 602 is used to extract emotion features from classical literature texts to obtain emotion features of the literature texts;

[0148] The second emotion feature extraction module 603 is used to extract emotion features from the real-time hot text to obtain emotion features of the hot text;

[0149] The feature fusion module 604 is used to fuse the emotional features of the literary text and the emotional features of the hot text to obtain a comprehensive emotional feature;

[0150] The text generation module 605 is used to obtain target music elements from the preset music materials according to the comprehensive emotional characteristics and classical literature texts, and generate music description text according to the target music elements;

[0151] The second acquisition module 606 is used to acquire a scene prompt instruction;

[0152] The music generation module 607 is used to generate target music according to the music description text and scene prompt instructions.

[0153] The specific implementation of the music generation device is basically the same as the specific embodiment of the above-mentioned music generation method, and will not be repeated here.

[0154] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned music generation method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0155] See also Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0156] The processor 701 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0157] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called by the processor 701 to execute the music generation method of the embodiments of this application.

[0158] Input / output interface 703, used to implement information input and output;

[0159] Communication interface 704, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0160] Bus 705 , which transmits information between various components of the device (e.g., processor 701 , memory 702 , input / output interface 703 , and communication interface 704 );

[0161] The processor 701 , the memory 702 , the input / output interface 703 and the communication interface 704 are connected to each other in communication within the device via a bus 705 .

[0162] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned music generation method is implemented.

[0163] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0164] The music generation method, music generation device, electronic device and storage medium provided in the embodiment of the present application obtain classical literature text and real-time hot text, and extract emotional features respectively, and then fuse the emotional features of classical literature text and real-time hot text to obtain comprehensive emotional features. Compared with the traditional music generation model that relies on limited open source data sets for training, the method of this embodiment utilizes the emotional features of literary texts of classical literature texts to enrich the emotional features and cultural heritage of real-time hot texts. During the generation process, the method of this embodiment obtains target music elements from music materials based on the comprehensive emotional features, and generates music description text based on the target music elements, which can ensure the deep matching of music generation and textual conception, and further enhance the artistry and diversity of the target music. In addition, music generation is performed by combining specific scene prompts and music description texts, so that the target music can be highly consistent with the needs of users. Therefore, the method of this embodiment can better meet the user's demand for diverse music styles, and overcome the shortcomings of the existing technology in music complexity and emotional richness, thereby improving the quality of music generation.

[0165] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0166] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0168] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0169] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0170] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0171] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0172] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0173] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0175] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A music generation method, characterized in that: The method comprises: Obtain classical literature texts and real-time hot texts; Extracting sentiment features from the classical literature text to obtain sentiment features of the literature text; Extracting emotional features from the real-time hot text to obtain emotional features of the hot text; Performing feature fusion on the emotional features of the literary text and the emotional features of the hot text to obtain a comprehensive emotional feature; Acquire target music elements from preset music materials according to the comprehensive emotional characteristics and the classical literature text, and generate music description text according to the target music elements; Get scene prompt instructions; Target music is generated according to the music description text and the scene prompt instruction.

2. The method according to claim 1, characterized in that The comprehensive emotional feature includes multiple emotional categories and the emotional intensity of each emotional category; the target music element includes a target rhythm sub-element, a target scale sub-element, and a target instrument sub-element; obtaining the target music element from a preset music material based on the comprehensive emotional feature and the classical literature text, and generating a music description text based on the target music element include: Selecting a target emotion category from the plurality of emotion categories according to each emotion intensity; Querying the preset music material according to the target emotion category to obtain the target rhythm sub-element, target scale sub-element and candidate instrument sub-element; Get the number of candidate instrument sub-elements to get the number of instruments; Screening the candidate instrument sub-elements according to the number of instruments and the classical literature text to obtain a target instrument sub-element; The music description text is generated according to the target rhythm sub-element, the target scale sub-element and the target instrument sub-element.

3. The method according to claim 2, characterized in that The step of screening the candidate instrument sub-elements according to the number of instruments and the classical literature text to obtain the target instrument sub-element includes: If the number of musical instruments is greater than one, performing artistic conception feature extraction on the classical literature text to obtain text artistic conception features; The candidate instrument sub-elements are screened according to the text artistic conception feature to obtain the target instrument sub-element.

4. The method according to claim 1, wherein The acquisition of classical literature texts and real-time hot texts includes: Obtain original classical texts and real-time hot texts; Perform keyword extraction on the original classical text to obtain a first keyword; Perform keyword extraction on the real-time hotspot text to obtain a second keyword; Calculating the semantic similarity between the first keyword and the second keyword; The classical literature text is obtained according to the semantic similarity.

5. The method according to claim 4, characterized in that The calculating the semantic similarity between the first keyword and the second keyword includes: Calculating the similarity between the first keyword and the second keyword to obtain a first similarity; If the first similarity is greater than or equal to a preset similarity threshold, determining the first similarity as the semantic similarity; If the first similarity is less than the preset similarity threshold, performing attention transformation on the first keyword to obtain a first attention feature; Performing attention transformation on the second keyword to obtain a second attention feature; Calculating a similarity between the first attention feature and the second attention feature to obtain a second similarity; The semantic similarity is determined according to the first similarity and the second similarity.

6. The method according to any one of claims 1 to 5, characterized in that Generating target music according to the music description text and the scene prompt instruction includes: Performing semantic analysis on the scene prompt instruction to obtain scene semantic text; updating the music description text according to the scene semantic text to obtain a supplementary description text; The target music is generated according to the supplementary description text.

7. The method according to any one of claims 1 to 5, characterized in that The feature fusion of the literary text sentiment feature and the hot text sentiment feature to obtain a comprehensive sentiment feature includes: Performing weighted calculation on the emotional features of the literary text and the emotional features of the hot text to obtain an initial emotional feature; The initial emotional features are normalized to obtain the comprehensive emotional features.

8. A music generating device, characterized in that: The device comprises: The first acquisition module is used to acquire classical literature texts and real-time hot texts; A first emotion feature extraction module is used to extract emotion features from the classical literature text to obtain emotion features of the literature text; A second emotion feature extraction module is used to extract emotion features from the real-time hot text to obtain emotion features of the hot text; A feature fusion module is used to fuse the emotional features of the literary text and the emotional features of the hot text to obtain a comprehensive emotional feature; a text generation module, configured to obtain target music elements from preset music materials based on the comprehensive emotional characteristics and the classical literature text, and generate a music description text based on the target music elements; The second acquisition module is used to obtain scene prompt instructions; The music generation module is used to generate target music according to the music description text and the scene prompt instruction.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Emotional music generation method based on deep neural network and music element driving

    CN113299255A

  • Method and device for generating multimedia product of music and medium

    CN118427372A

  • Music generation method and device, electronic equipment and storage medium

    CN119400134A

  • Process to provide audio / video / literature files and / or events / activities ,based upon an emoji or icon associated to a personal feeling

    US20180025004A1

Cited By

  • Cross-culture melody feature extraction and musical instrument matching system based on deep learning

    CN121459847A