Data generation method and apparatus for music works, and storage medium and program product
By extracting audio features and analyzing lyrics of musical works, tag data and metadata are generated. Machine learning models are then used for comprehensive analysis, solving the accuracy problem of music resource management and recommendation systems, and achieving more efficient music library management and more accurate recommendations.
Patent Information
- Application Number
- PCT/CN2024/094251
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-11-27
AI Technical Summary
How to effectively manage and accurately classify music resources, especially to generate data on valuable music works, in order to improve user experience and the efficiency of music recommendation systems.
By extracting audio features and analyzing lyrics of musical works, tag data and metadata of musical works are generated. Machine learning models are then used for comprehensive analysis, combining analysis results from multiple perspectives to improve the accuracy and richness of the data.
It improves the accuracy of music recommendation systems and the richness of music products, enhances the accuracy and convenience of music library management, and improves the recommendation accuracy of music recommendation systems.
Smart Images

Figure CN2024094251_27112025_PF_FP_ABST
Abstract
Description
Data generation method and device of music works, storage medium, program product TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular to a data generation method and device of music works, a computer readable storage medium and a computer program product. BACKGROUND
[0002] With the popularity of digital music services, music libraries have rapidly expanded. How to effectively manage these music resources, especially how to accurately classify and generate valuable data of music works, has become the key to improving user experience and the efficiency of music recommendation systems.
[0003] SUMMARY
[0004] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0005] According to a first aspect of some embodiments of the present disclosure, a data generation method of music works is provided, comprising:
[0006] extracting audio features of music in the music works, the audio features reflecting an understanding of the music in the music works;
[0007] performing lyrics analysis on lyrics of the music works to obtain lyrics analysis results, the lyrics analysis results reflecting an understanding of the content of the lyrics;
[0008] generating at least one of tag data and metadata of the music works according to the audio features and the lyrics analysis results.
[0009] According to a second aspect of some embodiments of the present disclosure, a data generation device of music works is provided, comprising:
[0010] an extraction module configured to extract audio features of music in the music works, the audio features reflecting an understanding of the music in the music works;
[0011] a lyrics analysis module configured to perform lyrics analysis on lyrics of the music works to obtain lyrics analysis results, the lyrics analysis results reflecting an understanding of the content of the lyrics;
[0012] a generation module configured to generate at least one of tag data and metadata of the music works according to the audio features and the lyrics analysis results.
[0013] According to a third aspect of some embodiments of the present disclosure, there is provided a data generation apparatus of a musical work, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the data generation method of any of the embodiments described in the present disclosure based on instructions stored in the memory.
[0014] According to a fourth aspect of some embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, performs the data generation method of any of the embodiments described in the present disclosure.
[0015] According to a fifth aspect of some embodiments of the present disclosure, there is provided a computer program product which, when running on a computer, causes the computer to implement the data generation method of any of the embodiments.
[0016] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments of the present disclosure with reference to the following drawings. BRIEF DESCRIPTION OF DRAWINGS
[0017] The preferred embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. The accompanying drawings are used in the present disclosure to provide a further understanding of the present disclosure and are incorporated and constitute a part of the description of the present disclosure. It should be understood that the accompanying drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure. In the drawings:
[0018] FIG. 1 is a flowchart illustrating a data generation method according to some embodiments of the present disclosure;
[0019] FIG. 2A is a flowchart illustrating a data generation method according to some other embodiments of the present disclosure;
[0020] FIG. 2B is a flowchart illustrating a data generation method according to some other embodiments of the present disclosure;
[0021] FIG. 2C is a flowchart illustrating a data generation method according to some other embodiments of the present disclosure;
[0022] FIG. 2D is a flowchart illustrating a data generation method according to some other embodiments of the present disclosure;
[0023] FIG. 2E is a flowchart illustrating a data generation method according to some other embodiments of the present disclosure;
[0024] FIG. 3 is a diagram illustrating prompt information for lyrics analysis according to some embodiments of the present disclosure;
[0025] FIG. 4 is a diagram illustrating prompt information for musical work analysis according to some embodiments of the present disclosure;
[0026] FIG. 5 is a diagram illustrating a result of analysis of a musical piece according to some embodiments of the present disclosure;
[0027] FIG. 6 is a block diagram illustrating a data generation apparatus according to some embodiments of the present disclosure;
[0028] FIG. 7 is a block diagram illustrating a data generation apparatus according to other embodiments of the present disclosure;
[0029] FIG. 8 is a block diagram illustrating an electronic device according to other embodiments of the present disclosure.
[0030] It is to be understood that the dimensions of the various portions shown in the drawings are not necessarily to scale and that the various embodiments can have different relative dimensions than those depicted in the drawings. Like or similar designations in the drawings have been used to indicate like and / or similar functionality utilizing like or similar designations drawn to like or similar characters throughout the several views. DETAILED DESCRIPTION
[0031] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the embodiments below is actually only illustrative, but not as any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein.
[0032] It should be understood that the various steps in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of the components and steps illustrated in these embodiments, numerical expressions, and numerical values should be interpreted as examples only, and not limiting the scope of the present disclosure.
[0033] The term "comprise" and variations of the term, such as "comprising," "includes," "including," and "contains," "containing," as used herein, are open-ended terms that mean "including, but not limited to." In addition, the term "comprise" and variations of the term, such as "comprising," "includes," "including," and "contains," "containing," as used herein, are open-ended terms that mean "including, but not limited to." Thus, "comprising" and "including" are synonymous. The term "based on" means "based, at least in part, on."
[0034] Reference throughout this specification to "one embodiment", "an embodiment", or "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment, although it can. Furthermore, the term "comprising" is used throughout the specification to mean that various described features, structures, or characteristics can be included in, but are not limited to, the present application. Furthermore, the terms "first", "second", and the like are used herein merely to distinguish one element from another, and are not necessarily used to denote a particular order or sequence.
[0035] It should be noted that the terms "first", "second", and the like, used in the description and in the claims of this disclosure are used for distinguishing between similar elements and do not necessarily have an ordinal number meaning by themselves. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the present disclosure. The terms "first", "second", and the like, are used herein merely to differentiate one element from another in a specific context and do not necessarily indicate a temporal or chronological order or sequence.
[0036] It should be noted that the terms "one", "another", "an", and "the", as used in this disclosure, are used in their broadest context to mean "one or more" unless otherwise indicated.
[0037] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the present disclosure are used only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0038] The embodiments of the present disclosure will be described in detail below with reference to the drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics can be combined by any suitable means from the present disclosure that will be clear to those skilled in the art.
[0039] Some technical terms in the art are provided below.
[0040] Large Language Models (LLM): These are advanced language processing systems based on artificial intelligence, trained using deep learning techniques and large-scale datasets. LLMs are capable of understanding and generating natural language text, suitable for various tasks such as text generation, translation, summarization, question answering, etc. LLMs like the GPT (Generative Pre-trained Transformer) series and BERT (Bidirectional Encoder Representations from Transformers) learn language rules and patterns through pre-training and fine-tuning methods, providing high-level language understanding and generation capabilities.
[0041] Agent: In the field of artificial intelligence, an agent refers to a software entity that can perform tasks automatically. These agents can operate in specific environments, respond to environmental changes, and make decisions based on their programmed objectives. They can be used for automating processes, data analysis, real-time monitoring, etc. Agents can be designed based on rules, algorithms, or machine learning techniques, enabling them to perform tasks independently within their working environment.
[0042] LLM-based Agent: Combining the capabilities of large language models with the autonomy of agents, LLM-based agents utilize the language processing capabilities of LLMs to understand and generate natural language, enabling more human-like and efficient interactions with users. They can be used for chatbots, personal assistants, automated customer service, etc. LLM-based agents can understand complex queries, generate fluent responses, and learn and adapt to new contexts and information when needed.
[0043] Time Series Database (TSDB): A database specifically designed to handle time series data, which are data points arranged in chronological order. Time series data is common in various domains, such as financial market data, Internet of Things (IoT) sensor data, application performance monitoring data, etc. TSDBs provide optimized storage, querying, and processing capabilities for this type of data.
[0044] Cooperative Engagement: Refers to the interaction and collaboration between different agents for a common goal or task. This collaboration can include information sharing, decision coordination, task allocation, and solution negotiation, etc.
[0045] Disordered agents do not have a fixed structure or order when collaborating. Their interactions are more free-form and dynamic, not following strict rules or predefined processes. They have strong autonomy, with each agent maintaining high autonomy in collaboration, making decisions based on its own judgment and ability; dynamic adaptability, this type of agent can flexibly adjust its behavior and strategy according to environmental and task requirements; nonlinear interaction, the interaction mode between agents is not fixed and can change at any time according to real-time conditions. Complex problem solving scenarios, suitable for complex problems without fixed solutions, agents can explore possible solutions from multiple perspectives. Innovation and exploration task scenarios, good at tasks that require creative thinking and exploration of new fields.
[0046] Mainly based on unstructured, flexible and dynamic interaction mode: in this setting, agents work independently, but can also interact and collaborate with other agents as needed and according to the scene.
[0047] Ordered agents: these agents work in a structured and ordered environment, usually following predetermined rules and processes, suitable for tasks with clear rules and fixed processes.
[0048] Adversarial interactions: refers to the adversarial interactions between agents, where agents try to optimize their own goals through strategic behavior, which may sometimes conflict with the goals of other agents.
[0049] Adversarial interactions in LLM-based agents involve competitive or adversarial interactions between agents based on large language models (LLM). This type of interaction simulates real-world adversarial environments such as negotiations, debates or strategy games, where agents need to adopt competing strategies to achieve their own goals.
[0050] The present disclosure provides a technical solution that can improve the accuracy of data generated for music works.
[0051] Figure 1 is a flowchart showing a data generation method according to some embodiments of the present disclosure.
[0052] As shown in FIG. 1, the data generation method of the musical work comprises: step S10, extracting audio features of music in the musical work, the audio features reflecting the understanding of the music in the musical work; step S20, performing lyrics analysis on lyrics of the musical work to obtain lyrics analysis results, the lyrics analysis results reflecting the understanding of the content of the lyrics; and step S30, generating at least one of tag data and metadata of the musical work according to the audio features and the lyrics analysis results.
[0053] In the above embodiment, the audio features of music in the musical work and the lyrics analysis results of the lyrics are combined to generate at least one of the tag data and the metadata of the musical work, the musical work can be comprehensively and multi-angle analyzed, thereby improving the accuracy of the generated data of the musical work, and improving the accuracy of the music recommendation system and the richness of the music product. The tag data and the metadata obtained by any embodiment of the present disclosure can be used for music library management and music recommendation system, improve the accuracy and convenience of music library management, and improve the recommendation accuracy of the music recommendation system.
[0054] For example, the audio features include but are not limited to pitch, rhythm, harmony, timbre, dynamics, harmonic features, spectral centroid, and zero crossing rate.
[0055] Pitch refers to the high and low of music sound, and is a basic component of music melody. In audio analysis, pitch extraction can be used to identify melody lines and harmonic content.
[0056] Rhythm is the way of time organization of music, including the time value of notes and the arrangement of strong and weak beats. Rhythm analysis can reveal the speed (BPM, beats per minute) and rhythm pattern of the song.
[0057] Harmony involves the simultaneous sounding of multiple notes, forming chords or harmonic progressions. Harmony analysis can help identify the chord types and harmonic structures used in the song.
[0058] Timbre describes the texture of sound, even if the same pitch and volume of sound, different instruments or sound sources will have different timbres. Timbre analysis can identify the main instruments and sound characteristics in the song.
[0059] Dynamics refers to the range of volume changes, from the lightest to the loudest. Dynamic analysis can reveal the expressive intensity of the song, such as soft passages and intense explosions.
[0060] Harmonic features refer to the harmonic content in musical sound, which determines the richness and complexity of music. Harmonic analysis helps to understand the timbre and texture of a song.
[0061] Spectral centroid is an indicator of the "brightness" of sound, with high centroid values usually implying a brighter sound and low centroid values a darker sound.
[0062] Zero-crossing rate refers to the rate at which the waveform of an audio signal crosses zero, and is often used to identify the texture and rhythmic feel of a sound.
[0063] The data generation method in other embodiments of the present disclosure will be described in detail below in combination with FIGS. 2A-5.
[0064] FIG. 2A is a flow diagram illustrating a data generation method according to some embodiments of the present disclosure. FIG. 2A differs from FIG. 1 in that steps S31-S33 in FIG. 2A are an implementation of step S30 in FIG. 1. Only the differences between FIG. 2A and FIG. 1 will be described below, and the same parts will not be described again.
[0065] As shown in FIG. 2A, in step S31, prompt information for music work analysis is generated, the prompt information for music work analysis including description information of a work analysis angle, the audio features, and the lyrics analysis result, the work analysis angle focusing on the whole music work including the music and the lyrics.
[0066] In step S32, according to the prompt information for music work analysis, a machine learning model is used to perform music work analysis on the music work, obtaining an analysis result of the music work. For example, the prompt information for music work analysis can be input into a machine learning model to obtain an analysis result of the music work.
[0067] In step S33, at least one of label data and metadata of the music work is generated according to the analysis result of the music work.
[0068] In this embodiment, by using the prompt information for music work analysis including description information of a work analysis angle, the audio features, and the lyrics analysis result to guide the machine learning model to perform overall analysis on the music work, the direction of overall analysis of the music work can be clearly defined, and the accuracy of the generated data of the music work can be further improved, thereby improving the accuracy of the music recommendation system and the richness of the music product.
[0069] In some embodiments, there are multiple work analysis angles, and the analysis result of the music work includes analysis results for multiple work analysis angles. According to the analysis result of the music work, generating at least one of the tag data and the metadata of the music work includes: according to the analysis results for the multiple work analysis angles, generating at least one of the following: multiple tag dimensions based on the multiple work analysis angles and tags of the music work corresponding to each tag dimension as the tag data of the music work; and multiple metadata fields based on the multiple work analysis angles and field values corresponding to each metadata field as the metadata of the music work.
[0070] In this embodiment, at least one of the multiple tag dimensions and the multiple metadata fields is generated based on the multiple work analysis angles, so that the data of the music work presents multiple dimensions, improves the richness of the data of the music work, and thus improves the accuracy of the music recommendation system and the richness of the music product, and improves the accuracy of the music recommendation system and the richness of the user experience.
[0071] FIG. 2B is a flow diagram illustrating a data generation method according to some other embodiments of the present disclosure. FIG. 2B differs from FIG. 1 in that steps S21-S22 in FIG. 2B are an implementation of step S20 in FIG. 1. Hereinafter, only the differences between FIG. 2B and FIG. 1 will be described, and the same parts will not be described again.
[0072] As shown in FIG. 2B, in step S21, prompt information for lyrics analysis is generated, the prompt information for lyrics analysis including lyrics and description information of a lyrics analysis angle, the lyrics analysis angle focusing on the lyrics. In step S22, according to the prompt information for lyrics analysis, a machine learning model is used to perform lyrics analysis on the lyrics of the music work to obtain the lyrics analysis result.
[0073] In this embodiment, for the lyrics analysis part, the prompt information for lyrics analysis is used to guide the machine learning model to perform lyrics analysis, which can further improve the accuracy of lyrics analysis, and thus further improve the accuracy of the data of the music work, and thus improve the accuracy of the music recommendation system and the richness of the music product.
[0074] In some embodiments, there are multiple lyrics analysis angles, and according to the prompt information for lyrics analysis, a machine learning model is used to perform lyrics analysis on the lyrics of the music work to obtain the lyrics analysis result, which includes: according to the prompt information for lyrics analysis, a machine learning model is used to perform lyrics analysis on the lyrics of the music work to obtain analysis results for multiple lyrics analysis angles as the lyrics analysis result.
[0075] FIG. 2C is a flow diagram illustrating a data generation method according to some embodiments of the present disclosure. FIG. 2C differs from FIG. 1 in that FIG. 2C illustrates other steps S40-S70 of the data generation method according to some embodiments of the present disclosure. Hereinafter, only the differences between FIG. 2C and FIG. 1 will be described, and the same parts will not be described again.
[0076] As shown in FIG. 2C, in step S40, feedback information of at least one of the tag data and the metadata of the music work is obtained. For example, the feedback information can come from an agent expert, for example, a system that combines an artificial intelligence agent and human expert knowledge.
[0077] In step S50, in response to a difference between the feedback information and at least one of the tag data and the metadata of the music work, it is determined that the difference is associated with the music and the lyrics.
[0078] In step S60, in response to the association between the difference and the music satisfying a preset condition, parameters in a stage of extracting the audio features and a stage of generating at least one of the tag data and the metadata of the music work are adjusted. For example, the parameters in the stage of extracting the audio features can include parameters of a model for extracting the audio features, or can be the dimension of the audio features. For example, the parameters in the stage of generating at least one of the tag data and the metadata of the music work can be parameters of a machine learning model, or can be the generated prompt information related to the music itself. For example, the parameters of the machine learning model are adjusted, and the machine learning model is retrained.
[0079] In step S70, in response to the association between the difference and the lyrics satisfying a preset condition, parameters in a stage of analyzing the lyrics and a stage of generating at least one of the tag data and the metadata of the music work are adjusted. For example, the parameters in the stage of analyzing the lyrics can include parameters of a machine learning model for analyzing the lyrics, or can be the prompt information for analyzing the lyrics itself. For example, the parameters in the stage of generating at least one of the tag data and the metadata of the music work can be parameters of a machine learning model, or can be the generated prompt information related to the lyrics itself.
[0080] In this embodiment, by analyzing the feedback information of the data generation result of the music work, the difference is located, the data generation process of the music work is automatically optimized, the accuracy of the generated data of the music work is further improved, and the precision of the music recommendation system and the richness of the music product are improved.
[0081] In some embodiments, after performing the adjustment operation, the tag data and the metadata of the music work can be updated according to the adjusted method of the present disclosure.
[0082] FIG. 2D is a flow diagram illustrating a data generation method according to some embodiments of the present disclosure. FIG. 2D differs from FIG. 2A in that steps S331-S333 of FIG. 2D are an implementation of step S33 of FIG. 2A. Hereinafter, only the differences between FIG. 2D and FIG. 2A will be described, and the same parts will not be described again.
[0083] As shown in FIG. 2D, in step S331, the analysis result of the music work is subjected to text analysis to obtain an analysis result. In step S332, according to the analysis result, information extraction is performed on the analysis result of the music work using a pre-set matching instruction or template to obtain key information. In step S333, the key information is subjected to structured data conversion to obtain at least one of label data and metadata of the music work.
[0084] For example, the text analysis includes analyzing the semantics, contextual relationships, and key information points of the analysis result of the music work, i.e., content understanding of the analysis result of the music work. In some embodiments, natural language processing techniques such as entity recognition, sentiment analysis, and topic extraction can be used for text analysis. Key information points include, for example, music segmentation, sentiment orientation, and theme content, and the text analysis process can identify these key information points.
[0085] For example, according to the analysis result from the text "fast rhythm and full of energy", the labels "fast rhythm" and "high energy" are extracted. Through data classification, the label dimensions to which different labels belong can be determined.
[0086] For example, the structured data conversion can convert the text content into a structured format such as a database table or a JSON (JavaScript Object Notation) object, facilitating storage and retrieval.
[0087] FIG. 2E is a flow diagram illustrating a data generation method according to some embodiments of the present disclosure. FIG. 2E differs from FIG. 2A in that steps S331'-S332' of FIG. 2E are another implementation of step S33 of FIG. 2A. Hereinafter, only the differences between FIG. 2E and FIG. 2A will be described, and the same parts will not be described again.
[0088] As shown in FIG. 2E, in step S331', prompt information for the label data and / or metadata extraction of the music work is generated, where the prompt information for the label data and / or metadata extraction of the music work includes the analysis result of the music work and description information for the extraction target and presentation format, the extraction target including label dimensions based on the work analysis angle and / or metadata fields, and the presentation format including the format of the correspondence between the label dimensions and labels and / or the correspondence between the metadata fields and field values. In step S332', at least one of the label data and metadata of the music work is generated according to the prompt information for the label data and / or metadata extraction of the music work using a machine learning model.
[0089] For example, the prompt information for the label data and / or metadata extraction of the music work is input into a machine learning model, and at least one of the label data and metadata of the music work is output after processing by the machine learning model. The present disclosure does not limit the specific form of the prompt information for the label data and / or metadata extraction of the music work.
[0090] In some embodiments, generating prompt information for music work analysis includes:
[0091] Inserting the audio features and the lyrics analysis result into a music work analysis template, where the music work analysis template includes description information of the work analysis angle.
[0092] In some embodiments, generating prompt information for lyrics analysis includes:
[0093] Obtaining lyrics content of the music work;
[0094] Inserting the lyrics content into a lyrics analysis template to obtain the prompt information for lyrics analysis, where the lyrics analysis template includes description information of the lyrics analysis angle.
[0095] In some embodiments, the audio features include numerical representations of audio features and natural language description information corresponding to the numerical representations; and / or the audio features include physical features of the audio signal of the music and music features reflecting music properties based on the physical features. The audio features can be extracted from the music file by an audio analysis system using various signal processing and machine learning techniques.
[0096] In some embodiments, the data generation method in any of the above embodiments is performed by an agent. The machine learning model in the above embodiments may, for example, be a large language model.
[0097] FIG. 3 is a schematic diagram illustrating prompt information for lyrics analysis according to some embodiments of the present disclosure.
[0098] As shown in FIG. 3, the prompt information for lyrics analysis includes lyrics and description information of at least one lyrics analysis angle. The description information of at least one lyrics analysis angle includes description information of lyrics analysis angle 1 and description information of lyrics analysis angle 2. Referring to FIG. 3, the prompt information for lyrics analysis also includes some information that helps machine learning model to understand, such as “Please analyze in depth based on the following lyrics, explore its multi-dimensional meanings and artistic qualities”, “When analyzing, please consider the following dimensions carefully” and “Please provide a comprehensive analysis report covering all the above dimensions and revealing any hidden meanings or deep cultural connections”.
[0099] For example, the lyrics analysis angle includes but is not limited to emotional spectrum, theme depth, cultural and historical background, narrative technique, language creativity, emotional atmosphere and style, interaction with musical elements, target audience and resonance, etc.
[0100] The description information of lyrics analysis angle “emotional spectrum” includes, for example, “Describe in detail the range of emotions exhibited in the lyrics, including the primary emotions and their change process. Does the lyrics smoothly transition from one emotion to another? How is this change achieved?”
[0101] The description information of lyrics analysis angle “theme depth” includes, for example, “Identify and explain the central themes in the lyrics, including how they reflect human experiences, social issues or personal stories. Are there secondary themes, and how are they interwoven with the main theme?”
[0102] The description information of lyrics analysis angle “cultural and historical background” includes, for example, “Analyze whether the lyrics reference specific cultures, historical events or figures. How do these references enhance the meaning and depth of the song?”
[0103] The description information of lyrics analysis angle “narrative technique” includes, for example, “Explore the narrative techniques used in the lyrics, such as first-person or third-person narration, and how these choices affect the listener's feelings and understanding.”
[0104] The description information of lyrics analysis angle “language creativity” includes, for example, “Analyze the language innovations in the lyrics, including unique metaphors, symbolism, personification and other rhetorical devices. How do these techniques enrich the symbolic level of the song?”
[0105] The description information of lyrics analysis angle “emotional atmosphere and style” includes, for example, “Discuss how the lyrics create a specific emotional atmosphere and style through vocabulary selection, rhythm and rhythm. How does this atmosphere and style interact with the rest of the music?”
[0106] The description information of the lyrics analysis angle "interaction with musical elements" may include, for example, "evaluate the relationship between the lyrics and elements such as melody, rhythm, and harmony of the music. How do the lyrics and musical elements complement each other to convey the emotions and themes of the song?"
[0107] The description information of the lyrics analysis angle "target audience and resonance" may include, for example, "infer the target audience of this song based on the content and style of the lyrics. How do the lyrics resonate with its audience and touch the emotions of the listeners?"
[0108] It should be understood that the prompt information for lyrics analysis shown in FIG. 3 is only an example, and the present disclosure does not specifically limit the format or auxiliary content of the prompt information for lyrics analysis.
[0109] FIG. 4 is a schematic diagram illustrating prompt information for music work analysis according to some embodiments of the present disclosure.
[0110] As shown in FIG. 4, the prompt information for music work analysis includes lyrics analysis results, audio features, and description information of at least one work analysis angle.
[0111] The lyrics analysis results include lyrics analysis results corresponding to at least one lyrics analysis angle. For example, referring to FIG. 4, the lyrics analysis results include lyrics analysis angle 1 and lyrics analysis result 1 corresponding to lyrics analysis angle 1, lyrics analysis angle 2 and lyrics analysis result 2 corresponding to lyrics analysis angle 2.
[0112] For example, the lyrics analysis angle and its corresponding lyrics analysis result in the lyrics analysis results can be represented in key-value pairs, including "theme: freedom and adventure", "emotional tendency: excitement and optimism", "language style: direct and lyrical", "symbolic meaning: uses travel and flight as symbols of freedom".
[0113] For example, referring to FIG. 4, the audio features include feature values corresponding to at least one feature. For example, referring to FIG. 4, the audio features include feature 1 and feature value 1 of feature 1, feature 2 and feature value 2 of feature 2.
[0114] For example, the feature and its corresponding feature value in the audio features can be represented in key-value pairs, including "rhythm: fast-paced, BPM is 140", "main instruments: electronic synthesizers and drum machines", "timbre features: clear and futuristic", "dynamic range: from medium to high, especially in the chorus part", "harmonic structure: mainly based on minor chords, occasionally interspersed with major chords to add color".
[0115] For example, referring to FIG. 4, the description information of the at least one work analysis angle includes description information of work analysis angle 1 and description information of work analysis angle 2.
[0116] In some embodiments, referring to FIG. 4, the prompt information for the work analysis further includes some information to assist the machine learning model to understand, such as “Based on the following lyrics analysis results and audio features, conduct a comprehensive music work analysis. Please consider the music style, emotional tendency, theme content, and how the music elements and lyrics work together to convey the core information and feelings of the music work in the report.” “Please include the following aspects in the analysis report” and “Please provide a detailed analysis report that deeply explores each of the above aspects to reveal the complexity and richness of the music work.” and the like.
[0117] For example, the work analysis perspectives can include music style positioning, emotional comprehensive analysis, musical expression of song theme, interaction between music and lyrics, target audience and use scenario, creative elements and artistic value, etc.
[0118] The description information of the work analysis perspective “music style positioning” includes, for example, “determine the possible music style or sub-style of the song based on audio features and lyrics content”.
[0119] The description information of the work analysis perspective “emotional comprehensive analysis” includes, for example, “describe the overall emotion and atmosphere that the song tries to convey by combining the rhythm, harmony and lyrics content of the music”.
[0120] The description information of the work analysis perspective “musical expression of song theme” includes, for example, “explore how musical elements (such as rhythm, harmony, timbre) support and enhance the theme in the lyrics”.
[0121] The description information of the work analysis perspective “interaction between music and lyrics” includes, for example, “analyze the interaction between musical elements and lyrics, how to create a coherent narrative and emotional experience together”.
[0122] The description information of the work analysis perspective “target audience and use scenario” includes, for example, “speculate the target audience of this song and what occasions they might most enjoy listening to this song”.
[0123] The description information of the work analysis perspective “creative elements and artistic value” includes, for example, “evaluate the characteristics of the song in terms of creative expression and artistic value, including any unique musical innovations or profound insights in the lyrics”.
[0124] It should be understood that the prompt information for music work analysis shown in FIG. 4 is only as an example, and the disclosure does not specifically limit the format or auxiliary content of the prompt information for music work analysis.
[0125] FIG. 5 is a schematic diagram showing the analysis results of a music work according to some embodiments of the disclosure.
[0126] As shown in FIG. 5, the analysis result of the music work includes at least one work analysis angle of the work analysis result, for example, including work analysis angle 1 and work analysis result 1 corresponding to work analysis angle 1, work analysis angle 2 and work analysis result 2 corresponding to work analysis angle 2. For example, the analysis result of the music work can also include summary information of the music work (not shown in FIG. 5). For example, the analysis result of the music work also includes some other description information, for example, "Based on the provided lyrics and audio features, the following is the comprehensive analysis result of the music work".
[0127] In some embodiments, the work analysis angle in the analysis result of the music work includes but is not limited to music style positioning, emotional comprehensive analysis, music expression of song theme, interaction between music and lyrics, target audience and use scenario, creative element and artistic value, etc. The work analysis result corresponding to each work analysis angle is not described in detail in the present disclosure.
[0128] It should be understood that the music work analysis result shown in FIG. 5 is only an example, and the present disclosure does not specifically limit the format or content of the music work analysis result.
[0129] In some embodiments, the label data and metadata generated based on FIGS. 3-5 are as follows.
[0130] For example, the multiple label dimensions can include music style, rhythm, main instrument, etc., and the labels can include electronic pop, 140 and electronic synthesizer, etc. For example, the multiple metadata fields can include title, artist, etc., and the field values corresponding to the metadata fields can include "Free Voice", "Electronic Dreamer", etc.
[0131] Table 1 shows the correspondence between the label dimensions and labels in some embodiments of the present disclosure. For example, the label dimensions can also be referred to as analysis dimensions, and the labels can be referred to as descriptions (i.e., descriptions of analysis dimensions).
[0132] Table 1 Correspondence between label dimensions and labels
[0133] Table 2 shows the correspondence between the metadata fields and field values in some embodiments of the present disclosure. For example, the metadata fields can also be referred to as metadata categories, and the field values can be referred to as descriptions (i.e., descriptions of metadata categories).
[0134] Table 2 Correspondence between metadata fields and field values
[0135] The above is the data generation method provided by some embodiments of the present disclosure. The data generation device in some embodiments of the present disclosure will be described below with reference to FIG. 6.
[0136] FIG. 6 is a block diagram illustrating a data generation apparatus according to some embodiments of the present disclosure.
[0137] As shown in FIG. 6, the data generation apparatus 6 of a musical work includes an extraction module 61, a lyrics analysis module 62, and a generation module 63.
[0138] The extraction module 61 is configured to extract an audio feature of music in a musical work, the audio feature reflecting an understanding of the music in the musical work. The lyrics analysis module 62 is configured to perform lyrics analysis on lyrics of the musical work to obtain a lyrics analysis result, the lyrics analysis result reflecting an understanding of content of the lyrics. The generation module 63 is configured to generate at least one of tag data and metadata of the musical work according to the audio feature and the lyrics analysis result.
[0139] The data generation apparatus 6 can be used to perform steps S10-S30 of FIG. 1. In some embodiments, the data generation apparatus 6 can also perform any of the steps shown in FIGS. 2A-2E.
[0140] It should be noted that each of the above modules is only a logical module according to the specific function it implements, and is not intended to limit the specific implementation manner, for example, it can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, each of the above modules can be implemented as an independent physical entity, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, each of the above modules is shown in the figure with a dashed line indicating that these modules can not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.
[0141] The above is a data generation apparatus in some embodiments of the present disclosure.
[0142] FIG. 7 is a block diagram illustrating a data generation apparatus according to some other embodiments of the present disclosure.
[0143] As shown in FIG. 7, the data generation apparatus 7 of a musical work includes a memory 71 and a processor 72 coupled to the memory 71, the processor 72 being configured to perform the data generation method of a musical work according to any of the preceding embodiments based on instructions stored in the memory 71.
[0144] The memory 71 is configured to store one or more computer-readable instructions. The memory 71 can include any combination of various types of computer-readable storage media, such as volatile memory and / or non-volatile memory including, but not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 71 may, for example, store an operating system, application programs, a boot loader, a database, and other programs, and can also store various application programs and various data.
[0145] The processor 72 is configured to run the computer-readable instructions to implement the data generation method of the musical work according to any one of the preceding embodiments. For specific implementation of each step of the data generation method of the musical work, reference can be made to the above-described embodiments, and repeated descriptions are not repeated here.
[0146] The processor 72 and the memory 71 can directly or indirectly communicate with each other. For example, the processor 72 and the memory 71 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 72 and the memory 71 can also communicate with each other through a system bus, and the present disclosure does not limit the communication between the processor 72 and the memory 71.
[0147] It should be noted that the components of the electronic device 7 shown in FIG. 7 are only exemplary and not limiting, and the electronic device 7 can also have other components according to actual application needs. The processor 72 can communicate with other components of the electronic device 7 to perform the desired functions.
[0148] The data generation apparatus can be implemented by software, firmware, and / or hardware, and can be integrated into an electronic device installed with a related application program.
[0149] FIG. 8 is a block diagram of an electronic device according to some embodiments of the present disclosure.
[0150] The electronic device 8 shown in FIG. 8 can be a computer system with a special hardware structure, which can perform corresponding functions when a related application program is installed.
[0151] The electronic device includes, but is not limited to, a mobile terminal such as a smartphone, a notebook computer, a personal digital assistant (PDA), a tablet computer (Tablet PC), a PMP (portable multimedia player), a vehicle terminal (such as a vehicle navigation terminal), a wearable device, and the like, and a fixed terminal such as a digital television, a desktop computer, and the like.
[0152] As shown in FIG. 8, a central processing unit (CPU) 81 executes various processes in accordance with a program stored in a read only memory (ROM) 82 or a program loaded from a storage section 88 to a random access memory (RAM) 83. In the RAM 83, data required when the CPU 81 executes various processes and the like is stored as necessary. The central processing unit is merely exemplary, and can be other types of processors, such as the various processors described above. The ROM 82, the RAM 83, and the storage section 88 can be various forms of computer readable storage media. Note that although the ROM 82, the RAM 83, and the storage section 88 are shown separately in FIG. 8, one or more of them can be combined, or located in the same or different memory or storage modules.
[0153] The CPU 81, the ROM 82, and the RAM 83 are connected to each other via a bus 84. An input / output interface 85 is also connected to the bus 84.
[0154] The following components are connected to the input / output interface 85: an input section 86 including a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output section 87 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage section 88 including a hard disk, a magnetic tape, and the like; and a communication section 89 including a network interface card such as a LAN card, a modem, and the like. The communication section 89 allows communication processing to be performed via a network such as the Internet. It is easily understood that although the various devices or modules in the electronic device 8 are shown in FIG. 8 as communicating via the bus 84, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0155] A drive 810 is also connected to the input / output interface 85 as necessary. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 810 as necessary, so that a computer program read therefrom is installed in the storage section 88 as necessary.
[0156] In the case where the above series of processes are implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium 811 or the like.
[0157] According to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the data generation method described in any of the preceding embodiments. The computer program product includes a computer program carried on a computer-readable medium, which contains program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from the network by the communication section 89, or installed from the storage section 88, or installed from the ROM 82. When the computer program is executed by the CPU 81, the data generation method of the embodiments of the present disclosure is executed.
[0158] Note that, in the context of the present disclosure, the computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0159] The computer-readable medium can be a computer-readable storage medium, or a computer-readable signal medium, or any combination of the two.
[0160] The computer-readable storage medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the data generation method described in any of the preceding embodiments.
[0161] The computer readable medium can include a computer-readable signal medium and / or computer-readable storage medium. A computer readable signal medium can include a propagated data signal with computer executable instructions. A computer readable storage medium can include any non-transitory medium that can store computer executable instructions such as volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, or a combination of the two. The computer readable storage medium can include, but is not limited to, volatile memory (e.g., DRAM, SRAM, etc.), non-volatile memory (e.g., ROM, PROM, and EPROM, EEPROM, Flash memory, etc.), and memory hardware separate from the processor, such as a ROM, a RAM, a flash memory, etc. The computer readable storage medium can also include, but is not limited to, tapes, papers, cards, etc. with holes or other patterns (e.g., a punch card, a computer punch tape, exposed film, etc.). The computer readable storage medium can also include one or more components of the electronic device (e.g., one or more components of the processor). Specifically, the computer readable storage medium can include the RAM alone or in combination with a non- volatile memory, and / or the ROM alone or in combination with the non-volatile memory, and / or other components of the electronic device 1000. The computer readable storage medium can also include distributed storage used to store data, consisting of storages physically located at different locations, and logically connected together to form a single storage. The computer readable storage medium can also include a combination of one or more of the above components.
[0162] The computer readable medium described above can be included in the electronic device; or exist separately from the electronic device.
[0163] In some embodiments, a computer program product is also provided, which, when running on a computer, causes the computer to implement the data generation method described in any of the above embodiments.
[0164] In some embodiments, a computer program is also provided, which includes instructions, which, when executed by a processor, cause the processor to perform the data generation method of any of the above embodiments. For example, the instructions can be embodied in computer program code.
[0165] In embodiments of the present disclosure, the computer program code for carrying out operations of the present disclosure can be written in one or more programming languages, or combinations of languages, including object oriented, scripting and / or procedural languages, such as Java, Smalltalk, C++, or the like. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). The computer program code can also be embodied in a computer readable storage medium that can be any media that can store code for use by a computer.
[0166] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing device, cause the one or more processors to perform the operations of the method of the first aspect. The one or more non-transitory computer-readable media can include one or more of the following: a magnetic storage device, an optical storage device, a solid-state storage device, a hard disk drive, a flash drive, a RAM, a ROM, a database, and a cache. The one or more processors can include one or more of the following: a central processing unit (CPU), a microprocessor, a microcontroller, a microcomputer, a microprocessor-based or microcontroller-based system, a programmable logic unit (PLU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system-on-chip (SoC), and a complex programmable logic device (CPLD).
[0167] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non-limiting examples of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0168] While certain aspects of the disclosure have been described with reference to particular examples, those of ordinary skill in the art will understand that various other modifications can be made to the examples without departing from the scope and spirit of the disclosure. The range of the disclosure is defined by the appended claims.
Claims
1. A method for generating data of a music work, comprising: extracting audio features of music in the music work, the audio features reflecting an understanding of the music in the music work; performing lyrics analysis on lyrics of the music work to obtain lyrics analysis results, the lyrics analysis results reflecting an understanding of content of the lyrics; generating at least one of tag data and metadata of the music work according to the audio features and the lyrics analysis results.
2. The data generating method according to claim 1, wherein The generating at least one of tag data and metadata of the music work according to the audio features and the lyrics analysis results comprises: generating prompt information for music work analysis, the prompt information for music work analysis comprising description information of a work analysis perspective, the audio features and the lyrics analysis results, the work analysis perspective focusing on the music work as a whole including the music and the lyrics; performing music work analysis on the music work by using a machine learning model according to the prompt information for music work analysis to obtain analysis results of the music work; generating at least one of tag data and metadata of the music work according to the analysis results of the music work.
3. The data generating method according to claim 2, wherein The work analysis perspective exists in multiple, and the analysis results of the music work comprise analysis results for multiple work analysis perspectives. The generating at least one of tag data and metadata of the music work according to the analysis results of the music work comprises: generating at least one of the following according to the analysis results for the multiple work analysis perspectives: based on multiple tag dimensions of the multiple work analysis perspectives and tags of the music work corresponding to each tag dimension, as the tag data of the music work; based on multiple metadata fields of the multiple work analysis perspectives and field values corresponding to each metadata field, as the metadata of the music work.
4. The data generating method according to any one of claims 1 to 3, wherein The performing lyrics analysis on the lyrics of the music work to obtain lyrics analysis results comprises: generating prompt information for lyrics analysis, the prompt information for lyrics analysis comprising the lyrics and description information of a lyrics analysis perspective, the lyrics analysis perspective focusing on the lyrics; performing lyrics analysis on the lyrics of the music work by using a machine learning model according to the prompt information for lyrics analysis to obtain the lyrics analysis results.
5. The data generating method according to claim 4, wherein The lyrics analysis perspective exists in multiple, and the performing lyrics analysis on the lyrics of the music work by using a machine learning model according to the prompt information for lyrics analysis to obtain the lyrics analysis results comprises: performing lyrics analysis on the lyrics of the music work by using a machine learning model according to the prompt information for lyrics analysis to obtain analysis results for multiple lyrics analysis perspectives as the lyrics analysis results. 6.The method of any one of claims 1-3, further comprising: obtaining feedback information for at least one of tag data and metadata of the music work. in response to a difference between the feedback information and at least one of the tag data and the metadata of the musical work, determining that the difference is associated with a relationship between the music and the lyrics; in response to the relationship between the difference and the music satisfying a preset condition, adjusting parameters in a stage of extracting the audio features and a stage of generating at least one of the tag data and the metadata of the musical work; in response to the relationship between the difference and the lyrics satisfying a preset condition, adjusting parameters in a stage of the lyrics analysis and a stage of generating at least one of the tag data and the metadata of the musical work.
7. The data generating method according to claim 2 or 3, wherein generating at least one of the tag data and the metadata of the musical work according to an analysis result of the musical work includes: performing text analysis on the analysis result of the musical work to obtain an analysis result; according to the analysis result, using a pre-set matching instruction or template to perform information extraction on the analysis result of the musical work to obtain key information; performing structured data conversion on the key information to obtain at least one of the tag data and the metadata of the musical work. generating at least one of the tag data and the metadata of the musical work according to an analysis result of the musical work includes:
8. The data generating method according to claim 2 or 3, wherein generating prompt information for the tag data and / or the metadata extraction, wherein the prompt information for the tag data and / or the metadata extraction includes the analysis result of the musical work and description information for an extraction target and a presentation format, the extraction target includes a tag dimension and / or a metadata field based on the work analysis angle, and the presentation format includes a format of a corresponding relationship between the tag dimension and a tag and / or a corresponding relationship between the metadata field and a field value; according to the prompt information for the tag data and / or the metadata extraction, generating at least one of the tag data and the metadata of the musical work using a machine learning model. generating prompt information for musical work analysis includes:
9. The data generating method of claim 2, wherein, inserting the audio features and the lyrics analysis result into a musical work analysis template, the musical work analysis template including description information of the work analysis angle. generating prompt information for lyrics analysis includes:
10. The data generating method of claim 4, wherein, obtaining lyrics content of the musical work; inserting the lyrics content into a lyrics analysis template to obtain the prompt information for lyrics analysis, the lyrics analysis template including description information of the lyrics analysis angle.
11. The data generation method of claim 1, wherein: the audio features include a numerical representation of the audio features and natural language description information corresponding to the numerical representation; and / or the audio features include physical features of an audio signal of the music and music features reflecting music properties based on the physical features.
12. A data generation apparatus for a musical work, comprising: an extraction module configured to extract audio features of music in a musical work, the audio features reflecting an understanding of the music in the musical work; a lyrics analysis module configured to perform lyrics analysis on lyrics of the music work to obtain lyrics analysis results reflecting an understanding of content of the lyrics; a generation module configured to generate at least one of tag data and metadata of the music work according to the audio features and the lyrics analysis results.
13. A data generation apparatus for a music work, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the data generation method according to any one of claims 1 to 11 based on instructions stored in the memory.
14. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the data generation method according to any one of claims 1 to 11.
15. A computer program product which, when running on a computer, causes the computer to implement the data generation method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Music genre classification method and device and storage medium
CN111414513A
Music structure analysis method, electronic equipment and storage medium
CN112989109A
Music recognition method, music recognition device, electronic equipment and storage medium
CN116645974A
Music feature processing method and device, equipment, medium and program product
CN116662781A