Music recognition model training method, music recognition method and related device
By adjusting the parameters of the feature mapping sub-model of the music recognition model, the problem of low recognition accuracy in the training of multimodal music recognition models is solved, and efficient and accurate music recognition is achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2026-03-05
AI Technical Summary
Because multimodal music recognition models contain a large number of model parameters, existing music recognition models trained using open-source music datasets have low recognition accuracy and cannot meet practical needs.
The music recognition model was trained iteratively using a sample dataset, including the already trained feature extraction sub-model and text recognition sub-model. Only the feature mapping sub-model was trained, and the recognition accuracy of the model was improved by adjusting the parameters of the feature mapping sub-model.
This reduces training difficulty and time, avoids problems such as insufficient data and overfitting, and improves the recognition accuracy of the music recognition model.
Smart Images

Figure CN2025107370_05032026_PF_FP_ABST
Abstract
Description
A training method for a music recognition model, a music recognition method, and related equipment.
[0001] Related applications
[0002] This application claims priority to Chinese patent application filed on August 26, 2024, application number 202411170712.8, entitled "A training method for a music recognition model, a music recognition method and related equipment", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of computer technology, and in particular to a training method for a music recognition model, a music recognition method, and related equipment. Background Technology
[0004] With the continuous development of technology, more and more devices can identify music through multimodal music recognition models, achieving the goal of describing music in text form. This allows devices to possess richer music recognition functions. For example, when music is playing in an environment, the device can identify the music through a music recognition model and describe it for the user from multiple aspects such as the music title, performer, and instruments. Thus, the device can provide users with music-related Q&A services, music search services for similar music styles, or music-based social networking services.
[0005] In related technologies, in order to describe music in text form, the model structure of music recognition models usually adopts a multimodal large language model structure, and the method of training music recognition models usually involves using open-source music datasets from online resources to perform multiple rounds of iterative training on the music recognition model to be trained.
[0006] In each iteration of training, the music recognition model to be trained is used to extract features from the music data, obtain music features, and identify the obtained music features to obtain corresponding text descriptions. Based on the differences between the text descriptions and the sample descriptions associated with the music data, all model parameters of the music recognition model are adjusted.
[0007] However, because multimodal music recognition models involve a larger number of model parameters to recognize music features as text descriptions compared to single-modal recognition models, multimodal music recognition models require a much larger amount of music data for training.
[0008] However, since the use cases of open-source music datasets on the Internet are relatively niche, the amount of data in open-source music datasets is insufficient to support the training of music recognition models with such a large number of model parameters, resulting in low recognition accuracy of music recognition models trained on open-source music datasets.
[0009] It is evident that the methods used to train music recognition models under related technologies result in music recognition models with low accuracy. Summary of the Invention
[0010] This application provides a training method for a music recognition model, a music recognition method, and related equipment to solve the problem of low recognition accuracy of the trained music recognition model.
[0011] Firstly, a method for training a music recognition model, executed by a computer device, is provided, including:
[0012] Obtain a sample dataset; wherein each sample dataset includes: a sample music segment and a sample music description corresponding to the sample music segment, the sample music description containing multiple sample music attributes of the sample music segment;
[0013] Based on the aforementioned sample dataset, the music recognition model to be trained undergoes multiple rounds of iterative training; wherein, the music recognition model includes: a trained feature extraction sub-model and a text recognition sub-model, and a feature mapping sub-model to be trained; each round of iterative training includes:
[0014] The feature extraction sub-model is used to extract the spectral features of the sample music fragments;
[0015] The obtained spectral features are mapped to descriptive text features using the feature mapping sub-model; wherein the descriptive text features are used to describe various training music attributes of the sample music fragment;
[0016] Using the text recognition sub-model, the features of the descriptive text are recognized to obtain the training music description; and
[0017] The model parameters of the feature mapping sub-model are adjusted based on the differences between the obtained training music description and the sample music description corresponding to the training music description.
[0018] Secondly, a music recognition method is provided, including:
[0019] Obtain the music to be identified;
[0020] A trained music recognition model is used to extract the spectral features of the music to be recognized, and the spectral features are mapped to text features to be recognized; wherein the music recognition model is trained using the method described in the first aspect.
[0021] The music recognition model is used to identify the features of the text to be identified and obtain a target music description; wherein the target music description is used to: introduce various target music attributes of the music to be identified in text form.
[0022] Thirdly, a training device for a music recognition model is provided, comprising:
[0023] Acquisition module: used to acquire sample datasets; wherein each sample dataset includes: a sample music segment and a sample music description corresponding to the sample music segment, the sample music description containing multiple sample music attributes of the sample music segment;
[0024] Processing module: used to perform multiple rounds of iterative training on the music recognition model to be trained based on the sample dataset; wherein, the music recognition model includes: a trained feature extraction sub-model and a text recognition sub-model, and a feature mapping sub-model to be trained; each round of iterative training includes:
[0025] The processing module is used to: extract the spectral features of the sample music segment using the feature extraction sub-model;
[0026] The processing module is used to: map the obtained spectral features into descriptive text features using the feature mapping sub-model; wherein the descriptive text features are used to describe various training music attributes of the sample music fragment;
[0027] The processing module is used to: employ the text recognition sub-model to recognize the features of the descriptive text and obtain a training music description; and
[0028] The processing module is used to adjust the model parameters of the feature mapping sub-model based on the difference between the obtained training music description and the sample music description corresponding to the training music description.
[0029] Fourthly, a music recognition device is provided, comprising:
[0030] Acquisition module: Used to acquire the music to be identified;
[0031] Processing module: used to extract the spectral features of the music to be identified using a trained music recognition model, and map the spectral features to be identified into text features to be identified; wherein the music recognition model is trained using the method described in the first aspect;
[0032] The processing module is further configured to: use the music recognition model to identify the features of the text to be identified and obtain a target music description; wherein the target music description is used to: describe various target music attributes of the music to be identified in text form.
[0033] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method as described in the first aspect or the second aspect.
[0034] Sixthly, a computer device is provided, comprising:
[0035] Memory, used to store program instructions;
[0036] A processor is configured to invoke program instructions stored in the memory and execute the method as described in the first aspect or the second aspect according to the obtained program instructions.
[0037] A seventh aspect provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the method as described in the first aspect or the second aspect.
[0038] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.
[0040] Figure 1A is a schematic diagram of an application field of the training method of the music recognition model provided in the embodiment of this application;
[0041] Figure 1B is a schematic diagram of an application field of the training method of the music recognition model provided in the embodiments of this application;
[0042] Figure 1C illustrates an application scenario of the training method for the music recognition model provided in this embodiment of the application;
[0043] Figure 2 is a schematic flowchart of a training method for a music recognition model provided in an embodiment of this application;
[0044] Figure 3A is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application;
[0045] Figure 3B is a schematic diagram of the principle of a training method for a music recognition model provided in an embodiment of this application;
[0046] Figure 4A is a schematic diagram of the principle of a training method for a music recognition model provided in an embodiment of this application;
[0047] Figure 4B is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0048] Figure 4C is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0049] Figure 4D is a schematic diagram of the principle of a training method for a music recognition model provided in an embodiment of this application;
[0050] Figure 5A is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0051] Figure 5B is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0052] Figure 5C is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0053] Figure 5D is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application;
[0054] Figure 6 is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0055] Figure 7A is a schematic flowchart of a training method for a music recognition model provided in an embodiment of this application.
[0056] Figure 7B is a schematic diagram illustrating the principle of a training method for a music recognition model provided in an embodiment of this application.
[0057] Figure 8A is a schematic diagram of a training device for a music recognition model provided in an embodiment of this application;
[0058] Figure 8B is a schematic diagram of a music recognition device provided in an embodiment of this application;
[0059] Figure 9 is a schematic diagram of another structure of the training device or music recognition device for the music recognition model provided in the embodiments of this application. Detailed Implementation
[0060] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0061] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.
[0062] (1) Multimodal:
[0063] In the field of machine learning, multimodal learning refers to learning tasks and techniques that involve two or more different types of data modalities. These data modalities can include, but are not limited to, text, images, audio, and video. Multimodal learning aims to improve the performance of models by combining different types of input data.
[0064] (2) Pre-training:
[0065] The basic idea of pre-training is to first train a base model on a large amount of unlabeled data, and then apply this base model to a specific task. It usually requires further fine-tuning to adapt to the specific downstream task.
[0066] (3) Large Language Model (LLM):
[0067] Large language models refer to language models with a large number of parameters and high complexity. These models are typically built using deep learning techniques and trained on large-scale datasets. Large language models can process and generate natural language text, possessing powerful language understanding and generation capabilities.
[0068] (4) Music Large Language Model (MusicLLM or musicLLM):
[0069] Music large language models refer to a class of multimodal large language models specifically designed for music generation and understanding. These models typically utilize deep learning techniques, especially neural network architectures, to learn complex patterns in music data and are capable of generating new musical compositions or performing music-related tasks.
[0070] (5) Waveform:
[0071] Waveforms are used to describe the shape of a signal or wave as it changes over time. They are used in fields such as physics, signal processing, audio engineering, and electronics. A waveform can be a sound wave, an electrical signal, or any fluctuation that can be represented as a function of time. Waveforms can be observed and recorded using various devices (such as an oscilloscope). Common waveforms include sine waves, square waves, triangle waves, and sawtooth waves.
[0072] (6) Transformer structure:
[0073] The transformer architecture can include an encoder-decoder architecture, where the encoder is responsible for converting the input sequence into an intermediate representation, and the decoder is responsible for using the representation generated by the encoder to generate the output sequence.
[0074] The transformer architecture can also include a self-attention mechanism layer, where multi-head attention allows the model to learn information from different representation subspaces; positional encoding is used to preserve the positional information of the input sequence.
[0075] The transformer architecture can also include a feedforward neural network, which can be placed after each self-attention mechanism layer for further feature processing.
[0076] The transformer structure can also include residual connections, placed before and after each sub-layer, which helps to alleviate the vanishing gradient problem and accelerate the training process.
[0077] The transformer structure can also include layer normalization applied to the output of each sub-layer to help stabilize the training process.
[0078] It should be noted that the embodiments of this application involve operations such as obtaining sample datasets. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0079] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0080] The training method of the music recognition model and the application fields of the music recognition method provided in the embodiments of this application are briefly introduced below.
[0081] With the continuous development of technology, more and more devices can recognize music through multimodal music recognition models (such as music large language models) to achieve the purpose of describing music in the form of text. As a result, devices can have richer music recognition functions.
[0082] For example, referring to Figure 1A, when music is playing in the environment, the device can identify the music through a music recognition model and describe the music for the user from multiple aspects such as music name, performer, and instrument.
[0083] For example, referring to Figure 1B, when a device plays a piece of music, it can identify the music using a music recognition model. When the device collects a question or text related to the music, it can generate a corresponding response. For instance, if the question is "In what year was the music that was just played composed?", the response would include at least "The music was composed in 1895".
[0084] Through music recognition models, devices can provide music Q&A services for specific music users, as well as music search services for similar music styles or music-based dating services, without any specific restrictions.
[0085] In related technologies, in order to describe music in text form, the model structure of music recognition models usually adopts a multimodal large language model structure, and the method of training music recognition models usually involves using open-source music datasets from online resources to perform multiple rounds of iterative training on the music recognition model to be trained.
[0086] In each iteration of training, the music recognition model to be trained is used to extract features from the music data, obtain music features, and identify the obtained music features to obtain corresponding text descriptions. Based on the differences between the text descriptions and the sample descriptions associated with the music data, all model parameters of the music recognition model are adjusted.
[0087] However, because multimodal music recognition models involve a larger number of model parameters to recognize music features as text descriptions compared to single-modal recognition models, multimodal music recognition models require a much larger amount of music data for training.
[0088] However, since the use cases of open-source music datasets on the Internet are relatively niche, the amount of data in open-source music datasets is insufficient to support the training of music recognition models with such a large number of model parameters, resulting in low recognition accuracy of music recognition models trained on open-source music datasets.
[0089] It is evident that the methods used to train music recognition models under related technologies result in music recognition models with low accuracy.
[0090] To address the issue of low recognition accuracy in trained music recognition models, this application proposes a training method for such models. This method involves acquiring a sample dataset, where each sample dataset includes a sample music segment and a corresponding music description. The music description contains various music attributes of the sample music segment. Based on the sample dataset, the music recognition model to be trained undergoes multiple rounds of iterative training. The music recognition model includes a trained feature extraction sub-model and a text recognition sub-model, as well as a feature mapping sub-model to be trained.
[0091] Each round of iterative training includes:
[0092] A feature extraction sub-model is used to extract the spectral features of the sample music fragments. A feature mapping sub-model is then used to map the obtained spectral features to descriptive text features. These descriptive text features are used to describe various training music attributes of the sample music fragments. A text recognition sub-model is then used to recognize the descriptive text features, obtaining training music descriptions. Based on the differences between the obtained training music descriptions and the corresponding sample music descriptions, the model parameters of the feature mapping sub-model are adjusted.
[0093] In this embodiment, a pre-trained text recognition sub-model, which is relatively easier to obtain, is used to build the model structure of the music recognition model to be trained. It is not necessary to train the large number of model parameters contained in the multimodal recognition model structure, nor is it necessary to train the large model such as the text recognition sub-model. This reduces the training difficulty and shortens the training time from multiple perspectives, thereby improving the training efficiency.
[0094] Furthermore, since the training process of the music recognition model to be trained does not require training the large language model of the text recognition sub-model in any music recognition scenario, it avoids the problems of low training accuracy of the music recognition model due to insufficient sample data or overfitting, and low recognition accuracy of the trained music recognition model.
[0095] Furthermore, when training the music recognition model, only the feature mapping sub-model used to map spectral features to text features needs to be trained. This aligns the feature space of the music modality with the feature space of the text modality, greatly reducing the amount of data required to train the model parameters in the music recognition model. As a result, a trained music recognition model with high recognition accuracy can be obtained without too much sample data. This reduces the difficulty of obtaining sample data while ensuring the recognition accuracy of the trained music recognition model.
[0096] The following describes the application scenarios of the training method for the music recognition model provided in this application.
[0097] Please refer to Figure 1C, which is a schematic diagram of an application scenario for the training method of the music recognition model provided in this application. This application scenario includes a client 101 and a server 102. The client 101 and the server 102 can communicate with each other. The communication method can be wired, such as through a network cable or serial cable; or wireless, such as through Bluetooth or Wireless Fidelity (WIFI). No specific limitation is imposed.
[0098] Client 101 generally refers to devices capable of capturing music and displaying corresponding music descriptions, such as terminal devices, third-party applications accessible by the terminal devices, or web pages accessible by the terminal devices. Terminal devices include, but are not limited to, mobile phones, computers, smart medical devices, smart home appliances, vehicle terminals, or aircraft. Server 102 generally refers to devices capable of training and using music recognition models, such as terminal devices or servers. Servers include, but are not limited to, cloud servers, local servers, or associated third-party servers. Both client 101 and server 102 can utilize cloud computing to reduce the consumption of local computing resources; similarly, they can also utilize cloud storage to reduce the consumption of local storage resources.
[0099] As one embodiment, the client 101 and the server 102 can be the same device or different devices, and there is no specific limitation.
[0100] The training method of the music recognition model provided in this application embodiment will be described in detail below based on Figure 1C. Please refer to Figure 2, which is a flowchart illustrating one method for training the music recognition model provided in this application embodiment.
[0101] S201, Obtain the sample dataset; wherein each sample dataset includes: a sample music segment and a sample music description corresponding to the sample music segment, and the sample music description contains various sample music attributes of the sample music segment.
[0102] A sample dataset is a collection of multiple sample data. Each sample data contains a sample music segment and its corresponding sample music description. The sample music description contains various sample music attributes of the sample music segment. It can be constructed by collecting music multimedia files and metadata, performing segment segmentation, music attribute identification, and description generation.
[0103] Before training the music recognition model, you can first obtain a sample dataset, such as obtaining a sample dataset from online resources; or you can construct a sample dataset based on music resources in online resources, etc. There are no specific restrictions.
[0104] Each sample data includes: a sample music segment and a corresponding sample music description. The sample music description contains various sample music attributes of the sample music segment. The sample music segment is a piece of music of a certain duration, which can be represented by data points arranged in chronological order or by symbols.
[0105] For example, it can be represented in the form of a waveform diagram, please refer to (1) in Figure 3A; or it can be represented in the form of a digital multimedia file, such as MP3 format, AAC (Advanced Audio Coding) format or FLAC (Free Lossless Audio Codec) format, please refer to (2) in Figure 3A; or it can be represented in the format defined by Musical Instrument Digital Interface (MIDI), please refer to (3) in Figure 3A; or it can be represented in the form of ASCII characters, please refer to (4) in Figure 3A.
[0106] The duration of a sample music segment can be preset, such as 30 seconds; or it can be a duration that ensures the integrity of the lyrics, melody, or rhythm, etc., with no specific restrictions. For example, a complete sample music can be divided into multiple sample music segments according to the punctuation or paragraph divisions in the lyrics; or, for another example, a complete sample music can first be divided into multiple music segments according to the punctuation or paragraph divisions in the lyrics, and then each music segment can be divided into multiple sample music segments according to the rhythm, etc., with no specific restrictions.
[0107] The sample music description includes various sample music attributes of the sample music fragment. For example, the sample music description uses fluent language to describe the various sample music attributes of the sample music fragment; for another example, the sample music description uses classical Chinese, prose, or poetry to describe the various sample music attributes of the sample music fragment; for yet another example, the sample music description uses a combination of text and images to describe each sample music attribute more intuitively and vividly, etc., and there are no specific restrictions.
[0108] For example, the sample music description could be: "This is a piece of pop music, recorded live, and incorporates elements of a less common language. The music features a large number of string instruments, including cello, double bass, violin, and fiddle. The tempo is 71 bpm, and the string instruments create a melodious atmosphere. The climax uses a string section, creating a highly infectious and captivating musical experience."
[0109] Among them, "pop", "live", "minority language music elements", "Cello", "double bass", "Violin", "fiddle", "71bpm", "string instruments" and "string section" are various sample music attributes.
[0110] As an example, considering that the datasets in online resources are mainly music descriptions in English, and since the number of model parameters that need to be trained in the music recognition model provided in this application embodiment is relatively small, music resources can be obtained from online resources to construct a sample dataset to obtain music descriptions in any language, providing flexibility in training the music recognition model and improving the applicability of the trained music recognition model.
[0111] Multiple multimedia files of different music tracks can be collected from online resources or other devices, along with their respective metadata. Each multimedia file can be any of the representations of the sample music clips described earlier, or other representations; there are no specific restrictions. The metadata of the multimedia files includes, for example, the music title, performer, singer, lyrics, comments, and tags added to the music when it was uploaded to online resource platforms; there are no specific restrictions.
[0112] Taking sample music segments of the same duration as an example, such as 30 seconds, we can divide each multimedia file into segments based on the preset segment duration, obtaining multiple sub-files for each multimedia file, and using each sub-file as a sample music segment. For example: Based on the preset segment duration, we divide each multimedia file into segments, obtaining multiple sub-files for each multimedia file, and using each sub-file as a sample music segment. For audio multimedia files, we can calculate the number of sampling points in each segment based on the audio sampling rate and the preset segment duration, and then divide the audio file segment by segment according to this number of sampling points. For example, if the audio sampling rate is 44100Hz and the preset segment duration is 30 seconds, then each segment contains 44100 × 30 sampling points. Starting from the beginning of the audio file, every 44100 × 30 sampling points constitutes one sub-file. For video multimedia files, we first extract the audio stream, and then divide it into segments according to the above audio segmentation method.
[0113] There are various methods for determining the musical attributes of sample music fragments. For example, a trained multimodal recognition model can be used. This model analyzes the spectral and temporal features of the sample music fragments, and uses predefined classification rules and feature matching algorithms to identify the musical attributes of each fragment, thus obtaining multiple musical attributes for each fragment. The trained multimodal recognition model is used to identify input music fragments as having multiple musical attributes. This model can be trained on open-source databases available online, such as 1-second music fragments and their corresponding English language musical attributes; it can also be a large open-source multimodal model available online, and there are no specific limitations.
[0114] For example, there are pre-stored reference music fragments corresponding to various reference music attributes. By determining the music similarity between the sample music fragment and each reference music fragment, the reference music attribute corresponding to the reference music fragment with a music similarity greater than a preset threshold is used as the sample music attribute of the sample music fragment.
[0115] Musical similarity is an indicator used to measure the degree of similarity between a sample music segment and a reference music segment. By comparing the musical features of the two, the reference music attributes corresponding to the reference music segment with a similarity greater than a preset threshold are used as the sample music attributes of the sample music segment. Various reference music attributes can include musical attributes measured from multiple perspectives such as style, theme, rhythm, meter, instruments, voice type, pitch, and genre. Voice type, for example, includes male voice, female voice, rain sound, wind sound, ambient sound, and animal sounds.
[0116] After obtaining each sample music fragment, its metadata, and sample music attributes, a trained large language model can be used. This model takes the obtained metadata and sample music attributes as input and, through its internal language generation mechanism, combines pre-trained language patterns and semantic understanding to generate sample music descriptions for each sample music fragment. Thus, a sample dataset can be obtained based on each sample music fragment and its corresponding sample music description.
[0117] The trained large language model is used to arrange input words into phrases in a specified language. If the input words are in a language other than the specified language, the trained large language model can translate the words in that other language into the specified language and then arrange the translated words into phrases. Therefore, regardless of the language in which the obtained sample music attributes are represented, they can be converted to the specified language when generating sample music descriptions, reducing the language requirements for obtaining sample music attributes and thus lowering the difficulty of obtaining these attributes. An example of a large language model is the Chinese large language model, used to generate high-quality Chinese text, including dialogues, articles, and poems.
[0118] The process of building a sample dataset does not require manual annotation of sample data, which greatly reduces the cost of data annotation. Furthermore, the use of open-source models makes the process of building a sample dataset more efficient and convenient, and it can build a sample dataset containing a large number of sample data according to the usage requirements.
[0119] Please refer to Figure 3B. Taking a piece of music as an example, after collecting the multimedia file and metadata of the music, divide the multimedia file into multiple sub-files according to 30-second sample music segments, such as 6 sub-files. Each sub-file is a sample music segment, that is, obtain sample music segment A, sample music segment B, sample music segment C, sample music segment D, sample music segment E and sample music segment F.
[0120] Each sample music segment is input into a trained multimodal recognition model to obtain multiple sample music attributes output by the multimodal recognition model for that sample music segment. The metadata and multiple sample music attributes of each sample music segment are input into a trained large language model to obtain the sample music description of that sample music segment output by the large language model, i.e., sample music description A, sample music description B, sample music description C, sample music description D, sample music description E, and sample music description F.
[0121] Large language models can perform semantic recognition based on lyrics in metadata, and integrate the atmosphere created by the semantics of the lyrics into the description of various sample music attributes. For example, if a sample music attribute is "orchestral music" and the lyrics have a light and cheerful semantic, then the sample music description may include statements such as "a light and cheerful atmosphere is created through the performance of orchestral music".
[0122] Large language models can also perform person identification based on the singer in the metadata, and integrate the singer's timbre into the description of various sample music attributes. For example, if a sample music attribute is "pop" and the singer's timbre is relatively soft, then the sample music description may include statements such as "singing this Cpop music with a warm and soft voice, like a summer evening breeze, both sweet and full of nostalgia".
[0123] Large language models can also describe the attributes of sample music based on the characteristics of the attributes themselves. For example, if the attribute of sample music is "irregular beat", then the description of sample music can include statements such as "using irregular beats to create tension and mystery, conveying complex or unconventional emotions".
[0124] This allows each sample music fragment and its corresponding sample music description to be added to the sample dataset, thus achieving the goal of establishing the sample dataset. Through vivid and figurative sample music descriptions, when training a music recognition model based on the sample dataset, the trained music recognition model can learn the vivid and figurative descriptive methods in the sample music descriptions. This enables the music recognition model to output more readable and vivid descriptive text, showcasing the charm of Chinese language and reaching the level of a writer's literary creation.
[0125] S202, based on the sample dataset, performs multiple rounds of iterative training on the music recognition model to be trained.
[0126] After obtaining the sample dataset, the music recognition model to be trained can be iteratively trained multiple times based on the sample dataset. The music recognition model to be trained includes the already trained feature extraction sub-model and text recognition sub-model, as well as the feature mapping sub-model to be trained.
[0127] Since the feature extraction sub-model and text recognition sub-model are already trained, using the sample dataset, only the feature mapping sub-model in the music recognition model needs to be trained. Compared to training the entire music recognition model or a large-scale model structure, the sample dataset contains fewer samples, making it easier and faster to construct and more efficient in the training process.
[0128] Please refer to Figure 4A. After inputting the sample data into the trained feature extraction sub-model, the sample data is processed layer by layer using the trained feature extraction sub-model, the feature mapping sub-model to be trained, and the trained text recognition sub-model. By adjusting the model parameters of the feature mapping sub-model to be trained, this round of iterative training for the music recognition model to be trained is completed.
[0129] As one example, before iterating through multiple rounds of training on the music recognition model to be trained based on a sample dataset, the music recognition model to be trained can be built first. Since the music recognition model to be trained can include a trained feature extraction sub-model and a text recognition sub-model, as well as a feature mapping sub-model to be trained, for example, the trained feature extraction sub-model and text recognition sub-model can be obtained first, and then the network layer structure of the feature mapping sub-model to be trained can be built.
[0130] For example, one could first build a feature extraction sub-model to be trained, then perform multiple rounds of iterative training on this sub-model to obtain a trained feature extraction sub-model. After obtaining the trained text recognition sub-model, and finally, build the network layer structure of the feature mapping sub-model to be trained. Thus, if the sample dataset is constructed from collected multimedia music files, using music similar to or collected from the sample dataset to train the feature extraction sub-model can enable it to more accurately extract the features of each sample in the sample dataset, thereby improving the accuracy of the trained feature mapping sub-model and resulting in a more accurate mapping model.
[0131] Based on the structure building strategy of the music recognition model, this paper establishes a feature extraction sub-model and a feature mapping sub-model to be trained, as well as obtains a trained text recognition sub-model. This structure building strategy can indicate the network layer structure of the feature extraction sub-model, the network layer structure of the feature mapping sub-model, and the location where the text recognition sub-model is obtained, such as a download address from an online resource; no specific restrictions are imposed.
[0132] For example, referring to Figure 4B, the structure building strategy indicates that the network layer structure of the feature extraction sub-model is the model structure of the AudioMAE model, or a 24-layer cascaded transformer structure, etc.; and indicates that the network layer structure of the feature mapping sub-model is a linear transformation layer; and indicates that the text recognition sub-model is the internlm2-chat-7b Chinese large model.
[0133] After collecting multiple audio data sets, and establishing a feature extraction sub-model and a feature mapping sub-model to be trained, as well as obtaining a trained text recognition sub-model, the feature extraction sub-model to be trained can be iteratively trained multiple times based on the collected audio data sets, outputting a trained feature extraction sub-model. Thus, a music recognition model to be trained can be established based on the trained feature extraction sub-model, the text recognition sub-model, and the feature mapping sub-model to be trained.
[0134] Multiple audio data can include individual music segments from the sample dataset or open-source audio datasets from online resources. Since the feature extraction sub-model can be trained unsupervised and does not require labeling, the number of collected multiple audio data can be much greater than the number of sample data contained in the sample dataset. For example, the sample dataset contains 20 million sample data, while the multiple audio data contains 170 million, in order to ensure the accuracy of feature extraction by the trained feature extraction sub-model.
[0135] As an example, when training the feature extraction sub-model, we will take one round of iterative training as an example. The training process for each other round of iterative training is similar and will not be described in detail here.
[0136] Subdata masking is performed on the audio data to obtain masked data that masks random subdata.
[0137] Data masking refers to the process of randomly selecting a certain proportion of sub-data blocks for masking during the training of audio data. This results in masked data containing only sub-data other than the masked random sub-data. The original positions of these random sub-data are marked with specific placeholders to ensure that the audio data can be processed in a uniform format. This avoids data processing errors caused by complex and diverse data formats and improves the accuracy of data processing.
[0138] For example, setting the masking probability to 0.2 means randomly selecting 20% of the sub-data blocks for masking, resulting in masked data that conceals the random sub-data. Audio data can also be converted to a specified format, such as spectrogram format (Mel frequency cepstral coefficients MFCC or log-Mel spectrum), and then sub-data masking can be performed on the converted audio data to ensure that the audio data can be processed in a uniform format. This avoids data processing errors caused by complex and diverse data formats and improves the accuracy of data processing.
[0139] Audio data can be divided into multiple sub-data. Taking spectrogram format audio data as an example, audio data can be divided into multiple audio patches, each representing spectral information within a continuous time window. Therefore, when performing sub-data masking on audio data, any one of the multiple sub-data segments—that is, a random sub-data segment—can be masked, so that the resulting masked data contains only the sub-data segments other than the random sub-data segment, where the original location of the random sub-data segment is marked with a specific placeholder.
[0140] Feature extraction is performed on masked data to obtain data features. Data features refer to the features obtained after feature extraction of masked data. These features can be obtained by converting the masked data into a spectrogram format and then using a convolutional neural network to extract features from the spectrogram. Alternatively, the masked data can be encoded into a data vector format first, and then feature extraction can be performed to capture the complex and diverse spectral features in the audio data.
[0141] Based on the obtained data features, data prediction can be performed on the masked random sub-data in the masked data to obtain predicted sub-data. Based on the difference between the obtained predicted sub-data and the random sub-data, the model parameters of the feature extraction sub-model can be adjusted. Specifically, based on the difference between the obtained predicted sub-data and the random sub-data, the stochastic gradient descent algorithm can be used to adjust the model parameters of the feature extraction sub-model.
[0142] Predicted sub-data refers to the data obtained by predicting the random sub-data that was masked in the masked data based on the obtained data features. The model parameters of the feature extraction sub-model are adjusted by comparing the differences between the predicted sub-data and the random sub-data.
[0143] The greater the difference between the predicted sub-data and the random sub-data, the less accurately the feature extraction sub-model can extract the data features of the audio data. Therefore, the model parameters of the feature extraction sub-model can be adjusted, and the next round of iteration training can begin. The smaller the difference between the predicted sub-data and the random sub-data, the more accurately the feature extraction sub-model has learned to extract the data features of the audio data. Therefore, when the difference between the predicted sub-data and the random sub-data is less than a preset value, the model parameters of the feature extraction sub-model in this round can be used as the final model parameters, and the trained feature extraction sub-model can be output.
[0144] By masking sub-data, the feature extraction sub-model can be trained unsupervised without data labeling, thus obtaining a trained feature extraction sub-model with accurate feature extraction capabilities.
[0145] Please refer to Figure 4C. Taking a 30-second piece of music as an example, a 30-second piece of music can more completely present the melody, rhythm and other musical attributes of the music compared to a shorter piece of music such as 1 second. Therefore, in this embodiment, various music segments (such as sample music segments) and audio data are all music of a relatively long duration of 30 seconds. The training method of the music recognition model provided in this embodiment can recognize such long music and give corresponding music descriptions.
[0146] In each round of training the feature extraction sub-model, the 30-second audio data is divided into 128 sub-data points arranged in chronological order. Each of the 128 sub-data points is then vector-encoded to obtain 128 sub-data vectors of size (1, 1024) arranged in the aforementioned chronological order, denoted as sub-data vector_1, sub-data vector_2, ..., and sub-data vector_128.
[0147] The feature extraction sub-model is used to randomly replace a sub-data vector (e.g., sub-data vector_1) with a specific identifier (e.g., X) to obtain a sequence of sub-data vectors representing the masked data; the feature extraction sub-model is used to extract features from the sub-data vector sequence, and the sub-data vector_1 is predicted based on the obtained data features to obtain the predicted vector.
[0148] Therefore, based on the difference between the predicted vector and the sub-data vector_1, the model parameters of the feature extraction sub-model can be adjusted, and the next round of iterative training can be entered until the difference reaches the training objective, at which point the trained feature extraction sub-model is output.
[0149] The training process of the feature mapping sub-model is described in detail below. Please refer to S2021 to S2024. Taking one round of iterative training as an example, the training process of each other round of iterative training is similar and will not be described in detail here.
[0150] S2021 employs a feature extraction sub-model to extract the spectral features of the sample music segments.
[0151] The trained feature extraction sub-model has the ability to extract the spectral features of music. Therefore, it can be used to extract the spectral features presented by sample music fragments. If the sample music fragments are in a uniform format, feature extraction can be performed directly using the feature extraction sub-model. If they are not in a uniform format, the sample music fragments can first be converted to a specific format, such as a spectrogram format, before feature extraction is performed on the converted fragments. This specific format can be determined based on the format used during the training of the feature extraction sub-model, and there are no specific restrictions.
[0152] Spectral features are used to characterize information about musical attributes such as frequency components, timbre, and rhythm in music. Spectral features include, for example, spectrograms, Mel-frequency cepstral coefficients (MFCCs), fundamental frequency, harmonics, frequency bandwidth, energy, zero crossover rate, spectral roll-off, spectral peaks, spectral entropy, spectral flatness, spectral centroid, spectral slope, spectral band ratio, frequency resolution, and temporal features, etc., without any specific limitations.
[0153] As one embodiment, when extracting the spectral features of a sample music segment using a feature extraction sub-model, the sample music segment can first be divided into multiple sub-segments. The spectral sub-features of each sub-segment are extracted, as well as the spectral correlation features between these sub-features, such as the spectral correlation features between any two sub-features. Then, the obtained spectral sub-features and spectral correlation features are fused to obtain the spectral features presented by the sample music segment. Therefore, not only can the spectral sub-features of each sub-segment be extracted from a microscopic perspective, but the spectral correlation features between these sub-features can also be extracted from a macroscopic perspective, ensuring the accuracy of feature extraction.
[0154] Among them, spectral sub-features refer to the spectral characteristics of each sub-segment after the sample music segment is divided into multiple sub-segments. They can characterize at least one of the following: pitch, timbre, recording environment, rhythm, instrument, and emotion of the sub-segment, as well as at least one of the spectral features. Spectral correlation features refer to the correlation features between the various spectral sub-features of the sample music segment. They can characterize at least one of the following: similarity, rhythmic relationship, and pitch transition between the sub-segments, as well as at least one of the spectral features.
[0155] The obtained spectral sub-features and spectral correlation features can be fused using a weighted summation method to obtain the spectral features presented by the sample music segment. Specifically, this can be achieved using the formula... Where F represents the spectral feature, m represents the number of spectral sub-features, and w i For the spectral sub-feature f iThe weights, k represents the number of spectral correlation features, v j For spectral correlation features g j The weight.
[0156] Referring to Figure 4D, the sample music segment is divided into 128 sub-segments, from which 128 spectral sub-features can be extracted, denoted as sub-feature vector_1, sub-feature vector_2, ..., and sub-feature vector_128. For example, the extracted spectral correlation features include: correlation feature vector_12 between sub-feature vector_1 and sub-feature vector_2; correlation feature vector_24 between sub-feature vector_2 and sub-feature vector_4; correlation feature vector_57 between sub-feature vector_5 and sub-feature vector_7, etc. During fusion, sub-feature vector_1 and correlation feature vector_12 are weighted and summed to obtain feature vector_1; sub-feature vector_2, correlation feature vector_24, and correlation feature vector_12 are weighted and summed to obtain feature vector_2; and so on. Finally, the obtained feature vectors_1, feature vector_2, ..., feature vector_128 are arranged sequentially to obtain the spectral features.
[0157] As one embodiment, the spectral sub-features of each sub-segment represent at least one of the following: pitch, timbre, recording environment, rhythm, instrument, and emotion of the sub-segment. They may also represent at least one of the spectral features described above, without any specific limitation. Spectral correlation features represent at least one of the following: similarity, rhythmic relationship, and pitch transition between sub-segments. They may also represent at least one of the spectral features described above, without any specific limitation.
[0158] S2022 employs a feature mapping sub-model to map the obtained spectral features into descriptive text features; whereby the descriptive text features are used to describe various training music attributes of the sample music segments.
[0159] After obtaining the spectral features of the sample music fragment, a feature mapping sub-model can be used to map the obtained spectral features into descriptive text features. These descriptive text features are used to describe various training musical attributes of the sample music fragment. Converting spectral features into descriptive text features enables the understanding of the sample music fragment in textual form, allowing the music recognition model to learn the ability to describe a piece of music through textual descriptions.
[0160] The feature mapping sub-model maps spectral features to the text feature space, and aggregates the dense spectral features extracted by the feature extraction sub-model into several learnable text tokens. A text token is a learnable text unit that aggregates the dense spectral features extracted by the feature extraction sub-model. Compared with spectral features, it is more abstract and semantic, thus obtaining a higher level of semantic representation. At the same time, the amount of data is reduced, thus reducing the amount of computation.
[0161] In the field of unimodal data processing, such as law, medicine, coding, and mathematics, training a unimodal recognition model is equivalent to using a joint probability distribution P(X, Y) combined with text data from both general and specific domains. By training on general domain data, the model can better understand and generate text from specific domains. Here, X represents general domain text corpus, such as broad text data from the internet, books, news articles, etc.; Y represents specific domain text corpus, such as legal documents, case studies, and regulations in the legal field.
[0162] In the field of multimodal data processing, such as audio and video processing, training a multimodal recognition model is equivalent to using a conditional probability distribution P(X|Y) to train the model to transform information from one modality to another. For example, a multimodal recognition model can extract information from non-textual modal data and convert it into text. Here, X represents general text, i.e., ordinary natural language text; Y represents new modal data. New modal data refers to non-textual data that differs from general text, such as images and audio. In the field of multimodal data processing, when training a multimodal recognition model, the conditional probability distribution P(X|Y) enables the model to convert information from the new modal data into general text. P(X|Y) represents the conditional probability distribution that produces X given Y. General text refers to ordinary natural language text. In the field of multimodal data processing, when training a multimodal recognition model, the general text (X) and the new modal data (Y) are connected through the conditional probability distribution P(X|Y), enabling the model to extract information from non-textual modal data and convert it into text.
[0163] It can be seen that, in both unimodal and multimodal domains, the training of such recognition models focuses on learning basic abilities and new knowledge, and the amount of corpus used is usually large. Therefore, in this embodiment, the unimodal or multimodal recognition process of the model is not trained; the recognition process directly uses the text recognition sub-model already trained based on general text for data processing. In this embodiment, the training process is set on the feature mapping sub-model, allowing the feature mapping sub-model to learn the ability to map features from the music dimension to features in the text dimension. This simplifies the training process of the multimodal recognition model to the training process of the feature mapping sub-model. If the feature mapping sub-model is a linear transformation layer, it can be seen that this embodiment greatly reduces the number of model parameters that need to be trained and improves training efficiency.
[0164] Please refer to Figure 5A. Taking the feature mapping sub-model as a linear transformation layer as an example, the linear transformation layer can be represented as a weight matrix, such as a weight matrix of size (1024×4096). The values of each weight element in the weight matrix are the model parameters of the feature mapping sub-model. Through iterative training, the values of each weight element in the weight matrix can be adjusted so that the linear transformation layer can transform the spectral features into descriptive text features that can be accurately recognized by the trained text recognition model.
[0165] For example, if the size of the spectral feature is (128×1024), then matrix multiplication of the spectral feature and the weight matrix yields the descriptive text feature, which has a size of (128×4096). Through iterative training, the obtained descriptive text feature is recognized by the trained text recognition model as a training music description containing multiple training music attributes, which is very similar to the corresponding sample music description.
[0166] As one example, by continuously adjusting the model parameters of the feature mapping sub-model in each round of iterative training, the feature mapping sub-model has the ability to map features from the music dimension to features from the text dimension. In each round of iterative training, the model parameters of the feature mapping sub-model can represent a first mapping relationship between the spectral features of each attribute and the text features of each attribute that the feature mapping sub-model has learned. Thus, by continuously adjusting the model parameters of the feature mapping sub-model through multiple rounds of iterative training, the first mapping relationship is continuously adjusted, making the first mapping relationship more accurate, thereby making the music recognition model trained more accurate.
[0167] Therefore, when mapping the obtained spectral features to descriptive text features, we can actually obtain the first mapping relationship between the spectral features of each attribute and the textual features of each attribute learned by the feature mapping sub-model in this round of training, based on the model parameters of the feature mapping sub-model. Then, based on the current first mapping relationship, the spectral features are mapped to multiple attribute text features, which represent various training music attributes. Specifically, for each attribute spectral feature in the spectral features, the corresponding attribute text feature is found in the first mapping relationship, and these found attribute text features are combined to obtain multiple attribute text features. Thus, the obtained multiple attribute text features can be fused to generate descriptive text features. The specific fusion method can be a weighted sum of multiple attribute text features, with the weights set according to the importance or relevance of the attribute text features; or the multiple attribute text features can be arranged sequentially to form a text sequence as descriptive text features.
[0168] Each attribute text feature represents one of multiple reference music attributes; these multiple reference music attributes include at least the union of multiple sample music attributes and multiple training music attributes. These multiple reference music attributes can be all music attributes contained in the music domain. The model parameters of the feature mapping sub-model in this round of training learn a first mapping relationship between the spectral features of each attribute and the attribute text features of all music attributes contained in the music domain. Therefore, after obtaining the spectral features, the attribute text features corresponding to each of the multiple attribute spectral features included in the spectral features can be determined within the current first mapping relationship. The reference music attributes corresponding to each of the determined multiple attribute text features are denoted as multiple training music attributes.
[0169] Please refer to Figure 5B. During this round of iterative training, the model parameters of the feature mapping sub-model can be represented as the mapping relationship between attribute spectral feature_1 and attribute text feature_1, the mapping relationship between attribute spectral feature_2 and attribute text feature_2, ..., the mapping relationship between attribute spectral feature_N and attribute text feature_N. After the feature mapping sub-model obtains a spectral feature, it determines that the spectral feature includes attribute spectral feature_1, attribute spectral feature_4, and attribute spectral feature_12. Then, according to the first mapping relationship, the attribute text feature_1 corresponding to attribute spectral feature_1, the attribute text feature_4 corresponding to attribute spectral feature_4, and the attribute text feature_12 corresponding to attribute spectral feature_12 can be determined.
[0170] Then, attribute text features_1, attribute text features_4, and attribute text features_12 are fused to obtain descriptive text features. Fusion can be achieved by sequentially arranging attribute text features_1, attribute text features_4, and attribute text features_12, by performing a weighted sum of attribute text features_1, attribute text features_4, and attribute text features_12, or by concatenating attribute text features_1, attribute text features_4, and attribute text features_12 from the beginning and end, etc. There are no specific restrictions.
[0171] As one example, each iteration of training can process only one sample music segment, or it can process multiple sample music segments in a batch. When processing multiple sample music segments in a batch, each sample music segment can be processed in parallel on the same device, or it can be processed in a distributed manner on different devices, etc., without any specific limitations. For example, if a device is configured with 8 Graphics Processing Units (GPUs), and each GPU can process 4 sample music segments in parallel, then one device can process 32 sample music segments in parallel. If a batch of sample music segments contains 128 sample music segments, then it can be processed in a distributed manner on 4 devices to improve the efficiency of training the music recognition model.
[0172] As one example, during the training of a music recognition model, various data can be represented in multiple formats, such as single-precision floating-point numbers (FP32), half-precision floating-point numbers (FP16), and lower precision (e.g., BF16). Different data formats require different amounts of data; higher precision formats require more data. To further improve model training efficiency, it is advisable to convert the data format to a smaller format during some stages of the training process to accelerate processing and improve model training efficiency.
[0173] Before mapping the obtained spectral features to descriptive text features using a feature mapping sub-model, a target data format with a smaller data size than the reference data size can be selected from the data formats based on a preset second mapping relationship between each data format and its data size. For example, if the preset second mapping relationship between each data format and its data size is as follows: single-precision floating-point numbers (FP32) occupy 32 bits of data, half-precision floating-point numbers (FP16) occupy 16 bits of data, and lower precision (such as BF16) occupy 16 bits of data, then based on this second mapping relationship, a target data format with a smaller data size than the reference data size can be selected from the data formats.
[0174] The target data format refers to the data format selected from each data format with a corresponding data size smaller than the reference data size, based on the second mapping relationship between each preset data format and each data size. In the process of training the music recognition model, converting the spectral features into the target data format for processing can reduce the amount of data that needs to be processed in the feature mapping process and improve the model training efficiency.
[0175] The reference data volume is used as a reference standard to select the target data format in the second mapping relationship between the preset data formats and the data volume occupied. That is, the data format with a corresponding data volume occupied less than the reference data volume is selected as the target data format.
[0176] Specifically, the data formats and their corresponding data sizes in the second mapping relationship are traversed, and data formats with data sizes smaller than the reference data size are selected as candidates. If multiple candidate data formats exist, one can be randomly selected as the target data format; if there is only one candidate, it is determined as the target data format. The parameter data size is the data size corresponding to the data format used by the model parameters of the feature mapping sub-model. The spectral features are converted into the target data format to obtain the converted spectral features. Then, the feature mapping sub-model is used to map the converted spectral features into descriptive text features, thereby reducing the amount of data that needs to be processed in the feature mapping process and improving the model training efficiency.
[0177] After obtaining the descriptive text features, a scaling factor can be used to convert the data format of the descriptive text features from the target data format back to the data format corresponding to the reference data volume, thereby ensuring data accuracy during model training. The scaling factor is a coefficient used to scale the data. In this application, after converting the spectral features to the target data format for processing, a scaling factor can be used to convert the data format of the descriptive text features back to the data format corresponding to the reference data volume in order to ensure data accuracy during model training.
[0178] As one example, before converting the data format, the system can determine whether conversion is necessary based on the GPU memory usage during the current training iteration. If the GPU memory usage is low, it indicates that there are still many idle processing resources available, so data format conversion is unnecessary and model training can remain efficient. Conversely, if the GPU memory usage is high, it indicates that idle processing resources are limited, so data format conversion can be performed to improve model training efficiency. Therefore, by using GPU memory usage as a guide, the system can selectively choose whether to perform data format conversion, thus improving the flexibility of model training.
[0179] GPU memory usage refers to the GPU memory consumption during the current training iteration. By obtaining GPU memory usage, it's possible to determine whether data format conversion to a target format is necessary for processing, thus improving model training efficiency. When GPU memory usage exceeds a preset threshold, the target data format is selected for data processing; when GPU memory usage is not greater than the preset threshold, data processing can proceed directly. The preset threshold is a threshold used to determine whether data format conversion to a target format is necessary. When GPU memory usage exceeds the preset threshold, the target data format is selected from various data formats based on a second preset mapping relationship between each data format and its memory usage; when GPU memory usage is not greater than the preset threshold, data processing can proceed directly.
[0180] When training using deep learning frameworks (such as PyTorch or TensorFlow), you can utilize the framework's provided APIs to obtain GPU memory usage. For example, in PyTorch, you can use the `torch.cuda.memory_allocated()` function to get the currently allocated GPU memory size, and the `torch.cuda.max_memory_allocated()` function to get the maximum GPU memory allocation during the current training iteration. When using TensorFlow, you can use `tf.config.experimental.get_memory_info('GPU:0')` to obtain GPU memory usage information for a specific GPU.
[0181] After obtaining the GPU memory usage for this round of training iterations, if the GPU memory usage is greater than a preset limit, a target data format with a smaller data usage than a reference data size is selected from the data formats based on the second mapping relationship between each preset data format and each data usage size. If the GPU memory usage is not greater than the preset limit, the feature mapping sub-model can be directly used to map the obtained spectral features to descriptive text features.
[0182] S2023 employs a text recognition sub-model to identify descriptive text features and obtain training music descriptions.
[0183] After obtaining the descriptive text features, a trained text recognition sub-model can be used to identify the descriptive text features and obtain the training music description, which contains multiple training music attributes of the sample music fragments.
[0184] S2024, Based on the differences between the obtained training music description and the sample music description corresponding to the training music description, adjust the model parameters of the feature mapping sub-model.
[0185] After obtaining the training music description, the difference between the obtained training music description and the corresponding sample music description can be used to determine whether the training objective has been achieved. If the training objective has not been achieved, the model parameters of the feature mapping sub-model are adjusted; if the training objective has been achieved, the trained music recognition model can be output based on the current model parameters of the feature mapping sub-model.
[0186] The training objective refers to the criteria used to determine whether to stop training during the training of a music recognition model. For example, the difference between the obtained training music description and the corresponding sample music description is less than a preset difference value; or the difference between the obtained training music description and the corresponding sample music description is within a preset difference range, and the number of training iterations reaches a preset number. When the training objective is achieved, the trained music recognition model can be output based on the current model parameters of the feature mapping sub-model. The preset difference value is a criterion within the training objective. When the difference between the obtained training music description and the corresponding sample music description is less than the preset difference value, the training objective is considered achieved, and the trained music recognition model can be output based on the current model parameters of the feature mapping sub-model. The preset difference range is a criterion within the training objective. Combined with the number of training iterations, when the difference between the obtained training music description and the corresponding sample music description is within a preset difference range, and the number of training iterations reaches a preset number, the training objective is considered achieved, and the trained music recognition model can be output based on the current model parameters of the feature mapping sub-model.
[0187] As one embodiment, the obtained training music description contains multiple training music attributes. Therefore, the mapping training loss can be determined based on the differences between these attributes and the sample music attributes contained in the corresponding sample music description. Based on the mapping training loss, the model parameters of the feature mapping sub-model are adjusted. For example, if the mapping training loss does not reach the training objective, the model parameters of the feature mapping sub-model are adjusted; if the mapping training loss has reached the training objective, the trained music recognition model can be output based on the current model parameters of the feature mapping sub-model.
[0188] Therefore, as long as the training music description introduces the same training music attributes as the sample music description, or if the training music description introduces the same training music attributes as the sample music description, and the ratio of the number of each training music attribute to the number of multiple sample music attributes is not less than a preset ratio, it can be determined that the mapping training loss has achieved the training objective, regardless of whether other text in the statements introducing the training music attributes is the same. Thus, when determining the differences, it is not necessary to understand the semantics of the music description, reducing the difficulty of determining the differences and improving the efficiency of determining the differences.
[0189] The mapping training loss can be calculated using the cross-entropy loss function, please refer to formula (1), which measures the difference between the probability distribution of the music recognition model and the probability distribution of the sample data.
[0190] Where exp(x[class]) represents the probability that the training music description output by the music recognition model contains any sample music attribute, ∑ j exp(x[j]) represents the sum of probabilities of multiple sample music attributes in the sample music description in the training music description output by the music recognition model.
[0191] In other embodiments, the cross-entropy loss function can also refer to formula (2) to measure the difference between the probability distribution recognized by the music recognition model and the probability distribution of the sample data.
[0192] Where L represents the mapping training loss, n represents the number of music attributes in the sample, and x i y represents the unnormalized score of the music attribute of the i-th sample in the training music description output by the music recognition model. i This represents the true label of the music attribute of the i-th sample in the sample data (y is the true label if the attribute is included). i =1, otherwise y i =0).
[0193] Please refer to Figure 5C. The sample music description is: "This music blends local characteristics, cleverly combining traditional and modern elements to showcase a unique artistic charm. Its BPM (beats per minute) is level 2, with a moderate tempo, making it perfect for dancing. The rhythmic quality is level 1, giving the entire piece a vibrant feel. The modern pop music elements add a fashionable touch. At the same time, the incorporation of opera elements imbues the entire piece with a rich local flavor. Furthermore, being sung in Cantonese provides listeners with a unique linguistic experience, making the song more regionally distinctive and showcasing the profoundness of Chinese culture."
[0194] The training music is described as follows: "This music falls into the BPM (beats per minute) range of 2 and rhythmic intensity of 1, possessing a strong sense of rhythm that can make listeners involuntarily dance along. The song belongs to the Cpop genre, blending modern pop elements with traditional musical elements to showcase a unique musical style. Furthermore, being sung in Cantonese gives it a distinctive musical style, combining pop music elements with characteristics of traditional music, making the song even more unique."
[0195] The sample music attributes include "traditional elements", "modern elements", "local characteristics", "BPM (beats per minute) level 2", "rhythm level 1", "pop elements", "opera elements" and "Cantonese", etc.
[0196] The training music attributes include "BPM (beats per minute) level 2", "rhythm level 1", "pop style", "Cpop style", "traditional elements", "modern elements" and "Cantonese", etc.
[0197] Therefore, the mapping training loss can be determined based on the differences between the various training music attributes contained in the obtained training music description and the various sample music attributes contained in the sample music description corresponding to the obtained training music description.
[0198] As one embodiment, to further improve training accuracy, a semantic training loss can be determined based on the obtained training music description and the text similarity between the obtained training music description and the corresponding sample music description, in addition to the determined mapping training loss. Based on the mapping training loss and the semantic training loss, if the training objective is not achieved, the model parameters of the feature mapping sub-model are adjusted; if the training objective is achieved, the trained music recognition model is output based on the current model parameters of the feature mapping sub-model.
[0199] The semantic training loss can also be calculated using the formula (1) above, where exp(x[class]) represents the probability that the training music description output by the music recognition model contains any word group from the sample music description, ∑ j exp(x[j]) represents the sum of probabilities of each word group in the sample music description in the training music description output by the music recognition model.
[0200] Semantic training loss can also be determined by calculating text similarity between the obtained text features of the training music description and the text features of the corresponding sample music description. The smaller the distance, the greater the text similarity, and the smaller the semantic training loss; conversely, the greater the distance, the smaller the text similarity, and the greater the semantic training loss.
[0201] In some embodiments, the total loss L is calculated by weighted summation based on the mapping training loss and the semantic training loss. total =αL map +βL sem L total This represents the total loss, where α and β are preset weighting coefficients, and α + β = 1. map The training loss for mapping can be calculated using the cross-entropy loss function. L sem The semantic training loss is represented by the cosine similarity used to calculate text similarity, which is then used to determine the loss. The calculation formula is as follows: in This represents the text feature vector used to train the music description. This represents the text feature vector describing the sample music. The total loss is calculated by weighted summation of the mapping training loss and the semantic training loss. It is used to determine whether the training objective has been achieved. When the total loss exceeds a preset loss threshold, it is determined that the training objective has not been achieved, and the model parameters of the feature mapping sub-model are adjusted. When the total loss L... total If the loss exceeds the preset threshold, it is determined that the training objective has not been achieved, and the model parameters of the feature mapping sub-model are adjusted.
[0202] As one example, if multiple sample data are processed in the same batch, when determining whether to adjust the model parameters of the feature mapping sub-model, the number of sample data that has not reached the training objective can be determined. If the ratio of the number of sample data that has not reached the training objective to the total number of sample data reaches a preset ratio, it indicates that a large number of sample data has not reached the training objective, and the model parameters of the feature mapping sub-model should be adjusted. If the preset ratio is not reached, it indicates that a large number of sample data has reached the training objective, and the trained music recognition model can be output. This provides a macro-level evaluation of the music recognition model's learned recognition ability, avoiding overfitting and other issues. The preset ratio is a standard used to determine whether to adjust the model parameters of the feature mapping sub-model when processing multiple sample data in the same batch. When the ratio of the number of sample data that has not reached the training objective to the total number of sample data reaches a preset ratio, it indicates that a large number of sample data has not reached the training objective, and the model parameters of the feature mapping sub-model should be adjusted. If the preset ratio is not reached, it indicates that a large number of sample data has reached the training objective, and the trained music recognition model can be output.
[0203] Please refer to Figure 5D. When processing two sample data in the same batch, for example, the first sample music description is: "This music blends local characteristics, cleverly combining traditional and modern elements to showcase a unique artistic charm. Its BPM (beats per minute) is level 2, with a moderate rhythm, making it very suitable for dancing. The rhythmic feel is level 1, making the whole song full of energy. The modern pop music elements make the song more fashionable. At the same time, the incorporation of opera elements gives the whole song a rich local flavor. In addition, being sung in Cantonese can bring listeners a unique linguistic beauty, making the song more regional and showcasing the profoundness of Chinese culture."
[0204] The first training music description states: "This music falls into the BPM (beats per minute) range of 2 and rhythmic intensity range of 1, possessing a strong sense of rhythm that can make listeners involuntarily dance along. The song belongs to the pop-Cpop style, blending modern pop elements with traditional musical elements, showcasing a unique musical style. Furthermore, being sung in Cantonese gives it a distinctive musical style, not only incorporating pop music elements but also integrating characteristics of traditional music, making the song even more unique."
[0205] The second sample music is described as follows: "This music belongs to the electronic music category, with a particular emphasis on its rhythm and dynamism. A BPM of 4 indicates a fast tempo, suitable for energizing and exciting music. As a type of electronic music, it is usually categorized as House Music or EDM (Electronic Dance Music)." This music, known for its strong rhythm and beat, is designed to energize and excite. It's described as "energy-boosting," indicating its powerful rhythm and vibrant energy—the kind of music that gets your heart racing and your blood pumping, perfect for promoting activities or sports, or creating a lively atmosphere at parties and social events. Furthermore, the "instrumental" tag means the music is lyrical, focusing on its expressive power and emotional delivery. Without lyrics, the music can directly engage the listener's senses and emotions. Finally, the "5-level rhythm" tag emphasizes its strong rhythm and beat, ideal for dancing and exercise, making it irresistibly catchy. Overall, this is a vibrant, rhythmic, and energetic electronic dance track, suitable for igniting energy and enthusiasm, creating a lively atmosphere, and can also be enjoyed purely as instrumental music to appreciate its inherent charm.
[0206] The second training music is described as follows: "This music is at BPM 4. It's a remix version with a strong rhythm and beat. It belongs to the Western music category, blending elements of Western music and electronic music to create a unique musical style. This song belongs to the electronic music and DJ dance music category, with a strong rhythm and beat, making it perfect for dancing. With its crisp and energetic rhythmic sound effects, this DJ dance music is a very passionate and vibrant style. It's full of various sonic characteristics, including a variety of unique sonic elements, making the whole music more lively and interesting. This musical style is perfect for parties, dancing, and other social events."
[0207] The training objective can be determined by first comparing the differences between the first sample music description and the first training music description. For example, the differences between the first sample music description ("This music incorporates local characteristics," "The rhythm is moderate," "It makes the whole piece energetic," "It makes the piece more fashionable," "It incorporates opera elements, making the whole piece full of rich local flavor," and "It makes the song more regional and showcases the profoundness of Chinese culture") and the differences between the first training music description ("It has a strong sense of rhythm" and "The song belongs to the pop Cpop style").
[0208] The training objective was determined based on the differences between the second sample music description and the second training music description. For example, the differences included in the second sample music description: "It is usually categorized as house or EDM," "This music is described as 'energetic and energetic,'" "It's music that makes your heart race and your blood boil," "The 'pure music' label means that this music has no lyrics, focusing on the expressiveness and emotional delivery of the music itself. Without the interference of lyrics, the music can more directly touch the listener's senses and emotions," and "It can also be appreciated as pure music, experiencing its inherent charm." The differences also included in the second training music description: "Remix version," "DJ dance music," "DJ dance music with crisp, rhythmic sound effects," and "It is full of various sonic characteristics, including a variety of unique sonic elements."
[0209] If one of the two sample data does not reach the training objective, adjust the model parameters of the feature mapping sub-model until both sample data reach the training objective, and output the trained music recognition model.
[0210] The training method of the music recognition model provided in the embodiments of this application will be described below as an example.
[0211] Please refer to Figure 6. The music recognition model includes a trained feature extraction sub-model, denoted as Audio encoder; a feature mapping sub-model to be trained, denoted as Adapter; and a trained text recognition sub-model, denoted as LLM.
[0212] The feature extraction sub-model has a 24-layer transformer structure, used to encode the input 30-second sample music clip to extract its spectral features, resulting in a spectral feature output of size (128, 1024). The model parameters of the feature extraction sub-model were obtained through unsupervised training using 170 million collected 30-second audio data points.
[0213] Sample music clips can be obtained by performing a log-mel transform on a sub-file of a music clip. This transforms the original audio signal into a log-Mel spectrogram, a two-dimensional array where each column represents a time point and each row represents a Mel frequency. The log-transformed Mel spectrogram is a log-Mel spectrogram, and the Mel frequency at each time point can be represented by an audio token. Therefore, sample music clips can be represented using multiple audio tokens.
[0214] The feature mapping sub-model has a linear transformation layer that maps the spectral features output by the feature extraction sub-model to the text feature space that the text recognition sub-model can recognize. For example, the linear transformation layer can be represented as a weight matrix of size (1024, 4096). Multiplying the spectral features by the weight matrix yields a descriptive text feature of size (128, 4096). This maps the dense spectral features into 128 more abstract and semantically richer text tokens.
[0215] The text recognition sub-model uses the open-source 7B Chinese large model internlm2 to recognize the features of the input descriptive text as training music descriptions.
[0216] Therefore, the model parameters of the feature extraction sub-model and the text recognition sub-model can be fixed. Based on the difference between the training music description and the sample music description corresponding to the training music description, the model parameters of the feature mapping sub-model can be adjusted until the training target is reached, thus obtaining the trained feature mapping sub-model and the trained music recognition model.
[0217] In this embodiment, by mining music tags, genres, instruments, pitch, metadata, etc., a massive sample music segment-sample music description pairs are constructed as a sample dataset. Multiple music attributes are fused, and an open-source Chinese large-scale model is used to generate music descriptions for these attributes. This can accumulate tens of millions of sample data points. Furthermore, no manual annotation is required, significantly reducing data annotation costs.
[0218] Music tags are identifiers used to mark the features, attributes, or categories of music. They can describe music from multiple dimensions, such as genre (pop, rock, classical, etc.), theme (love, inspirational, homesickness, etc.), rhythm (fast, slow, etc.), instrument used (piano, guitar, violin, etc.), and singing style (bel canto, popular, jazz, etc.). In this application, music tags can serve as part of the information used to construct the sample dataset, assisting in generating sample music descriptions and helping the music recognition model learn and recognize various attributes of music.
[0219] Pitch refers to the highness or lowness of a sound, and it depends on the frequency of the vibration of the sound-producing body. The higher the frequency, the higher the pitch; the lower the frequency, the lower the pitch. In music, pitch is one of the important elements in constructing melody and harmony, and different combinations of pitches can create different musical effects and emotional expressions. In this application, pitch can be used as part of the attributes of sample music to describe the characteristics of sample music segments, helping music recognition models to more accurately identify and describe music.
[0220] Furthermore, through large-scale pre-training by aligning spectral features with large-scale model text features, robust and highly accurate descriptive text features were obtained. The model using large-scale pre-training achieved better results than the single-modality model in multiple downstream tasks, representing a significant improvement.
[0221] Furthermore, by using a large model as a pre-trained text recognition model without training a large model, the music recognition model achieves better accuracy and generalization ability without increasing the training difficulty, and can be applied to a wider range of fields.
[0222] After obtaining the trained music recognition model using the above training method, the trained music recognition model can be used for music recognition. The music recognition method provided in this application embodiment will be further described below based on Figure 1C.
[0223] Please refer to Figure 7A, which is a flowchart of a music recognition method provided in an embodiment of this application.
[0224] S701, acquire the music to be identified.
[0225] There are several ways to acquire the music to be identified. For example, the music can be acquired through a sound acquisition device; or it can be received from other devices; or it can be generated based on received note playing instructions. There are no specific limitations.
[0226] S702 uses a trained music recognition model to extract the spectral features of the music to be recognized and maps these spectral features to the text features to be recognized.
[0227] The music recognition model is trained using the method described above. Therefore, the trained music recognition model includes a trained feature extraction sub-model, a trained feature mapping sub-model, and a trained text recognition sub-model. When the music to be recognized is input into the music recognition model, the feature extraction sub-model extracts features from the music to obtain the spectral features to be recognized. The feature mapping sub-model, for example, uses a linear transformation to map the spectral features to be recognized into text features to be recognized.
[0228] S703 uses a music recognition model to identify the features of the text to be recognized and obtain a description of the target music.
[0229] The text recognition sub-model in the music recognition model can identify the features of the text to be recognized as a target music description. The target music description is used to introduce various target music attributes of the music to be recognized in text form.
[0230] Referring to Figure 7B, in a music application running on a terminal device, when the music recognition button displayed on the application's interface is clicked, the music application invokes the music recognition method provided in this embodiment. The method continuously collects the music to be recognized using a recording device on the terminal device. If the collected music is less than 30 seconds long, it is padded to a 30-second length and input into the trained music recognition model. If the collected music is longer than 30 seconds, 30 seconds of that time is selected as the music to be recognized and input into the trained music recognition model. This allows the acquisition of the target music description output by the trained music recognition model.
[0231] When the collected music exceeds 30 seconds, 30 seconds can be directly selected as the input model for the music to be identified. This method is suitable for situations where there is a preliminary understanding of the overall characteristics of the music and the recognition result is desired quickly. On the other hand, the method of collecting and generating the input model for the music to be identified every 10 seconds for the same piece of music is more suitable for scenarios where detailed analysis of the music is required to obtain a more accurate description of the target music.
[0232] As one example, when inputting into a trained music recognition model, for the same piece of music, a new piece of music to be recognized can be generated every 10 seconds of music collected. Each time a piece of music to be recognized is generated, the collected music to be recognized is input into the trained music recognition model, so that the trained music recognition model can combine the obtained multiple text features to be recognized to generate a more accurate description of the target music.
[0233] Based on the same inventive concept, this application provides a training device for a music recognition model, capable of implementing the functions corresponding to the aforementioned training method for a music recognition model. Referring to Figure 8A, the device includes an acquisition module 81 and a processing module 82, wherein:
[0234] Acquisition Module 81: Used to acquire sample datasets; wherein each sample dataset includes: a sample music segment and a sample music description corresponding to the sample music segment, the sample music description containing various sample music attributes of the sample music segment;
[0235] Processing module 82: Used for multi-round iterative training of the music recognition model to be trained based on the sample dataset; wherein, the music recognition model includes: a trained feature extraction sub-model and a text recognition sub-model, and a feature mapping sub-model to be trained; each round of iterative training includes:
[0236] Processing module 82 is specifically used to: extract the spectral features of the sample music segments using a feature extraction sub-model;
[0237] The processing module 82 is specifically used to: use a feature mapping sub-model to map the obtained spectral features into descriptive text features; wherein, the descriptive text features are used to describe various training music attributes of the sample music fragments;
[0238] Processing module 82 is specifically used for: employing a text recognition sub-model to identify descriptive text features and obtain training music descriptions; and
[0239] The processing module 82 is specifically used to: adjust the model parameters of the feature mapping sub-model based on the differences between the obtained training music description and the sample music description corresponding to the training music description.
[0240] In one possible embodiment, the acquisition module 81 is specifically used for:
[0241] Collect the multimedia files of multiple music tracks, and collect the metadata of each multimedia file;
[0242] Based on the preset segment length, each multimedia file is divided into segments to obtain multiple sub-files corresponding to the multimedia file, and each obtained sub-file is used as a sample music segment.
[0243] A trained multimodal recognition model is used to identify the music attributes of each obtained sample music segment, thereby obtaining multiple sample music attributes for each sample music segment.
[0244] Using a trained large language model, based on the obtained metadata and the attributes of each sample music, a sample music description is generated for each sample music segment.
[0245] A sample dataset is obtained based on each sample music segment and its corresponding sample music description.
[0246] In one possible embodiment, the processing module 82 is further configured to:
[0247] Before conducting multiple rounds of iterative training on the music recognition model to be trained based on the sample dataset, a feature extraction sub-model and a feature mapping sub-model to be trained are established based on the structure building strategy of the music recognition model, as well as the already trained text recognition sub-model is obtained.
[0248] Based on the collected audio data, the feature extraction sub-model to be trained is iterated through multiple rounds, and the trained feature extraction sub-model is output.
[0249] Based on the trained feature extraction sub-model and text recognition sub-model, as well as the feature mapping sub-model to be trained, a music recognition model to be trained is established.
[0250] In one possible embodiment, the processing module 82 is specifically used for:
[0251] In each training iteration, perform the following operations:
[0252] Sub-data masking is performed on the audio data to obtain masked data that masks random sub-data.
[0253] Feature extraction is performed on the masked data to obtain data features;
[0254] Based on data characteristics, predict the random sub-data that is hidden in the masked data to obtain the predicted sub-data;
[0255] Based on the differences between the obtained predicted sub-data and random sub-data, the model parameters of the feature extraction sub-model are adjusted.
[0256] In one possible embodiment, the processing module 82 is specifically used for:
[0257] A feature extraction sub-model is used to extract the spectral sub-features of each sub-segment in the sample music segment, and to extract the spectral correlation features between the spectral sub-features.
[0258] By fusing the obtained spectral sub-features and spectral correlation features, the spectral features presented by the sample music fragments are obtained.
[0259] In one possible embodiment, the spectral sub-features of each sub-segment represent at least one of the following: pitch, timbre, recording environment, beat, instrument, and emotion of the sub-segment; the spectral association features represent at least one of the following: similarity, rhythmic relationship, and pitch transition between sub-segments.
[0260] In one possible embodiment, the processing module 82 is specifically used for:
[0261] Based on the model parameters of the feature mapping sub-model in this round of training, the first mapping relationship between the spectral features of each attribute and the text features of each attribute learned by the feature mapping sub-model is obtained; wherein, each attribute text feature represents one of multiple reference music attributes; the multiple reference music attributes include at least: the union of multiple sample music attributes and multiple training music attributes;
[0262] Based on the first mapping relationship, the spectral features are mapped to multiple attribute text features; wherein, the multiple attribute text features represent multiple training music attributes;
[0263] Multiple attribute text features obtained are fused to generate descriptive text features.
[0264] In one possible embodiment, the processing module 82 is further configured to:
[0265] Before mapping the obtained spectral features to descriptive text features using the feature mapping sub-model, a target data format with a smaller data size than the reference data size is selected from each data format based on the second mapping relationship between each preset data format and each data size; where the parameter data size is: the data size corresponding to the data format used by the model parameters of the feature mapping sub-model;
[0266] The spectral features are converted into the target data format to obtain the converted spectral features.
[0267] In one possible embodiment, the processing module 82 is specifically used for:
[0268] Get the GPU memory usage for running this round of training iterations;
[0269] When it is determined that the video memory usage is greater than the preset usage, based on the second mapping relationship between each preset data format and each occupied data amount, the target data format with a corresponding occupied data amount less than the reference data amount is selected from each data format.
[0270] In one possible embodiment, the processing module 82 is specifically used for:
[0271] Based on the differences between the various training music attributes contained in the obtained training music description and the various sample music attributes contained in the sample music description corresponding to the obtained training music description, the mapping training loss is determined.
[0272] The model parameters of the feature mapping sub-model are adjusted based on the mapping training loss.
[0273] In one possible embodiment, the processing module 82 is specifically used for:
[0274] Based on the text similarity between the obtained training music description and the corresponding sample music description, the semantic training loss is determined.
[0275] Based on the mapping training loss and semantic training loss, when the training objective is not achieved, the model parameters of the feature mapping sub-model are adjusted.
[0276] Based on the same inventive concept, this application provides a music recognition device capable of realizing the functions corresponding to the aforementioned music recognition method. Referring to Figure 8B, the device includes an acquisition module 801 and a processing module 802, wherein:
[0277] Acquisition module 801: Used to acquire the music to be identified;
[0278] Processing module 802: used to extract the spectral features of the music to be identified by using a trained music recognition model, and map the spectral features to be identified to the text features to be identified; wherein, the music recognition model is trained using the method described in the first aspect;
[0279] The processing module 802 is also used to: use a music recognition model to identify the features of the text to be identified and obtain a target music description; wherein, the target music description is used to: introduce the various target music attributes of the music to be identified in text form.
[0280] Please refer to Figure 9, which illustrates a computer device 900 provided in an embodiment of this application. This computer device 900 can be, for example, the client 101 or server 102 shown in Figure 1C. The computer device 900 includes a processor 980 and a memory 920. In some embodiments, the computer device 900 may include a display unit 940, which includes a display panel 941 for displaying a user-interactive interface, etc.
[0281] In one possible embodiment, the display panel 941 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0282] The processor 980 is used to read a computer program and then execute the methods defined by the computer program. For example, the processor 980 reads a data storage program or file, thereby running the data storage program on the computer device 900 and displaying the corresponding interface on the display unit 940. The processor 980 may include one or more general-purpose processors, and may also include one or more DSPs (Digital Signal Processors) for performing related operations to implement the technical solutions provided in the embodiments of this application. The current and previous versions of the data storage program and the application software corresponding to the data storage program can be installed on the computer device 900.
[0283] The memory 920 generally includes main memory and secondary storage. Main memory can be random access memory (RAM), read-only memory (ROM), and cache, etc. Secondary storage can be a hard disk, optical disk, USB flash drive, floppy disk, or magnetic tape drive, etc. The memory 920 is used to store computer programs and other data. The computer programs include applications corresponding to each client, and other data may include data generated after the operating system or applications are run, including system data (e.g., operating system configuration parameters) and user data. In this embodiment, the computer program is stored in the memory 920, and the processor 980 executes the computer program in the memory 920 to implement any of the methods described in the preceding figures.
[0284] The aforementioned display unit 940 is used to receive input digital information, character information, or contact touch operations / non-contact gestures, and to generate signal inputs related to user settings and function control of the computer device 900. Specifically, in this embodiment, the display unit 940 may include a display panel 941. The display panel 941, for example, is a touch screen, which can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or on the display panel 941), and drive corresponding connection devices according to a pre-set program.
[0285] In one possible embodiment, the display panel 941 may include two parts: a touch detection device and a touch controller. The touch detection device detects the player's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 980. It can also receive and execute commands from the processor 980.
[0286] The display panel 941 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 940, in some embodiments, the computer device 900 may also include an input unit 930. The input unit 930 may include an image input device 931 and other input devices 932, wherein the other input devices may include, but are not limited to, one or more of the following: a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick.
[0287] In addition to the above, the computer device 900 may also include a power supply 990 for powering other modules, an audio circuit 960, a near-field communication module 970, and an RF circuit 910. The computer device 900 may also include one or more sensors 950, such as an accelerometer, a light sensor, and a pressure sensor. The audio circuit 960 specifically includes a speaker 961 and a microphone 962, for example, the computer device 900 can use the microphone 962 to collect the user's voice and perform corresponding operations.
[0288] As one embodiment, the number of processors 980 can be one or more, and the processors 980 and the memory 920 can be coupled together or relatively independent.
[0289] As one embodiment, the processor 980 in FIG9 can be used to implement the functions of the acquisition module 81 and the processing module 82 in FIG8A; it can also be used to implement the functions of the acquisition module 801 and the processing module 802 in FIG8B.
[0290] As one embodiment, the processor 980 in Figure 9 can be used to implement the functions corresponding to the server or terminal device discussed above.
[0291] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the computer program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0292] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of software products, for example, through computer program products. These computer program products are stored in a storage medium and include computer programs used to cause a computer device to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0293] In summary, this application provides a training method for a music recognition model, a music recognition method, a training device for a music recognition model, a music recognition device, a computer program product, a computer device, and a computer-readable storage medium. The computer device acquires a sample dataset, where each sample data contains a sample music fragment and a corresponding sample music description, the sample music description covering multiple sample music attributes. Based on this sample dataset, a music recognition model containing a trained feature extraction sub-model, a text recognition sub-model, and a feature mapping sub-model to be trained is subjected to multiple rounds of iterative training. In each round of iterative training, the feature extraction sub-model first extracts the spectral features of the sample music fragment. The spectral features contain key information such as frequency components, timbre, and rhythm in the music, and are an important manifestation of music features. Next, the feature mapping sub-model maps the spectral features to descriptive text features. These descriptive text features are used to describe multiple training music attributes, realizing the conversion from music modality to text modality. Then, the text recognition sub-model recognizes the descriptive text features to obtain the training music description. Finally, the parameters of the feature mapping sub-model are adjusted based on the difference between the training music description and the sample music description.
[0294] From a technical perspective, using a pre-trained text recognition sub-model to build the structure avoids training the large number of model parameters and the large model itself, such as the text recognition sub-model, in multimodal recognition models. Multimodal recognition models have numerous parameters, making training difficult and time-consuming. This method only requires training the feature mapping sub-model, significantly reducing training difficulty and time, and improving training efficiency. Simultaneously, since there is no need to train the large language model of the text recognition sub-model, it avoids the problem of low training and recognition accuracy in music recognition models due to insufficient sample data or overfitting. Furthermore, training only the feature mapping sub-model reduces the amount of model parameter data that needs to be trained in the music recognition model. This allows for training a pre-trained music recognition model with high recognition accuracy without requiring excessive sample data, reducing the difficulty of obtaining sample data, ensuring the recognition accuracy of the trained model, and improving resource utilization.
[0295] Furthermore, when acquiring the sample dataset, multiple music multimedia files and metadata are collected. The multimedia files are then segmented based on preset segment durations to obtain sample music segments. A trained multimodal recognition model is used to identify various sample music attributes of the sample music segments. A trained large language model is then used to generate sample music descriptions based on metadata and sample music attributes, thus obtaining the sample dataset. This process eliminates the need for manual annotation of sample data, reducing labor and time costs. The use of open-source models in the construction process makes the process more efficient and convenient, enabling the creation of datasets with a large number of samples according to actual usage needs. This improves the efficiency and applicability of the sample dataset construction, providing richer and more accurate data support for subsequent model training.
[0296] Before conducting multiple rounds of iterative training on the music recognition model to be trained based on the sample dataset, a feature extraction sub-model and a feature mapping sub-model are established based on the structure building strategy of the music recognition model, and a trained text recognition sub-model is obtained. Then, based on multiple collected audio data, the feature extraction sub-model to be trained is subjected to multiple rounds of iterative training to obtain a trained feature extraction sub-model, and thus the music recognition model to be trained is established. In this way, using music collected from the sample dataset or similar music to train the feature extraction sub-model enables the feature extraction sub-model to extract the features of each sample data in the sample dataset more accurately. Because the accuracy of the feature extraction sub-model directly affects the effect of subsequent feature mapping and text recognition, improving the accuracy of the feature extraction sub-model can improve the accuracy of training the feature mapping sub-model, resulting in a higher mapping accuracy of the trained feature mapping sub-model, thereby improving the performance of the entire music recognition model.
[0297] When training a feature extraction sub-model based on multiple collected audio datasets through iterative rounds, sub-data masking is performed on the audio data in each iteration to obtain masked data that hides random sub-data. Data masking is an important technique that enables audio data to be processed in a uniform format, avoiding data processing errors caused by complex and diverse data formats. Feature extraction is performed on the masked data to obtain data features, and then the masked random sub-data is predicted based on these features to obtain predicted sub-data. Finally, the parameters of the feature extraction sub-model are adjusted based on the difference between the predicted sub-data and the random sub-data. This unsupervised training method does not require audio data annotation, reducing the workload and cost of data annotation. By continuously adjusting the parameters, the feature extraction sub-model can more accurately learn the ability to extract data features from audio data, ultimately resulting in a trained feature extraction sub-model with accurate feature extraction capabilities, improving the accuracy and stability of feature extraction.
[0298] When extracting spectral features from sample music segments using a feature extraction sub-model, the spectral sub-features of each of the multiple sub-segments within the sample music segment are extracted, along with the spectral correlation features between these sub-features. These features are then fused to obtain the overall spectral features of the sample music segment. From a microscopic perspective, spectral sub-features can reflect detailed information such as pitch, timbre, recording environment, rhythm, instruments, and emotion of each sub-segment. From a macroscopic perspective, spectral correlation features can reflect information such as similarity, rhythmic relationships, and pitch transitions between sub-segments. By fusing these two types of features, the spectral features of the sample music segment can be extracted more comprehensively and accurately, ensuring the accuracy of feature extraction. This provides richer and more accurate feature information for subsequent feature mapping and text recognition, contributing to improved performance of the music recognition model.
[0299] Each sub-segment's spectral sub-features represent at least one of the following: pitch, timbre, recording environment, rhythm, instrument, and emotion. Spectral correlation features represent at least one of the following: similarity, rhythmic relationship, and pitch transition between sub-segments. This meticulous feature representation method allows the extracted features to more accurately reflect the characteristics of the sample music segments. During music recognition, this rich feature information helps the model more accurately understand the connotation and characteristics of music, thereby improving the music recognition model's ability to understand and describe music, and enhancing the model's recognition accuracy and reliability.
[0300] When using a feature mapping sub-model to map the obtained spectral features to descriptive text features, the first mapping relationship between the spectral features and text features of each attribute learned by the feature mapping sub-model in this round of training is obtained based on the model parameters of the feature mapping sub-model. Each attribute text feature represents one of multiple reference music attributes, which includes at least the union of multiple sample music attributes and multiple training music attributes. Based on the first mapping relationship, the spectral features are mapped to multiple attribute text features, which represent multiple training music attributes. These attribute text features are then fused to generate descriptive text features. By continuously adjusting the model parameters of the feature mapping sub-model through multiple rounds of iterative training, the first mapping relationship can be continuously adjusted to make it more accurate. An accurate first mapping relationship enables the trained music recognition model to more accurately map spectral features to descriptive text features, achieving effective alignment between the feature space of the music modality and the feature space of the text modality, thereby improving the recognition accuracy of the music recognition model.
[0301] Before mapping the obtained spectral features to descriptive text features using a feature mapping sub-model, a target data format with a smaller data size than the reference data size is selected based on a pre-defined second mapping relationship between each data format and its data size. Different data formats have different data sizes; higher-precision data formats require more data. Selecting a target data format with a smaller data size reduces the amount of data that needs to be processed during feature mapping. Converting the spectral features to the target data format for processing speeds up the process and improves model training efficiency. After obtaining the descriptive text features, a scaling factor is used to convert the data format of the descriptive text features back from the target data format to the data format corresponding to the reference data size, ensuring data accuracy during model training. This data format conversion method improves training efficiency while ensuring data accuracy, achieving a balance between efficiency and accuracy.
[0302] When selecting a target data format based on a pre-defined second mapping relationship between various data formats and their respective data volumes, the GPU memory usage for the current training iteration is obtained. If the GPU memory usage exceeds the pre-defined limit, it indicates that the available processing resources are limited. In this case, selecting the target data format based on the second mapping relationship reduces GPU memory usage and improves model training efficiency. If the GPU memory usage is less than the pre-defined limit, it means there are still many available processing resources. In this case, data processing can be performed directly, avoiding unnecessary data format conversions. This method of selectively choosing whether to perform data format conversions based on GPU memory usage improves the flexibility of model training, optimizes the utilization of training resources, and avoids resource waste.
[0303] When adjusting the parameters of the feature mapping sub-model based on the differences between the obtained training music descriptions and the corresponding sample music descriptions, the mapping training loss is determined based on the differences between the various training music attributes contained in the training music descriptions and the various sample music attributes contained in the sample music descriptions. The parameters of the feature mapping sub-model are then adjusted based on this mapping training loss. This approach does not require understanding the semantics of the music descriptions when determining differences; it only focuses on the differences between music attributes. This reduces the difficulty of determining differences, avoids complex semantic analysis processes, improves the efficiency of difference determination, and accelerates model training. Furthermore, as long as the training music attributes described in the training music descriptions and the sample music attributes described in the sample music descriptions meet certain conditions, it can be determined that the mapping training loss has reached the training objective. This judgment method is more flexible and efficient.
[0304] When adjusting the parameters of the feature mapping sub-model based on the mapping training loss, the semantic training loss is determined based on the textual similarity between the obtained training music descriptions and the corresponding sample music descriptions. Based on both the mapping and semantic training losses, the parameters of the feature mapping sub-model are adjusted if the training objective is not achieved. By comprehensively considering both the mapping and semantic training losses, the training effect of the model can be evaluated more comprehensively. The mapping training loss mainly focuses on the differences between music attributes, while the semantic training loss focuses on the textual similarity of music descriptions. By combining these two losses, the model can learn both music attributes and semantic information of music descriptions, further improving the training accuracy and performance of the music recognition model.
[0305] In music recognition methods, the music to be recognized is acquired, and a trained music recognition model, obtained through the aforementioned training method, is used to extract the spectral features of the music and map them into text features. The model then uses these text features to obtain a description of the target music, which is used to introduce various target music attributes in text form. Because the trained music recognition model has high recognition accuracy and can accurately extract the features of the music and convert them into text descriptions, it can more accurately identify the music and provide users with more accurate and detailed music descriptions. This improves the quality of music recognition and user experience, enabling users to better understand and appreciate music.
[0306] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0307] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for training a music recognition model, executed by a computer device, comprising: Obtain a sample dataset; wherein each sample dataset includes: a sample music segment and a sample music description corresponding to the sample music segment, the sample music description containing multiple sample music attributes of the sample music segment; Based on the aforementioned sample dataset, the music recognition model to be trained undergoes multiple rounds of iterative training; wherein, the music recognition model includes: a trained feature extraction sub-model and a text recognition sub-model, and a feature mapping sub-model to be trained; each round of iterative training includes: The feature extraction sub-model is used to extract the spectral features of the sample music fragments; The obtained spectral features are mapped to descriptive text features using the feature mapping sub-model; wherein the descriptive text features are used to describe various training music attributes of the sample music fragment; Using the text recognition sub-model, the features of the descriptive text are recognized to obtain the training music description; and The model parameters of the feature mapping sub-model are adjusted based on the differences between the obtained training music description and the sample music description corresponding to the training music description.
2. The method according to claim 1, wherein obtaining the sample dataset comprises: Collect the multimedia files of multiple music tracks, and collect the metadata of each multimedia file; Based on the preset segment length, each multimedia file is divided into segments, and multiple sample music segments are obtained for each multimedia file. A trained multimodal recognition model is used to identify the music attributes of each obtained sample music segment, thereby obtaining multiple sample music attributes for each sample music segment. Using a trained large language model, based on the obtained metadata and sample music attributes, sample music descriptions are generated for each sample music segment. Based on each sample music segment and its corresponding sample music description, a sample dataset is obtained.
3. The method according to claim 1 or 2, further comprising, before performing multiple rounds of iterative training on the music recognition model to be trained based on the sample dataset: Based on the structure building strategy of the music recognition model, a feature extraction sub-model and a feature mapping sub-model to be trained are established, and a trained text recognition sub-model is obtained. Based on the collected audio data, the feature extraction sub-model to be trained is iterated through multiple rounds, and the trained feature extraction sub-model is output. Based on the trained feature extraction sub-model and text recognition sub-model, as well as the feature mapping sub-model to be trained, a music recognition model to be trained is established.
4. The method according to claim 3, wherein the step of performing multiple rounds of iterative training on the feature extraction sub-model to be trained based on the collected multiple audio data includes: In each training iteration, perform the following operations: Sub-data masking is performed on the audio data to obtain masked data that masks random sub-data; Feature extraction is performed on the masked data to obtain data features; Based on the data features, predict the random sub-data that is obscured in the obscured data to obtain the predicted sub-data; Based on the difference between the obtained predicted sub-data and the random sub-data, the model parameters of the feature extraction sub-model are adjusted.
5. The method according to any one of claims 1 to 4, wherein extracting the spectral features of the sample music segment using the feature extraction sub-model comprises: Using the aforementioned feature extraction sub-model, the spectral sub-features of each of the multiple sub-segments contained in the sample music segment are extracted, and the spectral correlation features between the various spectral sub-features are extracted; By fusing the obtained spectral sub-features and spectral correlation features, the spectral features presented by the sample music segment are obtained.
6. The method according to claim 5, wherein the spectral sub-features of each sub-segment represent at least one of the following: pitch, timbre, recording environment, rhythm, instrument, and emotion of the sub-segment; and the spectral correlation features represent at least one of the following: similarity, rhythmic relationship, and pitch transition between sub-segments.
7. The method according to any one of claims 1 to 6, wherein the step of using the feature mapping sub-model to map the obtained spectral features into descriptive text features includes: Based on the model parameters of the feature mapping sub-model in this round of training, a first mapping relationship is obtained between the spectral features of each attribute and the text features of each attribute learned by the feature mapping sub-model; wherein, each of the attribute text features represents one of a plurality of reference music attributes; the plurality of reference music attributes includes at least the union of the plurality of sample music attributes and the plurality of training music attributes; Based on the first mapping relationship, the spectral features are mapped to multiple attribute text features; wherein, the multiple attribute text features represent the multiple training music attributes; Multiple attribute text features obtained are fused to generate descriptive text features.
8. The method according to any one of claims 1 to 7, further comprising, before employing the feature mapping sub-model to map the obtained spectral features into descriptive text features: Based on the preset second mapping relationship between each data format and each data volume, a target data format with a corresponding data volume smaller than the reference data volume is selected from the data formats; wherein, the parameter data volume is: the data volume corresponding to the data format used by the model parameters of the feature mapping sub-model; The spectral features are converted into the target data format to obtain the converted spectral features.
9. The method according to claim 8, wherein selecting a target data format whose corresponding data occupation is less than a reference data volume from the data formats based on a preset second mapping relationship between each data format and each data volume comprises: Get the GPU memory usage for running this round of training iterations; When it is determined that the video memory usage is greater than the preset usage, based on the preset second mapping relationship between each data format and each occupied data amount, a target data format with a corresponding occupied data amount less than the reference data amount is selected from the data formats.
10. The method according to any one of claims 1 to 9, wherein adjusting the model parameters of the feature mapping sub-model based on the difference between the obtained training music description and the sample music description corresponding to the training music description includes: Based on the differences between the various training music attributes contained in the obtained training music description and the various sample music attributes contained in the sample music description corresponding to the obtained training music description, the mapping training loss is determined. Based on the mapping training loss, the model parameters of the feature mapping sub-model are adjusted.
11. The method according to claim 10, wherein adjusting the model parameters of the feature mapping sub-model based on the mapping training loss comprises: Based on the text similarity between the obtained training music description and the corresponding sample music description, the semantic training loss is determined. Based on the mapping training loss and the semantic training loss, when the training objective is not achieved, the model parameters of the feature mapping sub-model are adjusted.
12. A music recognition method, comprising: Obtain the music to be identified; A trained music recognition model is used to extract the spectral features of the music to be recognized, and the spectral features are mapped to text features to be recognized; wherein the music recognition model is trained using the method described in any one of claims 1 to 11. The music recognition model is used to identify the features of the text to be identified and obtain a target music description; wherein the target music description is used to: introduce various target music attributes of the music to be identified in text form.
13. A training device for a music recognition model, comprising: Acquisition module: used to acquire sample datasets; wherein each sample dataset includes: a sample music segment and a sample music description corresponding to the sample music segment, the sample music description containing multiple sample music attributes of the sample music segment; Processing module: used to perform multiple rounds of iterative training on the music recognition model to be trained based on the sample dataset; wherein, the music recognition model includes: a trained feature extraction sub-model and a text recognition sub-model, and a feature mapping sub-model to be trained; each round of iterative training includes: The processing module is used to: extract the spectral features of the sample music segment using the feature extraction sub-model; The processing module is used to: map the obtained spectral features into descriptive text features using the feature mapping sub-model; wherein the descriptive text features are used to describe various training music attributes of the sample music fragment; The processing module is used to: employ the text recognition sub-model to recognize the features of the descriptive text and obtain a training music description; and The processing module is used to adjust the model parameters of the feature mapping sub-model based on the difference between the obtained training music description and the sample music description corresponding to the training music description.
14. A music recognition device, comprising: Acquisition module: Used to acquire the music to be identified; Processing module: used to extract the spectral features of the music to be identified using a trained music recognition model, and map the spectral features to be identified into text features to be identified; wherein, the music recognition model is trained using the method described in any one of claims 1 to 11; The processing module is further configured to: use the music recognition model to identify the features of the text to be identified and obtain a target music description; wherein the target music description is used to: describe various target music attributes of the music to be identified in text form.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 12.
16. A computer device, comprising: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1 to 12 according to the obtained program instructions.
17. A computer-readable storage medium storing computer-executable instructions for causing a computer to perform the method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Music emotion recognition method and device, storage medium and electronic equipment
CN111858943A
Music recognition method and training method and device of music feature extraction model
CN114023289A
Audio feature extraction model training method and audio classification method
CN115148195A
Music emotion recognition method and system based on cross-modal fusion
CN116010902A
Audio processing method and device
CN117711427A