Base model training method, task execution method, equipment, medium and product
By introducing labeled training data and clustering operations into the base model, the base model is fine-tuned, and the base model is solved, and the base model is poorly performed in downstream music processing tasks is achieved, achieving better task understanding and effect.
Patent Information
- Application Number
- CN202510250746.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
AI Technical Summary
The existing pedestal models perform poorly in downstream music processing tasks, mainly due to the fully self-supervised training method that leads to the lack of in-depth understanding of specific scenarios.
By introducing labeled training data related to downstream music processing tasks, determining the cluster center in combination with clustering operations, fine-tuning the pre-trained pedestal model to improve the model's performance in downstream tasks.
Through fine-tuning, the model has achieved better understanding and effects in downstream music processing tasks, especially in language classification, genre classification and timbre classification tasks.
Smart Images

Figure CN120106173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a base model training method, a task execution method, a device, a medium and a product. Background Art
[0002] With the rapid rise and development of big model technology, big models can already meet people's needs in many aspects of life and work, but there is still a lot of room for development for deeper content understanding. For example, in the field of music, although there are some base models for music understanding active in the industry, the training of these models adopts a completely self-supervised approach. Since the model is not told the specific understanding goal, the model can only find patterns from a large number of music samples. Therefore, when the model is applied to downstream tasks, due to the lack of in-depth understanding of the specific scenarios involved in the downstream tasks, the model does not show good results in the specific scenarios involved in the downstream tasks. Summary of the invention
[0003] In view of this, the purpose of the present invention is to provide a base model training method, task execution method, device, medium and product, by introducing labeled training data related to downstream music processing tasks, and combining clustering operations to determine cluster centers, so as to fine-tune the base model of full self-supervised training, so that the model can achieve better results in downstream music processing tasks. The specific scheme is as follows:
[0004] In a first aspect, the present application provides a base model training method, comprising:
[0005] Input the training data into the pre-trained base model to obtain output representations corresponding to each of the training data; the training data includes a first training sample and a second training sample; the first training sample is an unlabeled sample used for pre-training the base model, and the second training sample is a labeled sample related to a downstream music processing task; the output representation includes a first output representation corresponding to the first training sample and a second output representation corresponding to the second training sample;
[0006] Clustering each of the second training samples using the second output representation and based on the label of the second training sample to obtain a number of initial cluster centers;
[0007] Clustering each of the first training samples using the initial cluster centers and based on the first output representation to obtain a plurality of target cluster centers;
[0008] The pre-trained base model is fine-tuned based on the target cluster center to obtain a target base model, so as to perform the downstream music processing task using the music representation of the music data to be processed output by the target base model.
[0009] Optionally, the number of the second training samples is smaller than the number of the first training samples.
[0010] Optionally, clustering each of the second training samples using the second output representation and based on the label of the second training sample to obtain a number of initial cluster centers includes:
[0011] Clustering the second training samples with the same label using the second output representation to obtain a number of initial cluster centers;
[0012] The number of the initial cluster centers is the same as the number of different labels corresponding to the second training samples.
[0013] Optionally, clustering the second training samples with the same label using the second output representation to obtain a number of initial cluster centers includes:
[0014] A mean is calculated based on the second output representation corresponding to the second training samples having the same label to obtain a plurality of initial cluster centers.
[0015] Optionally, when the pre-trained base model includes an encoder and a decoder, the number of the target cluster centers is consistent with the dimension of the output representation of the decoder.
[0016] Optionally, before inputting the training data into the pre-trained base model, the process further includes:
[0017] Freezing model parameters of the pre-trained base model;
[0018] Accordingly, before fine-tuning the pre-trained base model based on the target cluster center, the method further includes:
[0019] Unfreeze the model parameters of the encoder and decoder in the pre-trained base model.
[0020] Optionally, fine-tuning the pre-trained base model based on the target cluster center to obtain a target base model includes:
[0021] Determine a cluster where any training sample is located from the clusters corresponding to the target cluster center; the any training sample is any training sample in the first training samples;
[0022] Determine, from the target cluster centers, a cluster center corresponding to the cluster in which any one of the training samples is located, so as to obtain a cluster center corresponding to any one of the training samples;
[0023] A target loss is determined based on the cluster center corresponding to any one of the training samples and the first output representation, and the pre-trained base model is fine-tuned using the target loss to obtain a target base model.
[0024] In a second aspect, the present application provides a method for executing a downstream music processing task based on a base model, wherein the base model is a target base model trained based on the aforementioned method, and the method comprises:
[0025] Get the music data to be processed;
[0026] Inputting the music data to be processed into the target base model to obtain a music representation obtained by the target base model after encoding and decoding the music data to be processed in sequence;
[0027] The music representation is transmitted to a task execution node corresponding to a downstream music processing task, so that the task execution node completes the downstream music processing task based on the music representation.
[0028] In a third aspect, the present application provides an electronic device, including:
[0029] Memory, used to store computer programs;
[0030] A processor is used to execute the computer program to implement the aforementioned method.
[0031] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned method when executed by a processor.
[0032] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instructions, which implement the aforementioned method when executed by a processor.
[0033] In the present application, training data is input into a pre-trained base model to obtain output representations corresponding to each of the training data; the training data includes a first training sample and a second training sample; the first training sample is an unlabeled sample used for pre-training the base model, and the second training sample is a labeled sample related to a downstream music processing task; the output representation includes a first output representation corresponding to the first training sample and a second output representation corresponding to the second training sample; using the second output representation and based on the label of the second training sample, each of the second training samples is clustered to obtain a number of initial cluster centers; using the initial cluster centers and based on the first output representation, each of the first training samples is clustered to obtain a number of target cluster centers; based on the target cluster centers, the pre-trained base model is fine-tuned to obtain a target base model, so as to perform the downstream music processing task using the music representation of the music data to be processed output by the target base model. It can be seen that the present application introduces labeled training data related to downstream music processing tasks, combines clustering operations to determine cluster centers, and then uses cluster centers to fine-tune the fully self-supervised trained base model, so that the model has better understanding capabilities in downstream music processing tasks, thereby achieving better results in downstream music processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0035] Figure 1 A system architecture diagram disclosed in this application;
[0036] Figure 2 A flow chart of a base model training method disclosed in this application;
[0037] Figure 3 A flowchart of pre-training of a base model disclosed in this application;
[0038] Figure 4 A flow chart of fine-tuning a base model disclosed in this application;
[0039] Figure 5 A flow chart of a method for executing a downstream music processing task based on a base model disclosed in this application;
[0040] Figure 6 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] Although there are some pedestal models for music understanding active in the industry, the training of these models adopts a completely self-supervised approach. Since the model is not told the specific understanding goal, the model can only find patterns from a large number of music samples. Therefore, when the model is applied to downstream tasks, the model cannot show good results in downstream tasks. To this end, the present application provides a pedestal model training method, which introduces labeled training data related to downstream music processing tasks, combines clustering operations to determine cluster centers, and thus fine-tunes the pedestal model trained in a completely self-supervised manner, so that the model can achieve better results in downstream music processing tasks.
[0043] The system framework used in the method disclosed in this application can be specifically found in Figure 1 As shown, it may specifically include: a backend server 01 and one or more user terminals 02 connected to the backend server 01. The user terminal 02 may be a mobile terminal, such as a smart phone, a tablet computer, a laptop computer, etc., or a web terminal, which is not limited here.
[0044] In the present application, the backend server 01 is used to train the base model and process the music data to be processed through the trained target base model; the user terminal 02 is used to send the music data to be processed to the backend server 01.
[0045] Specifically, when the backend server 01 is used to train the base model, the steps of executing the base model training method include inputting the training data into the pre-trained base model to obtain the output representations corresponding to each training data; wherein the training data includes a first training sample and a second training sample; the first training sample is a sample without a label used for pre-training the base model, and the second training sample is a labeled sample related to the downstream music processing task; accordingly, the output representation includes a first output representation corresponding to the first training sample and a second output representation corresponding to the second training sample. Using the second output representation and based on the label of the second training sample, clustering each second training sample to obtain a number of initial clustering centers; using the initial clustering centers and based on the first output representation, clustering each first training sample to obtain a number of target clustering centers; fine-tuning the pre-trained base model based on the target clustering center to obtain a target base model, so as to perform the downstream music processing task using the music representation of the music data to be processed output by the target base model.
[0046] Furthermore, the background server 01 is used to execute the steps of the downstream music processing task execution method based on the base model when processing the music data to be processed through the trained target base model, including obtaining the music data to be processed sent by the user end 02, and inputting the music data to be processed into the target base model to obtain the music representation obtained by the target base model after encoding and decoding the music data to be processed in sequence; and then transmitting the music representation to the task execution node corresponding to the downstream music processing task, so that the task execution node completes the downstream music processing task based on the music representation.
[0047] See also Figure 2 As shown, an embodiment of the present invention discloses a base model training method, comprising:
[0048] Step S11, input the training data into the pre-trained base model to obtain the output representation corresponding to each of the training data; the training data includes a first training sample and a second training sample; the first training sample is an unlabeled sample used for pre-training the base model, and the second training sample is a labeled sample related to the downstream music processing task; the output representation includes a first output representation corresponding to the first training sample and a second output representation corresponding to the second training sample.
[0049] In this embodiment, a pre-trained base model is obtained, and training data is input into the pre-trained base model to obtain output representations corresponding to each training data. The pre-trained base model is a model obtained by full self-supervised training based on the first training sample without labels.
[0050] For example, pre-trained base models include but are not limited to EncodecMAE (Encodec Masked Autoencoders, a model that uses neural audio encoders to generate discrete targets and learns general audio representations based on masked autoencoders) model, MERT (Music Understanding Model with Large-Scale Self-supervised Training, a music understanding model with large-scale self-supervised training) model, and LAION-CLAP (LAION Contrastive Language-Audio Pretraining, a self-supervised pre-training model that connects language and audio developed by the LAION team) model.
[0051] Specifically, the training data is input into a pre-trained base model, so that the training data is encoded and decoded in sequence through the pre-trained base model to obtain an output representation corresponding to the training data.
[0052] It should be noted that in order to improve the effect of the base model in downstream music processing tasks, a labeled second training sample related to the downstream music processing task can be collected, and the first training sample and the second training sample can be input into the pre-trained base model to obtain a first output representation corresponding to the first training sample and a second output representation corresponding to the second training sample.
[0053] It should also be noted that the number of the second training samples is less than the number of the first training samples, that is, only a small number of labeled second training samples are needed in this embodiment. Moreover, in the field of music, since the input of most base models is music clips of 15 seconds or less, the first training samples and the second training samples can both use music clips of 15 seconds or less, which can be obtained by slicing songs, and of course, music clips of other lengths can also be used.
[0054] Step S12: clustering each of the second training samples using the second output representation and based on the labels of the second training samples to obtain a number of initial cluster centers.
[0055] In this embodiment, after obtaining the second output representation corresponding to the second training sample, the second training samples can be clustered using the second output representation and based on the labels of the second training samples to obtain a number of initial cluster centers. The clustering algorithm includes but is not limited to the K-means clustering algorithm, the spectral clustering algorithm, and the like.
[0056] Specifically, since each second training sample has a corresponding label, and the same label may correspond to one or more second training samples, the second output representation can be used to cluster the second training samples with the same label to obtain a number of initial cluster centers. It should be noted that the number of initial cluster centers is the same as the number of different labels corresponding to the second training samples.
[0057] Taking the second training sample including ten A-labeled training data and ten B-labeled training data as an example, the ten A-labeled training data are clustered using the second output representation corresponding to the ten A-labeled training data to obtain the first initial cluster center, and the ten B-labeled training data are clustered using the second output representation corresponding to the ten B-labeled training data to obtain the second initial cluster center. It can be seen that the number of initial cluster centers is the same as the number of different labels (A and B) corresponding to the second training sample.
[0058] When the clustering algorithm used is the K-means clustering algorithm, the initial cluster centers may be determined by calculating the mean based on the second output representation corresponding to the second training samples with the same label to obtain several initial cluster centers.
[0059] Taking the second training sample including ten A-labeled training data and ten B-labeled training data as an example, the mean is calculated based on the second output representation corresponding to the ten A-labeled training data to obtain the first initial cluster center, and the mean is calculated based on the second output representation corresponding to the ten B-labeled training data to obtain the second initial cluster center.
[0060] Step S13: clustering the first training samples using the initial cluster centers and based on the first output representation to obtain a plurality of target cluster centers.
[0061] In this embodiment, after obtaining the initial cluster centers and the first output representations corresponding to the first training samples, the first training samples are clustered using the initial cluster centers and based on the first output representations corresponding to the first training samples to obtain a number of target cluster centers.
[0062] It can be understood that, based on the initial cluster centers determined based on the labeled second training samples related to the downstream music processing task, each first training sample is clustered using the first output representation corresponding to the first training sample to obtain a number of target cluster centers.
[0063] It should be noted that when the pre-trained base model includes an encoder and a decoder, the number of target cluster centers is consistent with the dimension of the output representation of the decoder. For example, if the dimension of the output representation of the decoder is 1024, the number of target cluster centers is also 1024 accordingly.
[0064] Furthermore, when the algorithm used for clustering is the K-means clustering algorithm, a label prediction is performed on the first training sample without a label based on the initial cluster center to determine the predicted label corresponding to the first training sample. Then, the number of clusters is set based on the dimension of the output representation of the decoder, and the first training sample is clustered using the predicted label corresponding to the first training sample to obtain the same number of clusters as the number of clusters, and the mean is calculated based on the first output representation corresponding to the first training sample contained in each cluster to obtain the target cluster center corresponding to each cluster.
[0065] Step S14: fine-tune the pre-trained base model based on the target cluster center to obtain a target base model, so as to perform the downstream music processing task using the music representation of the music data to be processed output by the target base model.
[0066] In this embodiment, before inputting the training data into the pre-trained base model, the model parameters of the pre-trained base model can be frozen first. After obtaining the target cluster center, the model parameters of the encoder and decoder in the pre-trained base model are unfrozen, and then the encoder and decoder in the pre-trained base model are fine-tuned based on the target cluster center to obtain the target base model. In this way, this embodiment freezes the model parameters of the pre-trained base model first, and then unfreezes the model parameters of the encoder and decoder in the pre-trained base model, thereby fine-tuning only the encoder and decoder based on the model parameters of the pre-trained base model, so as to speed up the model fine-tuning efficiency and improve the model effect.
[0067] Specifically, the cluster cluster where any training sample is located is determined from the cluster clusters corresponding to the target cluster center; wherein any training sample is any training sample in the first training samples; then the cluster center corresponding to the cluster cluster where any training sample is located is determined from the target cluster center to obtain the cluster center corresponding to any training sample; the target loss is determined based on the cluster center corresponding to any training sample and the first output representation, and the pre-trained base model is fine-tuned using the target loss to obtain the target base model.
[0068] Among them, the target loss can be determined by MSE (Mean Squared Error), MAE (Mean Absolute Error), L1 loss function, etc. Taking the determination of the target loss by the mean square error as an example, the target loss is determined based on the square of the difference between the cluster center corresponding to any training sample and the first output representation. The formula involved is as follows:
[0069] ;
[0070] represents the target loss, M represents the total number of the first training samples, Indicates The first output representation corresponding to the training samples, Indicates The cluster centers corresponding to the training samples.
[0071] Taking the following music processing tasks as any one or a combination of language classification tasks, genre classification tasks and timbre classification tasks as an example, the training process of the base model is illustrated. The specific scheme is as follows:
[0072] A pre-trained base model is obtained, wherein the pre-trained base model is a model obtained by fully self-supervised training based on the first training sample without labels. Second training samples with labels related to the downstream music processing task are collected, for example, the second training samples include training samples with language labels, training samples with genre labels, and training samples with timbre labels, and there are 10 types of language labels, genre labels, and timbre labels.
[0073] The first unlabeled training sample and the second labeled training sample are input into the pre-trained base model to obtain the first output representation corresponding to the first training sample and the second output representation corresponding to the second training sample; then the mean is calculated based on the second output representation corresponding to the second training sample with the same label to obtain several initial cluster centers, and the number of initial cluster centers is 30 at this time. After obtaining the initial cluster centers, the first training samples are clustered using the first output representation corresponding to the first training samples based on the initial cluster centers to obtain several target cluster centers, and then the pre-trained base model is fine-tuned based on the target cluster centers to obtain the target base model.
[0074] Since this embodiment introduces a labeled second training sample related to the downstream music processing task to fine-tune the pre-trained base model in combination with clustering, the fine-tuned target base model can achieve better results in the downstream music processing task, that is, the fine-tuned target base model can achieve better results in the language classification task, the genre classification task and the timbre classification task than before fine-tuning.
[0075] It can be seen that the present application introduces labeled training data related to downstream music processing tasks, combines clustering operations to determine cluster centers, and then uses cluster centers to fine-tune the fully self-supervised trained base model, so that the model has better understanding capabilities in downstream music processing tasks, thereby achieving better results in downstream music processing tasks.
[0076] See also Figure 3 As shown in the figure, taking the EncodecMAE model used in the base model as an example, the pre-training process of the base model is described in detail, including:
[0077] Obtain a first training sample without a label, wherein the training sample is audio data. Input the first training sample into a base model, extract the Mel spectrum features of the first training sample through a feature extractor in the base model, and perform dimension conversion on the Mel spectrum features through a linear transformation layer in the base model to obtain a representation after dimension conversion. Then, perform a random masking of a preset proportion on the representation after dimension conversion, that is, randomly discard a preset proportion (e.g., 50%) of features to obtain a masked representation. Input the masked representation into a MAE (Masked Autoencoders) encoder of a Transformer structure to obtain an encoded representation; wherein the dimensions of the encoded representation and the masked representation are consistent. Perform an expansion operation on the encoded representation to obtain an expanded representation; the dimensions of the expanded representation and the dimension converted representation are consistent. Then, input the expanded representation into a MAE decoder of a Transformer structure to obtain an output representation.
[0078] After the first training sample is input into the base model, the first training sample is processed in sequence by the Encodec encoder (an open source audio compression encoder) and the residual vector quantization technology (RVQ) in the base model to obtain a processed representation.
[0079] The training loss is calculated for the output representation and the processed representation, so as to train the base model using the training loss to obtain a pre-trained base model. The training loss can be calculated using mean square error, mean absolute error, etc.
[0080] Taking the calculation of training loss using mean square error as an example, the training loss is calculated based on the square of the difference between the output representation and the processed representation. The formula involved is as follows:
[0081] ;
[0082] represents the training loss, M represents the total number of the first training samples, Indicates The output representation corresponding to the training samples is Indicates The processed representation corresponding to the training samples.
[0083] In this way, this embodiment allows the base model to learn the training loss between the processed representation and the output representation restored after masking, so that the base model can learn the partial representation discarded by the random mask, thereby forming an understanding of the music content.
[0084] See also Figure 4 As shown in the figure, taking the EncodecMAE model as an example, the fine-tuning process of the base model is described in detail, including:
[0085] The first unlabeled training sample used for pre-training the base model and the second labeled training sample related to the downstream music processing task are input into the pre-trained base model, so as to extract the Mel spectrum features of the training samples through the feature extractor in the base model, and perform dimension conversion on the Mel spectrum features through the linear transformation layer in the base model to obtain the dimension converted representation. Then, the dimension converted representation is randomly masked at a preset ratio, that is, the features of a preset ratio (for example, 50%) are randomly discarded to obtain the masked representation. The masked representation is input into the MAE encoder of the Transformer structure to obtain the encoded representation; wherein the dimensions of the encoded representation and the masked representation are consistent. The encoded representation is extended to obtain the extended representation; the dimensions of the extended representation and the dimension converted representation are consistent. Then, the extended representation is input into the MAE decoder of the Transformer structure to obtain the first output representation corresponding to the first training sample and the second output representation corresponding to the second training sample.
[0086] Using the second output representation and based on the label of the second training sample, K-means clustering is performed on each second training sample to obtain a number of initial cluster centers; that is, the mean is calculated based on the second output representation corresponding to the second training sample with the same label to obtain a number of initial cluster centers. It should be noted that when the downstream music processing task is a multi-scale task, the several initial cluster centers can also be recorded as multi-scale cluster centers. Then, based on the initial cluster centers, K-means clustering is performed on each first training sample using the first output representation corresponding to the first training sample to obtain a number of target cluster centers, which can also be recorded as virtual labels.
[0087] Then, for any training sample in the first training samples, the cluster cluster where any training sample is located is determined from the cluster cluster corresponding to the target cluster center, and the cluster center corresponding to the cluster cluster where any training sample is located is determined from the target cluster center to obtain the cluster center corresponding to any training sample, and then the mean square error is calculated for the cluster center corresponding to any training sample and the first output representation to obtain the target loss, and the pre-trained base model is fine-tuned using the target loss to obtain the target base model.
[0088] In the actual application process of the target base model, the music data to be processed is input into the target base model to obtain the music representation of the music data to be processed, and the music representation of the music data to be processed is used to perform downstream music processing tasks.
[0089] In this way, this embodiment introduces labeled training data related to the downstream music processing tasks, combines the clustering operation to determine the cluster center, and then uses the cluster center to fine-tune the pre-trained base model, so that the model has better understanding ability in the downstream music processing tasks, thereby achieving better results in the downstream music processing tasks.
[0090] See also Figure 5 As shown, an embodiment of the present invention discloses a method for executing a downstream music processing task based on a base model, wherein the base model is a target base model trained based on the above-mentioned embodiment, and the method comprises:
[0091] Step S21, obtaining music data to be processed;
[0092] Step S22, inputting the music data to be processed into the target base model to obtain a music representation obtained by the target base model after encoding and decoding the music data to be processed in sequence;
[0093] Step S23: transmitting the music representation to a task execution node corresponding to a downstream music processing task, so that the task execution node completes the downstream music processing task based on the music representation.
[0094] In this embodiment, during the application of the target base model, the music data to be processed is first acquired and input into the target base model, so that the music data to be processed is encoded and decoded in turn by the encoder and decoder in the target base model, thereby obtaining a music representation of the music data to be processed, and then the music representation of the music data to be processed is transmitted to the task execution node corresponding to the downstream music processing task, so that the task execution node completes the downstream music processing task based on the music representation.
[0095] Determination of the music representation of the music data to be processed may specifically include: inputting the music data to be processed into the target base model, extracting the Mel-spectrogram features of the music data to be processed through the feature extractor in the target base model, and performing dimension conversion on the Mel-spectrogram features through the linear transformation layer in the target base model to obtain the representation after dimension conversion. Then, the representation after dimension conversion is randomly masked at a preset ratio, that is, a preset ratio of features is randomly discarded to obtain the masked representation. The masked representation is input into the MAE encoder of the Transformer structure to obtain the encoded representation. The encoded representation is expanded to obtain the expanded representation. The expanded representation is then input into the MAE decoder of the Transformer structure to obtain the music representation of the music data to be processed.
[0096] It can be seen that this embodiment can better understand and process the music data to be processed by using the target base model that is fine-tuned based on labeled training data related to the downstream music processing tasks, so as to obtain a more accurate music representation, thereby better completing the downstream music processing tasks based on the music representation, so that the target base model can achieve better results in the downstream music processing tasks.
[0097] Furthermore, the present application also discloses an electronic device. Figure 6 This is a structural diagram of an electronic device 10 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0098] Figure 6 The present invention provides a schematic diagram of the structure of an electronic device 10 provided in an embodiment of the present application. The electronic device 10 may specifically include: at least one processor 11, at least one memory 12, a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. The memory 12 is used to store a computer program, which is loaded and executed by the processor 11 to implement the relevant steps in the method disclosed in any of the aforementioned embodiments. In addition, the electronic device 10 in this embodiment may specifically be an electronic computer.
[0099] In this embodiment, the power supply 13 is used to provide working voltage for each hardware device on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 15 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0100] In addition, the memory 12, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 121, a computer program 122, etc., and the storage method can be temporary storage or permanent storage.
[0101] The operating system 121 is used to manage and control the hardware devices on the electronic device 10 and the computer program 122, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the method performed by the electronic device 10 disclosed in any of the aforementioned embodiments, the computer program 122 can further include a computer program that can be used to complete other specific tasks.
[0102] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned disclosed method. The specific steps of the method can refer to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.
[0103] Furthermore, the present application also discloses a computer program product, including a computer program / instruction, wherein the computer program / instruction implements the aforementioned disclosed method when executed by a processor. The specific steps of the method can refer to the corresponding contents disclosed in the aforementioned embodiment, and will not be repeated here.
[0104] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0105] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0106] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0107] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0108] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A base model training method, characterized in that: include: Input the training data into the pre-trained base model to obtain output representations corresponding to each of the training data; The training data includes a first training sample and a second training sample; the first training sample is an unlabeled sample used for pre-training a base model, and the second training sample is a labeled sample related to a downstream music processing task; the output representation includes a first output representation corresponding to the first training sample and a second output representation corresponding to the second training sample; Clustering each of the second training samples using the second output representation and based on the label of the second training sample to obtain a number of initial cluster centers; Clustering each of the first training samples using the initial cluster centers and based on the first output representation to obtain a plurality of target cluster centers; The pre-trained base model is fine-tuned based on the target cluster center to obtain a target base model, so as to perform the downstream music processing task using the music representation of the music data to be processed output by the target base model.
2. The base model training method according to claim 1, characterized in that: The number of the second training samples is smaller than the number of the first training samples.
3. The base model training method according to claim 1, characterized in that: The clustering of each of the second training samples using the second output representation and based on the label of the second training sample to obtain a plurality of initial cluster centers includes: Clustering the second training samples with the same label using the second output representation to obtain a number of initial cluster centers; The number of the initial cluster centers is the same as the number of different labels corresponding to the second training samples.
4. The base model training method according to claim 3, characterized in that: The clustering of the second training samples having the same label by using the second output representation to obtain a plurality of initial cluster centers includes: A mean is calculated based on the second output representation corresponding to the second training samples having the same label to obtain a plurality of initial cluster centers.
5. The base model training method according to claim 1, characterized in that: When the pre-trained base model includes an encoder and a decoder, the number of the target cluster centers is consistent with the dimension of the output representation of the decoder.
6. The base model training method according to claim 5, characterized in that: Before inputting the training data into the pre-trained base model, the method further includes: Freezing model parameters of the pre-trained base model; Accordingly, before fine-tuning the pre-trained base model based on the target cluster center, the method further includes: Unfreeze the model parameters of the encoder and decoder in the pre-trained base model.
7. The base model training method according to any one of claims 1 to 6, characterized in that: The fine-tuning of the pre-trained base model based on the target cluster center to obtain a target base model includes: Determine a cluster where any training sample is located from the clusters corresponding to the target cluster center; the any training sample is any training sample in the first training samples; Determine, from the target cluster centers, a cluster center corresponding to the cluster in which any one of the training samples is located, so as to obtain a cluster center corresponding to any one of the training samples; A target loss is determined based on the cluster center corresponding to any one of the training samples and the first output representation, and the pre-trained base model is fine-tuned using the target loss to obtain a target base model.
8. A method for executing downstream music processing tasks based on a base model, characterized in that: The base model is a target base model trained based on the method according to any one of claims 1 to 7, the method comprising: Get the music data to be processed; Inputting the music data to be processed into the target base model to obtain a music representation obtained by the target base model after encoding and decoding the music data to be processed in sequence; The music representation is transmitted to a task execution node corresponding to a downstream music processing task, so that the task execution node completes the downstream music processing task based on the music representation.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, which implements the method according to any one of claims 1 to 8 when executed by a processor.
11. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.