Method and device for model training, equipment and storage medium
By extracting audio content from music samples and generating training sequences, this method solves the problems of high cost and data dependence in existing technologies, and achieves efficient training and quality improvement of song synthesis models.
Patent Information
- Application Number
- CN202411251956.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-10
AI Technical Summary
Existing song synthesis technologies rely on a large amount of manually annotated phoneme data, resulting in high modeling costs and disrupting the regularity of note durations in musical scores, making it difficult to generate high-quality songs.
By extracting audio content associated with the singing content from music samples, generating training sequences of text and melody information, a music generation model is constructed. The model is then trained using a text-digital hybrid encoder and a melody recognition model, reducing the reliance on high-quality data.
It reduced model training costs, improved the quality and performance of song synthesis, reduced the workload of manual annotation, and simplified the training process.
Smart Images

Figure CN121640959A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for model training. Background Technology
[0002] With the development of machine learning technology, machine learning models can now be used to perform tasks in various application environments. Singing Voice Synthesis (SVS) is a computer technology that attempts to simulate human singing. SVS can be seen as a special branch of Text-to-Speech (TTS). In synthesizing songs using SVS, it's not only necessary to maintain the intelligibility of the language, but also to replicate musical features such as timbre, pitch, duration, and singing style as accurately as possible. However, SVS technology still has some problems that affect the quality of the synthesized songs. Summary of the Invention
[0003] In a first aspect of this disclosure, a model training method is provided. The method includes: extracting first audio content associated with singing content from music samples; generating first annotation information based on the first audio content, the first annotation information including text content corresponding to the first audio content and first melody information of the first audio content; constructing a first training sequence based on the text content and the first melody information; inputting the first training sequence into a music generation model to generate a first set of music encoding representations; and training the music generation model based on the first set of music encoding representations and a second set of music encoding representations of music samples.
[0004] In a second aspect of this disclosure, an apparatus for model training is provided. The apparatus includes: an extraction module configured to extract first audio content associated with singing content from music samples; a first generation module configured to generate first annotation information based on the first audio content, the first annotation information including text content corresponding to the first audio content and first melody information of the first audio content; a construction module configured to construct a first training sequence based on the text content and the first melody information; a second generation module configured to input the first training sequence into a music generation model to generate a first set of music encoding representations; and a training module configured to train the music generation model based on the first set of music encoding representations and a second set of music encoding representations of music samples.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0010] Figure 2 An example architecture diagram of a model training system according to some embodiments of the present disclosure is shown;
[0011] Figure 3 An example architecture diagram of a model fine-tuning system according to some embodiments of the present disclosure is shown;
[0012] Figure 4 An example architecture diagram of another example of a model fine-tuning system according to some embodiments of the present disclosure is shown;
[0013] Figure 5 An example architecture diagram is shown, illustrating one example of a melody element according to some embodiments of the present disclosure;
[0014] Figure 6 An example architecture diagram of a music generation system according to some embodiments of the present disclosure is shown;
[0016] Figure 7 A flowchart illustrating a model training process according to some embodiments of the present disclosure is shown;
[0017] Figure 8 A block diagram of an apparatus for model training according to some embodiments of the present disclosure is shown; and
[0018] Figure 9 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0023] It is understood that the embodiments disclosed herein involve model training and inference, and the data involved in model training and inference (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0026] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0027] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0028] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.
[0029] As used herein, a “unit,” “operation unit,” or “subunit” can consist of any suitable machine learning model or network. As used herein, a set of elements or similar expressions can include one or more such elements. For example, “a set of convolutional units” can include one or more convolutional units.
[0030] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, model 130-1 with pre-training parameter values and model 130-2 with post-training parameter values can be collectively referred to as model 130 or referred to individually. Model 130 can be implemented or included in electronic device 140.
[0031] exist Figure 1 In environment 100, it is desired to train and use a machine learning model (i.e., model 130) configured for various application environments. For example, if the model is a speech synthesis model, it can generate speech corresponding to the text based on reference speech and text input by the user, or edit the speech information input by the user. For example, if model 130 is a song synthesis model, it can generate corresponding song audio based on sheet music information and singing information input by the user.
[0032] like Figure 1As shown, environment 100 includes electronic device 140 and / or electronic device 150. Electronic device 140 may contain a model training system, and electronic device 150 may contain a model application system. Figure 1 The upper part illustrates the model training phase, and the lower part illustrates the model application phase. Before training, the parameter values of model 130 can have initial values or pre-trained parameter values obtained through a pre-training process. Model 130-1 can be trained via forward and backward propagation, during which the parameter values of model 130-1 can be updated and adjusted. After training, model 130-2 is obtained. At this point, the parameter values of model 130-2 have been updated, and based on the updated parameter values, model 130-2 can be used to implement the song synthesis task in the model application phase.
[0033] During the model training phase, model 130 can be trained using a model training system based on a training sample set 110 comprising multiple training samples 112. Each training sample 112 can involve a binary format. For example, for a song synthesis task, training sample 112 can include training input 120 and training output for the song synthesis task. The training input in the song synthesis task can, for example, include reference audio and sheet music. Training sample 112, including training input 120 and training output 122, can be used to train model 130. Specifically, the training process can be performed iteratively using a large number of training samples. After training is complete, model 130 can include knowledge about the task to be processed. During the model application phase, model 130 (at this point, model 130 has the trained parameter values) can be used to perform the corresponding task. For example, model input 142 from the song synthesis task can be received, and corresponding model output 144 can be output.
[0034] exist Figure 1 In this context, electronic device 140 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0035] It should be understood that Figure 1 The components and arrangements shown in environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. Implementations of this disclosure are not limited in this respect.
[0036] As briefly mentioned earlier, song synthesis technology is used to generate songs based on input text and musical notes. Song synthesis technology includes splicing synthesis technology and artificial intelligence (AI) synthesis technology. AI synthesis technology uses machine learning models to learn musical characteristics of human voice samples, such as timbre, pitch, phoneme duration, and singing style.
[0037] Currently, the mainstream song synthesis technology is represented by the DiffSinger model. DiffSinger is a song synthesis method based on a diffusion model, which enhances the control over the singing voice by adding parameters such as pitch, dynamics, gender, and energy to generate songs that meet specific requirements.
[0038] However, the current DiffSinger model is based on phoneme modeling. This requires a large amount of manually annotated phoneme data, increasing modeling costs. Furthermore, the manual annotations in the phoneme data can disrupt the regularity of note durations in the musical score. Therefore, an additional phoneme duration prediction model is needed to predict the duration of each phoneme, further increasing the cost of training the model.
[0039] Embodiments of this disclosure propose a scheme for model training. According to various embodiments of this disclosure, first audio content associated with singing content is extracted from music samples. Based on the first audio content, first annotation information is generated, the first annotation information including text content corresponding to the first audio content and first melody information of the first audio content. Based on the text content and the first melody information, a first training sequence is constructed. The first training sequence is input into a music generation model to generate a first set of music encoding representations. Based on the first set of music encoding representations and a second set of music encoding representations of music samples, the music generation model is trained.
[0040] In this way, on the one hand, the music generation model is trained by constructing training sequences from text and melody content. This eliminates the need for processed phoneme data, reducing the training cost of the model. On the other hand, the music generation model is further improved by training a first set of music encoding representations generated by the model and a second set of music encoding representations from music samples.
[0041] Figure 2 An example architecture diagram of a model training system 200 according to some embodiments of the present disclosure is shown. Figure 2 As shown, the model training system can be implemented or included in the electronic device 140. The electronic device 140 is used to train the music generation model 223 based on the music samples 210 provided by the user, so as to update the parameters of the music generation model 223.
[0042] In some embodiments, the first audio content 212 is a vocal track associated with the singing content. The vocal track is provided to the music generation model 223 to train the music generation model 223 based on the first audio content 212.
[0043] In some embodiments, such as Figure 2 As shown, electronic device 140 pre-trains music generation model 223 based on the provided music sample 210. Music sample 210 is music data containing vocal track information. The music data can be existing music data including vocal track information and reverb, or music recorded in a high-standard recording environment containing only vocal track information. As stated above, music sample 210 (including but not limited to itself, its acquisition, and / or its use) complies with relevant laws, regulations, and rules.
[0044] If the music sample 210 is existing data containing reverb and vocal track information, then it is necessary to extract the first audio content 212 (i.e., vocal track singing) associated with the singing content from the music sample 210. In some embodiments, the audio source separation module 211 can be used to separate the music sample 210 to obtain the vocal track information in the music sample 210. The first audio content 212 may include multiple vocal tracks singing.
[0045] In some embodiments, the first audio content 212 is provided to the text recognition model 213 and the melody recognition model 214 to obtain the first annotation information 215 generated by the text recognition model 213 and the melody recognition model 214. The first annotation information 215 includes text content corresponding to the first audio content 212 and the first melody information of the first audio content 212. The text content includes text 216-1, text 216-2, etc., which can be referred to individually or collectively as text content 216. The first melody information includes first melody 217-1, first melody 217-2, etc., which can be referred to individually or collectively as first melody 217.
[0046] The text recognition model 213 can be a model built based on automatic speech recognition (ASR) technology. The text recognition model 213 converts the input acoustic features into a text sequence. Subsequently, it performs operations such as part-of-speech tagging, syntactic analysis, and semantic understanding on the text sequence to obtain the annotated text content 216 corresponding to the first audio content 212. The melody recognition model 214 can be built based on human voice MIDI recognition technology. The melody recognition model 214 is used to annotate the melody corresponding to the first audio content 212 to obtain the annotated first melody information.
[0047] The model training system constructs a first training sequence based on the text content 216 generated by the text recognition model 213 and the first melody information generated by the melody recognition model 214. The first training sequence 218 is used to train the music generation model 223. The first training sequence includes a first sequence part 219-1 corresponding to the text content 216 and a second sequence part 219-2 corresponding to the first melody information.
[0048] In order to generate high-quality music, in some embodiments, the model training system needs to determine the time information corresponding to each lyric element in the text content 216 based on the music sample 210. That is, to determine the time of appearance and duration of each lyric element in the text content 216 in the music sample 210.
[0049] The first sequence portion 219-1 is determined based on the text content 216 and the first time information 220 corresponding to the text content 216. The first sequence portion 219-1 includes the text content 216 and the time distribution corresponding to each lyric element in the text content 216. The time distribution of each lyric element in the text content 216 indicates the occurrence time and duration of each lyric element. For example, if the text content 216 corresponding to the first audio content 212 is "Today's Mood", the time distribution of the text content 216 is: "Today" appears from the 1st to the 3rd second, "Today" appears from the 3rd to the 5th second, "Heart" appears from the 5th to the 6th second, and "Feelings" appears from the 6th to the 7th second.
[0050] The second sequence portion 219-2 includes a first set of melodic elements corresponding to the first melodic information, and a temporal distribution corresponding to the melodic elements in the first set of melodic elements. The first set of melodic elements indicates the melodic elements included in the first melodic information. The second sequence portion 219-2 is determined by the model training system based on the first set of melodic information and the second temporal information 221 corresponding to the first set of melodic information. For example, the melodic element "do" in the first set of melodic information appears between the 3rd and 4th second.
[0051] The first sequence portion 219-1 and the second sequence portion 219-2 are used to construct the first training sequence 218. The first training sequence 218 can be represented as [word1, w_dur1, word2, w_dur2, ..., note1, n_dur1, note2, n_dur2, ...]. In the first training sequence 218, word1, word2, ... indicate each lyric element in the text content 216, and w_dur1, w_dur2 indicate the time distribution of each lyric element. note1, note2 indicate each melody element corresponding to the first melody information, and n_dur1, n_dur2 indicate the time distribution of each melody element in the first melody information.
[0052] After being encoded, the first training sequence 218 is provided to the music generation model 223 for training. Since the first training sequence 218 contains both text (e.g., text content 216 and first melody information) and numbers (e.g., time distribution), in some embodiments, a text-to-digital hybrid encoder 222 can be used to encode the first training sequence 218. The first training sequence 218 is provided to the text-to-digital hybrid encoder 222 for encoding, generating a hybrid encoded representation with generalization performance. The hybrid encoded representation can better support the processing of the music generation model 223. Therefore, the training effect of the music generation model 223 can be improved.
[0053] The hybrid encoding representation is provided to the music generation model 223. The music generation model 223 can be an autoregressive model comprising multiple decoder models, generating sequence data by using the output of the previous time step as the input of the next time step. Based on the provided hybrid encoding representation, the music generation model 223 generates a first set of music encoding representations 224 corresponding to the first training sequence 218. The first set of music encoding representations 224 includes multiple predicted token features 225. For example, predicted token features 225-1 and 225-2, which can be individually or collectively referred to as predicted token features 225. The first set of music encoding representations 224 can be converted into musical sound waves using a transformation model.
[0054] In some embodiments, the parameters of the music generation model 223 can be updated based on a first set of music encoding representations 224 produced by the music generation model 223 and a second set of music encoding representations 232 corresponding to the music sample 210, in order to train the music generation model 223. The second set of music encoding representations 232 indicates the true lexical features of the music sample 210.
[0055] In some embodiments, the second set of music encoding representations 232 can be the encoding obtained after processing the first audio content 212 using a trained discrete encoder 230. Specifically, the discrete encoder 230 extracts audio information from the first audio content 212 to obtain lexical features 231-1, lexical features 231-2, etc. The obtained lexical features can be referred to individually or collectively as lexical features 231.
[0056] The model training system compares the first set of music codes with the second set of music codes based on a loss function to determine the training loss of the music generation model 223. The parameters of the music generation model 223 are then updated based on the training loss.
[0057] Figure 2The training process shown trains the music generation model 223 based on low-quality music samples 210, eliminating the need for large amounts of high-quality data and reducing the training difficulty. Simultaneously, the music samples 210 are labeled using a text generation model and a melody recognition model 214 during training, eliminating the need for manual labeling. This approach reduces both the human resource costs of training and the difficulty of modeling by eliminating the need for an additional duration prediction model to predict phoneme durations.
[0058] To further improve the performance of the music generation model 223, supervised fine-tuning (SFT) can be performed on the pre-trained music generation model 223 using high-quality human vocal tracks that are manually annotated. Since the music generation model 223 has already been pre-trained, the amount of high-quality human vocal tracks and workload required for supervised fine-tuning can be significantly reduced.
[0059] Figure 3 An example architecture diagram of a model fine-tuning system 300 according to some embodiments of the present disclosure is shown. Figure 3 As shown, the model fine-tuning system can be implemented or included in the electronic device 140. The electronic device 140 is used to perform supervised fine-tuning of the music generation model 223 based on the music generation model 223 trained by the model training system, according to the high-quality vocal track provided by the user, in order to further improve the performance of the music generation model 223.
[0060] like Figure 3 As shown, the second audio content 310 indicates a high-quality vocal track. After manual annotation, the second audio content 310 generates second annotation information 321. The second annotation information 321 includes phoneme information corresponding to the second audio content 310 (e.g., phoneme information 323-1, phoneme information 323-2, which can be referred to individually or collectively as phoneme information 323) and second melody information of the second audio content 310. The second melody information includes second melodies 322-1 and 322-2, which can be referred to individually or collectively as second melody 322.
[0061] The second annotation information 321 indicates multiple phonemes corresponding to the second audio content 310 and a first time distribution 324 corresponding to the multiple phonemes, and / or a second set of melodic elements corresponding to the second audio content 310 and a second time distribution 325 corresponding to the second set of melodic elements. The second set of melodic elements indicates the melodic elements included in the second melodic information, and the second melodic information indicates the melodic elements corresponding to the second audio content 310. The method for determining the first time distribution 324 and the second time distribution 325 is the same as the model pre-training process, and will not be described again here.
[0062] In some embodiments, the second annotation information 321 can be manually generated information. The second annotation information 321, obtained by manually annotating high-quality vocal tracks, is provided to the music generation model 223 for training. The predicted lexical features 225 output by the music generation model 223 are compared with the lexical features corresponding to the second audio content 310 to determine the training loss. Based on the training loss, the parameters of the music generation model 223 are fine-tuned. It can be seen that by pre-training the music generation model 223 with relatively easy-to-obtain low-quality music samples 210 and fine-tuning the model parameters with high-quality music samples (i.e., the second audio content 310), a high-performance music generation model 223 can be obtained. In this way, the amount of high-quality music samples (i.e., the second audio content 310) used during training can be reduced, thus lowering the training cost.
[0063] Figure 4 An example architecture diagram of another example of a model fine-tuning system 400 according to some embodiments of the present disclosure is shown. Figure 4 As shown, the model fine-tuning system 400 can generate second annotation information 321 based on the melody recognition model 214.
[0064] In some embodiments, the melody recognition model 214 is used to annotate the melody information of the second audio content 310 to determine the third set of melody elements 520. For example... Figure 4 As shown, Melody 410-1 and Melody 410-2 can be referred to individually or collectively as the third melody 410.
[0065] Figure 5 An example architecture diagram of a melody element 500 according to some embodiments of the present disclosure is shown. Figure 5 As shown, compared to the manually labeled sequence 510, the third set of melody elements 520 generated by the melody recognition model 214 has problems such as disconnected boundaries and inaccurate duration. In order to ensure the training effect of the model, the third melody 410 needs to be aligned with the manually labeled phoneme information 323.
[0066] Specifically, for the target time segment corresponding to the target phoneme in the first time distribution 324, at least one melodic element corresponding to the target time segment is determined based on the third set of melodic elements 520. Based on at least one melodic element, the target melodic element corresponding to the target time segment is determined. Using the target melodic element, the third set of melodic elements 520 is updated. For example... Figure 5 As shown, melody element 1 and melody element 2 are not connected. Melody element 1 needs to be adjusted to align it with phoneme 1.
[0067] In some embodiments, the target melody element can be determined based on the fundamental frequency of the melody elements within each time segment. Specifically, the fundamental frequency of each melody element in each third group of melody elements 520 can be determined by a fundamental frequency predictor to generate melody element fundamental frequency information 530. The target melody element is determined based on the median fundamental frequency of the melody phonemes within each time segment. The melody elements are updated using the target melody element to generate an updated third group of melody elements 540. Figure 5 As shown, the target melody element is Figure 5 In the case of melody element 1, first determine the target time segment in which melody element 1 is located. Determine the fundamental frequency of elements within the target time segment, and determine melody element 1# based on the median of the fundamental frequencies of the elements.
[0068] continue Figure 4 Based on the updated third set of melody elements 540, a first training sequence 218 is constructed. The first training sequence 218 is encoded and provided to the music generation model 223. Based on the music generation model 223, a first set of music encoding representations 224 and a second set of music encoding representations 232 corresponding to the second audio content 310 are generated, and the training loss of the music generation model 223 is determined. Based on the training loss, the model parameters are fine-tuned. It can be seen that the melody elements are determined by the melody recognition model 214, and the melody elements are aligned with manually labeled data. Therefore, without affecting model performance, the workload of manual labeling and training costs can be reduced.
[0069] Figure 6 An example architecture diagram of a music generation system 600 according to some embodiments of the present disclosure is shown. Figure 6 As shown, the music generation system can be implemented or included in the electronic device 140. The music generation system is used to generate music based on a musical score input by a user.
[0070] like Figure 6 As shown, the musical score information 610 includes text and melody related to the song. Third annotation information 620 is determined by parsing the musical score information 610. The third annotation information 620 includes text content 216 related to the musical score, phoneme information 323, and melody information. The third annotation information 620, after being encoded by the text-digital hybrid encoder 222, is provided to the trained music generation model 223. The music generation model 223 generates a music encoding representation 630 including lexical features based on the input hybrid encoding. The music encoding representation 630 is provided to the sound wave conversion model 640 (token2wav), where the diffusion model of the sound wave conversion model 640 is combined with a vocoder to convert the audio features corresponding to the music encoding representation 630 into an audio waveform 650. The diffusion model is used to convert the music encoding representation 630 into implicit audio features with a higher sampling rate and more information. The vocoder is used to map the implicit audio features into the audio waveform 650 and output it.
[0071] Figure 7 A flowchart of a model training process 700 according to some embodiments of the present disclosure is shown. Process 700 can be implemented at an electronic device 140.
[0072] In box 710, the first audio content associated with the singing content is extracted from the music sample.
[0073] At box 720, first annotation information is generated based on the first audio content. The first annotation information includes text content corresponding to the first audio content and the first melody information of the first audio content.
[0074] In some embodiments, generating first annotation information includes: processing first audio content using a text recognition model to determine text content; and / or processing first audio content using a melody recognition model to determine first melody information.
[0075] At box 730, the first training sequence is constructed based on the text content and the first melody information.
[0076] In some embodiments, constructing a first training sequence based on text content and first melody information includes: constructing a first sequence portion corresponding to the text content based on the text content and first time information corresponding to the text content; constructing a second sequence portion corresponding to the first melody information based on the first melody information and second time information corresponding to the first melody information; and constructing a first training sequence based on the first sequence portion and the second sequence portion.
[0077] In some embodiments, the first sequence portion indicates multiple lyric elements and the time distribution corresponding to the multiple lyric elements, and / or the second sequence portion indicates a first group of melody elements and the time distribution corresponding to the first group of melody elements.
[0078] At box 740, the first training sequence is input into the music generation model to generate the first set of music encoding representations.
[0079] In some embodiments, inputting a first training sequence into a music generation model includes: encoding the first training sequence using a text-digit hybrid encoder to generate a hybrid encoded representation; and inputting the hybrid encoded representation into the music generation model.
[0080] At box 750, a music generation model is trained based on the first set of music encoding representations and the second set of music encoding representations of the music samples.
[0081] In some embodiments, training a music generation model based on a first set of music encoding representations and a second set of music encoding representations of music samples includes: determining a training loss based on a comparison of the first set of music encodings and the second set of music encodings; and updating the model parameters of the music generation model based on the training loss.
[0082] In some embodiments, process 700 further includes: processing the first audio content using a trained discrete encoder to determine a second set of music encoding representations.
[0083] In some embodiments, process 700 further includes: acquiring a second audio content that has been collected, the second audio content corresponding to a vocal track; acquiring second annotation information of the second audio content, the second annotation information including phoneme information corresponding to the second audio content and second melody information of the second audio content; and fine-tuning a music generation model based on the second annotation information of the second audio content.
[0084] In some embodiments, the second annotation information indicates: a plurality of phonemes corresponding to the second audio content and a first time distribution corresponding to the plurality of phonemes; and / or a second set of melodic elements corresponding to the second audio content and a second time distribution corresponding to the second set of melodic elements.
[0085] In some embodiments, process 700 further includes: processing the second audio content using a melody recognition model to generate a third set of melody elements; and updating the third set of melody elements based on a first temporal distribution of multiple phonemes to determine the second set of melody elements.
[0086] In some embodiments, updating the third set of melodic elements based on the first time distribution of multiple phonemes includes: for a target time segment in the first time distribution corresponding to a target phoneme, determining at least one melodic element corresponding to the target time segment based on the third set of melodic elements; determining a target melodic element corresponding to the target time segment based on the at least one melodic element; and updating the third set of melodic elements using the target melodic element.
[0087] In some embodiments, determining the target melody element corresponding to the target time segment based on at least one melody element includes: determining the median fundamental frequency corresponding to the target time segment based on at least one melody element; and determining the target melody element based on the median fundamental frequency.
[0088] Figure 8 A schematic structural block diagram of an apparatus 800 for model training according to certain embodiments of the present disclosure is shown. The apparatus 800 may be implemented as or included in an electronic device. Various modules / components in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0089] like Figure 8As shown, the device 800 includes an extraction module 810 configured to extract first audio content associated with singing content from music samples. The device 800 also includes a first generation module 820 configured to generate first annotation information based on the first audio content, the first annotation information including text content corresponding to the first audio content and first melody information of the first audio content. The device 800 further includes a construction module 830 configured to construct a first training sequence based on the text content and the first melody information. The device 800 also includes a second generation module 840 configured to input the first training sequence into a music generation model to generate a first set of music encoding representations. The device 800 further includes a training module 850 configured to train the music generation model based on the first set of music encoding representations and a second set of music encoding representations of music samples.
[0090] In some embodiments, the first generation module 820 is further configured to process the first audio content using a text recognition model to determine text content; and / or process the first audio content using a melody recognition model to determine first melody information.
[0091] In some embodiments, the construction module 830 is further configured to construct a first sequence portion corresponding to the text content based on the text content and the first time information corresponding to the text content; construct a second sequence portion corresponding to the first melody information based on the first melody information and the second time information corresponding to the first melody information; and construct a first training sequence based on the first sequence portion and the second sequence portion.
[0092] In some embodiments, the second generation module 840 is further configured to encode a first training sequence using a text-digit hybrid encoder to generate a hybrid encoded representation; and to input the hybrid encoded representation into a music generation model.
[0093] In some embodiments, the training module 850 is further configured to determine a training loss based on a comparison of the first set of music codes and the second set of music codes; and to update the model parameters of the music generation model based on the training loss.
[0094] In some embodiments, the apparatus 800 further includes a second set of music encoding representation generation modules configured to process the first audio content using a trained discrete encoder to determine the second set of music encoding representations.
[0095] In some embodiments, the device 800 further includes a fine-tuning module configured to acquire a second audio content that corresponds to a vocal track; acquire second annotation information of the second audio content, the second annotation information including phoneme information corresponding to the second audio content and second melody information of the second audio content; and fine-tune a music generation model based on the second annotation information of the second audio content.
[0096] In some embodiments, the apparatus 800 further includes a second set of melody element determination module, configured to process the second audio content using a melody recognition model to generate a third set of melody elements; and to update the third set of melody elements based on a first temporal distribution of multiple phonemes to determine the second set of melody elements.
[0097] In some embodiments, the apparatus 800 further includes a third set of melody element update module, configured to, for a target time segment in the first time distribution corresponding to the target phoneme, determine at least one melody element corresponding to the target time segment based on the third set of melody elements; determine a target melody element corresponding to the target time segment based on the at least one melody element; and update the third set of melody elements using the target melody element.
[0098] In some embodiments, the apparatus 800 further includes a target melody element determination module, configured to determine the median of the fundamental frequency corresponding to a target time segment based on at least one melody element; and to determine the target melody element based on the median of the fundamental frequency.
[0099] Figure 9 A block diagram is shown illustrating an electronic device 900 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 9 The electronic device 900 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 9 The electronic device 900 shown can be used to achieve Figure 1 140 electronic devices.
[0100] like Figure 9 As shown, electronic device 900 is in the form of a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 900.
[0101] Electronic device 900 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 900.
[0102] Electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 9 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 920 may include computer program product 925 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0103] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 900 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0104] Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 900 can also communicate with one or more external devices (not shown) via communication unit 940 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 900, or with any device that enables electronic device 900 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0105] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0106] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0107] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0108] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0110] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method of model training, comprising: extracting, from a music sample, first audio content associated with singing content; generating, based on the first audio content, first annotation information, the first annotation information comprising text content corresponding to the first audio content and first melody information of the first audio content; constructing, based on the text content and the first melody information, a first training sequence; inputting, to a music generation model, the first training sequence to generate a first set of music encoded representations; and training, based on the first set of music encoded representations and a second set of music encoded representations of the music sample, the music generation model.
2. The method of claim 1, wherein generating first annotation information comprises: processing the first audio content with a text recognition model to determine the text content; and / or processing the first audio content with a melody recognition model to determine the first melody information.
3. The method of claim 2, wherein constructing, based on the text content and the first melody information, a first training sequence comprises: constructing, based on the text content and first time information corresponding to the text content, a first sequence portion corresponding to the text content; constructing, based on the first melody information and second time information corresponding to the first melody information, a second sequence portion corresponding to the first melody information; and constructing, based on the first sequence portion and the second sequence portion, the first training sequence.
4. The method of claim 3, wherein: the first sequence portion indicates a plurality of lyric elements and a time distribution corresponding to the plurality of lyric elements, and / or the second sequence portion indicates a first set of melody elements and a time distribution corresponding to the first set of melody elements.
5. The method of claim 1, further comprising: processing the first audio content with a trained discrete encoder to determine the second set of music encoded representations.
6. The method of claim 1, wherein training, based on the first set of music encoded representations and a second set of music encoded representations of the music sample, the music generation model comprises: determining a training loss based on a comparison of the first set of music encoded and the second set of music encoded; and updating, based on the training loss, model parameters of the music generation model.
7. The method of claim 1, wherein inputting, to a music generation model, the first training sequence comprises: encoding, with a text numerical hybrid encoder, the first training sequence to generate a hybrid encoded representation; and inputting, to the music generation model, the hybrid encoded representation.
8. The method of claim 1, further comprising: obtaining a second audio content captured, the second audio content corresponding to a vocal track; obtaining second annotation information of the second audio content, the second annotation information comprising phoneme information corresponding to the second audio content and second melody information of the second audio content; and fine-tuning, based on the second annotation information of the second audio content, the music generation model.
9. The method of claim 8, wherein the second annotation information indicates: a plurality of phonemes corresponding to the second audio content and a first temporal distribution corresponding to the plurality of phonemes; and / or a second set of melody elements corresponding to the second audio content and a second temporal distribution corresponding to the second set of melody elements.
10. The method of claim 9, further comprising: processing the second audio content with a melody recognition model to generate a third set of melody elements; and updating the third set of melody elements based on the first temporal distribution of the plurality of phonemes to determine the second set of melody elements.
11. The method of claim 10, wherein updating the third set of melody elements based on the first temporal distribution of the plurality of phonemes comprises: determining, based on the third set of melody elements, at least one melody element corresponding to a target time segment of the first temporal distribution corresponding to a target phoneme; determining, based on the at least one melody element, a target melody element corresponding to the target time segment; and updating the third set of melody elements with the target melody element.
12. The method of claim 11, wherein determining, based on the at least one melody element, a target melody element corresponding to the target time segment comprises: determining, based on the at least one melody element, a median fundamental frequency corresponding to the target time segment; and determining the target melody element based on the median fundamental frequency.
13. An apparatus for model training, comprising: an extraction module configured to extract, from a music sample, first audio content associated with singing content; a first generation module configured to generate, based on the first audio content, first annotation information, the first annotation information comprising text content corresponding to the first audio content and first melody information of the first audio content; a construction module configured to construct, based on the text content and the first melody information, a first training sequence; a second generation module configured to input, to a music generation model, the first training sequence to generate a first set of music code representations; and a training module configured to train, based on the first set of music code representations and a second set of music code representations of the music sample, the music generation model.
14. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method of any one of claims 1-12.
15. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method of any one of claims 1-12.