Audio data generation method and device, storage medium and electronic equipment

By extracting and evolving the original audio data and generating target audio data, the problem of low accuracy of machine learning generation audio data and single style of music theory rules is solved, and high-precision and diversified audio data generation is achieved.

CN120279869APending Publication Date: 2025-07-08HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510645252.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing machine learning-based audio data generation method has the problem of low accuracy, and the audio data generated by the style imitation melody based on music theory rules is single, which is difficult to meet the customized needs of users.

Method used

By extracting the original audio data, the original melody skeleton features are obtained, and the melody trend characteristics are determined based on the melody skeleton features, and feature evolution is carried out, and the target audio data is finally generated, avoiding the blindness of machine learning and the style singularity of music theory rules.

Benefits of technology

It improves the accuracy and diversity of audio data, meets users' needs for customized song styles, and provides a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279869A_ABST
    Figure CN120279869A_ABST
Patent Text Reader

Abstract

The invention relates to an audio data generation method and device, a storage medium and electronic equipment, and relates to the technical field of audio data processing, and the method comprises the steps: carrying out the feature extraction of original audio data, and obtaining an original melody skeleton feature; according to the original melody skeleton features, melody trend features of the original audio data are determined; performing feature evolution on the original melody skeleton features to obtain target melody skeleton features; and generating target audio data according to the melody trend features and the target melody skeleton features. According to the invention, the accuracy of the obtained target audio data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of audio data processing. More specifically, embodiments of the present disclosure relate to a method for generating audio data, an apparatus for generating audio data, a computer-readable storage medium, and an electronic device. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure recited in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] The existing audio data generated based on machine learning models is blind; that is, in the actual process of generating audio data, when users have more specific requirements for the melody trend of the generated song, the uncontrollability of machine learning will be reflected, making it difficult to meet higher customization requirements, and thus the accuracy of the obtained audio data is relatively low. Summary of the Invention

[0004] However, in related technical solutions, on the one hand, the audio data obtained based on machine learning has the problem of relatively low accuracy; on the other hand, the audio data obtained by imitating the melody based on music theory rules has the problem of single style.

[0005] Therefore, there is a great need for an improved method for generating audio data, which extracts the features of the original audio data to obtain the original melody skeleton features; then determines the melody trend features of the original audio data according to the original melody skeleton features; further evolves the features of the original melody skeleton features to obtain the target melody skeleton features; and finally generates the target audio data according to the melody trend features and the target melody skeleton features, so as to realize the generation of audio data with diverse styles on the basis of improving the accuracy of the obtained audio data.

[0006] In this context, embodiments of the present disclosure are expected to provide an improved method for generating audio data, an apparatus for generating audio data, a computer-readable storage medium, and an electronic device.

[0007] According to one aspect of the present disclosure, there is provided a method for generating audio data, including:

[0008] Extracting features from the original audio data to obtain the original melody skeleton features;

[0009] Determining the melody trend features of the original audio data according to the original melody skeleton features;

[0010] Evolving the features of the original melody skeleton features to obtain the target melody skeleton features;

[0011] Generate target audio data according to the described melody trend feature and target melody skeleton feature.

[0012] In an exemplary embodiment of the present disclosure, the original melody skeleton feature is a feature composed of the most core notes and / or the most stable notes in the original audio data; the notes included in the original melody skeleton feature include at least one of the following: main melody notes, harmonic support notes, key rhythm points, and the starting note and ending note of the target musical phrase.

[0013] In an exemplary embodiment of the present disclosure, feature extraction is performed on the original audio data to obtain the original melody skeleton feature, including: segmenting the original audio data based on a preset segment segmentation model to obtain one or more musical phrase segments; segmenting the musical phrase segments based on a preset beat segmentation model to obtain one or more current musical phrase beats; extracting sub-skeleton notes in the current musical phrase beats, and generating the original melody skeleton feature according to the sub-skeleton notes.

[0014] In an exemplary embodiment of the present disclosure, extracting sub-skeleton notes in the musical phrase beats includes: constructing a first musical phrase beat set according to the current musical phrase beats; traversing the first musical phrase beat set, selecting any current musical phrase beat as the target musical phrase beat, and extracting the current notes included in the target musical phrase beat to obtain a first note extraction result; when it is determined that the first note extraction result is not empty, extracting the current chord in-note of the target musical phrase beat from the first note extraction result to obtain a second note extraction result, and determining the sub-skeleton note of the target musical phrase beat according to the second note extraction result; sequentially repeating the extraction process of the sub-skeleton note of the target musical phrase beat to obtain the sub-skeleton notes of other current musical phrase beats in the first musical phrase beat set except the target musical phrase beat.

[0015] In an exemplary embodiment of the present disclosure, determining the sub-skeleton note of the target musical phrase beat according to the second note extraction result includes: when it is determined that the second note extraction result is empty, taking the first note in the first note extraction result as the sub-skeleton note of the target musical phrase beat; when it is determined that the second note extraction result is not empty, determining the number of notes of the current chord in-note in the second note extraction result, and when it is determined that the number of notes is one, taking the current chord in-note as the sub-skeleton note of the target musical phrase beat; when it is determined that the number of notes is multiple, determining the target chord in-note from the current chord in-notes according to the volume of the current chord in-notes, and taking the target chord in-note as the sub-skeleton note of the target musical phrase beat.

[0016] In an exemplary embodiment of the present disclosure, the preset segment segmentation model includes an embedding mapping layer, an encoding layer, and a mixture-of-experts model layer; wherein, segmenting the original audio data based on the preset segment segmentation model to obtain one or more musical phrase segments includes: obtaining audio attribute information of the original audio data; wherein the audio attribute information includes at least one of name information of the original audio data, audio creator information, and lyric information corresponding to the original audio data; generating basic information to be predicted according to the original note information of the original audio data, and generating context information to be predicted according to the audio attribute information and preset parameter prompt information; performing an embedding mapping process on the basic information to be predicted based on the embedding mapping layer to obtain a first audio feature, and performing an embedding mapping process on the context information to be predicted based on the embedding mapping layer to obtain a first context flag sequence; performing an encoding process on the first audio feature and the first context flag sequence based on the encoding layer to obtain a first overall context representation, and performing segment segmentation on the first context flag sequence and the first overall context representation based on the mixture-of-experts model layer to obtain one or more musical phrase segments.

[0017] In an exemplary embodiment of the present disclosure, the mixture-of-experts model layer includes a gated network model and multiple expert neural network models; wherein, performing segment segmentation on the first context flag sequence and the first overall context representation based on the mixture-of-experts model layer to obtain one or more musical phrase segments includes: determining a first model weight of the expert neural network model in the note structure dimension and a second model weight in the music style dimension based on the first context flag sequence by the gated network model; determining a first target neural network model required to perform the segment segmentation task in the note structure dimension and a second target neural network model required to perform the segment segmentation task in the music style dimension from the multiple expert neural network models based on the first model weight and the second model weight; inputting the first context flag sequence and the first overall context representation into the first target neural network model and the second target neural network model respectively to obtain a first segment segmentation result in the note structure dimension and a second segment segmentation result in the music style dimension, and generating the one or more musical phrase segments according to the first segment segmentation result and the second segment segmentation result.

[0018] In an exemplary embodiment of the present disclosure, determining the melody trend feature of the original audio data according to the original melody skeleton feature includes: classifying the original notes in the original audio data to obtain first in-key notes and first out-of-key notes, and classifying the original characteristic notes in the original melody skeleton feature to obtain second in-key notes and second out-of-key notes; determining the first in-key pitch of the first in-key notes and the first out-of-key pitch of the first out-of-key notes, and determining the second in-key pitch of the second in-key notes and the second out-of-key pitch of the second out-of-key notes; determining the in-key trend feature according to the first in-key pitch and the second in-key pitch, and determining the out-of-key trend feature according to the first out-of-key pitch and the second out-of-key pitch; determining the melody trend feature of the original audio data according to the in-key trend feature and the out-of-key trend feature.

[0019] In an exemplary embodiment of the present disclosure, determining the in-key trend feature according to the first in-key pitch and the second in-key pitch includes: determining a first pitch difference between the first in-key pitch and the second in-key pitch, and classifying the first pitch difference to obtain a first zero-degree difference, a first second-degree difference, a first second-degree adjacent difference, and a first high-degree difference; identifying the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference with preset characters to obtain a difference identification result, and determining the in-key trend feature according to the difference identification result.

[0020] In an exemplary embodiment of the present disclosure, identifying the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference with preset characters includes: extracting a first positive second-degree difference and a first negative second-degree difference from the first second-degree difference, and extracting a first positive second-degree adjacent difference and a first negative second-degree adjacent difference from the first second-degree adjacent difference; extracting a first positive high-degree difference and a first negative high-degree difference from the first high-degree difference, and respectively configuring character identifications for the first negative high-degree difference, the first negative second-degree adjacent difference, the first negative second-degree difference, the first zero-degree difference, the first positive second-degree difference, the first positive second-degree adjacent difference, and the first positive high-degree difference to obtain a first difference identification result, …, a seventh difference identification result; splicing the difference identification results corresponding to the first in-key notes in sequence according to the order of the first in-key notes in the original audio data to obtain the in-key trend feature.

[0021] In an exemplary embodiment of the present disclosure, evolving the original melody skeleton features to obtain target melody skeleton features includes: extracting original beats to be evolved from the current phrase beats of the original audio data according to the skeleton similarity value between the original audio data and the target audio data, and constructing a second phrase beat set according to the extracted original beats to be evolved; traversing the second phrase beat set, extracting any beat from the original beats to be evolved as the target beat to be evolved, and determining the original characteristic note corresponding to the target beat to be evolved from the original melody skeleton features; extracting the chord notes included in the target beat to be evolved, and determining the evolved characteristic note of the target beat to be evolved according to the chord notes and the original characteristic note; repeating the evolution process of the evolved characteristic note of the target beat to be evolved in sequence to obtain the evolved characteristic notes of other beats to be evolved in the second phrase beat set except the target beat to be evolved, and generating the target melody skeleton features according to the evolved characteristic notes.

[0022] In an exemplary embodiment of the present disclosure, determining the evolved characteristic note of the target beat to be evolved according to the chord notes and the original characteristic note includes: constructing an original evolved note set according to the chord notes and the original characteristic note, and calculating the interval relationship between the notes in the original evolved note set and the original characteristic notes of the current phrase beats adjacent to the target beat to be evolved; filtering the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set, and extracting any note from the target evolved note set as the evolved characteristic note of the target beat to be evolved.

[0023] In an exemplary embodiment of the present disclosure, filtering the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set includes: if the interval relationship is that the interval between the note in the original evolved note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is greater than or equal to the preset interval threshold, deleting the note; if the interval relationship is that the interval between the note in the original evolved note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is less than the preset interval threshold, retaining the note, and generating the target evolved note set according to the retained notes.

[0024] In an exemplary embodiment of the present disclosure, generating target audio data according to the melody trend feature and the target melody skeleton feature includes: determining the original rhythm type of the original audio data, and performing rhythm restoration on the target melody skeleton feature according to the original rhythm type to obtain a target rhythm type; performing pitch restoration on the target melody skeleton feature according to the melody trend feature to obtain a target melody pitch, and generating the target audio data according to the target rhythm type, the target melody pitch, and the target melody skeleton feature.

[0025] According to one aspect of the present disclosure, there is provided an apparatus for generating audio data, including:

[0026] An audio feature extraction module, configured to extract features from the original audio data to obtain an original melody skeleton feature;

[0027] A trend feature determination module, configured to determine the melody trend feature of the original audio data according to the original melody skeleton feature;

[0028] A feature evolution module, configured to perform feature evolution on the original melody skeleton feature to obtain a target melody skeleton feature;

[0029] An audio data generation module, configured to generate target audio data according to the melody trend feature and the target melody skeleton feature.

[0030] In an exemplary embodiment of the present disclosure, the original melody skeleton feature is a feature composed of the most core notes and / or the most stable notes in the original audio data; the notes included in the original melody skeleton feature include at least one of the following: main melody notes, harmonic support notes, key rhythm points, and the starting note and ending note of the target musical phrase.

[0031] In an exemplary embodiment of the present disclosure, extracting features from the original audio data to obtain an original melody skeleton feature includes: segmenting the original audio data based on a preset segment segmentation model to obtain one or more musical phrase segments; segmenting the musical phrase segments based on a preset beat segmentation model to obtain one or more current musical phrase beats; extracting sub-skeleton notes from the current musical phrase beats, and generating an original melody skeleton feature according to the sub-skeleton notes.

[0032] In an exemplary embodiment of the present disclosure, extracting sub-skeleton notes from the phrase beat includes: constructing a first phrase beat set according to the current phrase beat; traversing the first phrase beat set, selecting any current phrase beat as the target phrase beat, and extracting the current notes included in the target phrase beat to obtain a first note extraction result; when it is determined that the first note extraction result is not empty, extracting the in-chord notes of the target phrase beat from the first note extraction result to obtain a second note extraction result, and determining the sub-skeleton notes of the target phrase beat according to the second note extraction result; sequentially repeating the extraction process of the sub-skeleton notes of the target phrase beat to obtain the sub-skeleton notes of other current phrase beats in the first phrase beat set except the target phrase beat.

[0033] In an exemplary embodiment of the present disclosure, determining the sub-skeleton notes of the target phrase beat according to the second note extraction result includes: when it is determined that the second note extraction result is empty, taking the first note in the first note extraction result as the sub-skeleton notes of the target phrase beat; when it is determined that the second note extraction result is not empty, determining the number of in-chord notes in the second note extraction result, and when it is determined that the number of notes is one, taking this in-chord note as the sub-skeleton notes of the target phrase beat; when it is determined that the number of notes is multiple, determining the target in-chord note from the in-chord notes according to the volume of the in-chord notes, and taking this target in-chord note as the sub-skeleton notes of the target phrase beat.

[0034] In an exemplary embodiment of the present disclosure, the preset segment segmentation model includes an embedding mapping layer, an encoding layer, and a mixture of experts model layer; wherein, performing segment segmentation on the original audio data based on the preset segment segmentation model to obtain one or more phrase segments includes: obtaining the audio attribute information of the original audio data; wherein, the audio attribute information includes at least one of the name information of the original audio data, the audio creator information, and the lyric information corresponding to the original audio data; generating the basic information to be predicted according to the original note information of the original audio data, and generating the context information to be predicted according to the audio attribute information and the preset parameter prompt information; performing embedding mapping processing on the basic information to be predicted based on the embedding mapping layer to obtain a first audio feature, and performing embedding mapping processing on the context information to be predicted based on the embedding mapping layer to obtain a first context flag sequence; performing encoding processing on the first audio feature and the first context flag sequence based on the encoding layer to obtain a first context overall representation, and performing segment segmentation on the first context flag sequence and the first context overall representation based on the mixture of experts model layer to obtain one or more phrase segments.

[0035] In an exemplary embodiment of the present disclosure, the mixture-of-experts model layer includes a gating network model and a plurality of expert neural network models; wherein, based on the mixture-of-experts model layer, segmenting the first context flag sequence and the first overall context representation to obtain one or more musical phrase segments, including: based on the gating network model, according to the first context flag sequence, determining a first model weight of the expert neural network model in the dimension of note structure and a second model weight in the dimension of musical style; based on the first model weight and the second model weight, determining a first target neural network model required to perform the segmenting task in the dimension of note structure and a second target neural network model required to perform the segmenting task in the dimension of musical style from the plurality of expert neural network models; inputting the first context flag sequence and the first overall context representation into the first target neural network model and the second target neural network model respectively, obtaining a first segmenting result in the dimension of note structure and a second segmenting result in the dimension of musical style, and generating one or more musical phrase segments according to the first segmenting result and the second segmenting result.

[0036] In an exemplary embodiment of the present disclosure, determining the melody trend feature of the original audio data according to the original melody skeleton feature includes: classifying the original notes in the original audio data to obtain first in-key notes and first out-of-key notes, and classifying the original feature notes in the original melody skeleton feature to obtain second in-key notes and second out-of-key notes; determining a first in-key pitch of the first in-key notes and a first out-of-key pitch of the first out-of-key notes, and determining a second in-key pitch of the second in-key notes and a second out-of-key pitch of the second out-of-key notes; determining an in-key trend feature according to the first in-key pitch and the second in-key pitch, and determining an out-of-key trend feature according to the first out-of-key pitch and the second out-of-key pitch; determining the melody trend feature of the original audio data according to the in-key trend feature and the out-of-key trend feature.

[0037] In an exemplary embodiment of the present disclosure, determining the in-key trend feature according to the first in-key pitch and the second in-key pitch includes: determining a first pitch difference between the first in-key pitch and the second in-key pitch, and classifying the first pitch difference to obtain a first zero-degree difference, a first second-degree difference, a first second-degree adjacent difference, and a first high-degree difference; identifying the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference with preset characters to obtain a difference identification result, and determining the in-key trend feature according to the difference identification result.

[0038] In an exemplary embodiment of the present disclosure, the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first height difference are identified with preset characters, including: extracting a first positive second-degree difference and a first negative second-degree difference from the first second-degree difference, and extracting a first positive second-degree adjacent difference and a first negative second-degree adjacent difference from the first second-degree adjacent difference; extracting a first positive height difference and a first negative height difference from the first height difference, and respectively configuring character identifiers for the first negative height difference, the first negative second-degree adjacent difference, the first negative second-degree difference, the first zero-degree difference, the first positive second-degree difference, the first positive second-degree adjacent difference, and the first positive height difference to obtain a first difference identification result, …, a seventh difference identification result; splicing the difference identification results corresponding to the first in-key notes in sequence according to the order of the first in-key notes in the original audio data to obtain the in-key trend feature.

[0039] In an exemplary embodiment of the present disclosure, evolving the original melody skeleton feature to obtain a target melody skeleton feature includes: extracting original beats to be evolved from the current phrase beats of the original audio data according to the skeleton similarity value between the original audio data and the target audio data, and constructing a second phrase beat set according to the extracted original beats to be evolved; traversing the second phrase beat set, extracting any beat from the original beats to be evolved as a target beat to be evolved, and determining an original feature note corresponding to the target beat to be evolved from the original melody skeleton feature; extracting chord notes included in the target beat to be evolved, and determining an evolved feature note of the target beat to be evolved according to the chord notes and the original feature note; sequentially repeating the evolution process of the evolved feature note of the target beat to be evolved to obtain the evolved feature notes of other beats to be evolved in the second phrase beat set except the target beat to be evolved, and generating the target melody skeleton feature according to the evolved feature notes.

[0040] In an exemplary embodiment of the present disclosure, determining the evolved feature note of the target beat to be evolved according to the chord notes and the original feature note includes: constructing an original evolved note set according to the chord notes and the original feature note, and calculating the interval relationship between the notes in the original evolved note set and the original feature notes of the current phrase beats adjacent to the target beat to be evolved; filtering the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set, and extracting any note from the target evolved note set as the evolved feature note of the target beat to be evolved.

[0041] In an exemplary embodiment of the present disclosure, filtering the notes in the original evolving note set according to the interval relationship to obtain a target evolving note set, including: if the interval between a note in the original evolving note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is greater than or equal to a preset interval threshold, deleting the note; if the interval between a note in the original evolving note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is less than the preset interval threshold, retaining the note, and generating the target evolving note set according to the retained notes.

[0042] In an exemplary embodiment of the present disclosure, generating target audio data according to the melody trend feature and the target melody skeleton feature, including: determining the original rhythm type of the original audio data, and performing rhythm restoration on the target melody skeleton feature according to the original rhythm type to obtain a target rhythm type; performing pitch restoration on the target melody skeleton feature according to the melody trend feature to obtain a target melody pitch, and generating the target audio data according to the target rhythm type, the target melody pitch, and the target melody skeleton feature.

[0043] According to one aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for generating audio data according to any one of the foregoing exemplary embodiments is implemented.

[0044] According to one aspect of the present disclosure, there is provided an electronic device, including:

[0045] a processor; and

[0046] a memory for storing executable instructions of the processor;

[0047] wherein the processor is configured to execute the method for generating audio data according to any one of the foregoing exemplary embodiments by executing the executable instructions.

[0048] The method for audio data and the audio data generation device according to the embodiments of the present disclosure can extract features from the original audio data to obtain the original melody skeleton features; determine the melody trend features of the original audio data according to the original melody skeleton features; then perform feature evolution on the original melody skeleton features to obtain the target melody skeleton features; and finally generate the target audio data according to the melody trend features and the target melody skeleton features, without generating audio data by using machine learning, thereby significantly reducing the problem of low accuracy of the audio data obtained by using machine learning, and reducing the problem of single style of the audio data obtained by imitating the melody according to music theory rules, bringing a better experience to users. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0050] Figure 1 Schematically shows a flowchart of a method for generating audio data according to an exemplary embodiment of the present disclosure;

[0051] Figure 2 Schematically shows a structural example diagram of a preset segment segmentation model according to an exemplary embodiment of the present disclosure;

[0052] Figure 3 Schematically shows a structural example diagram of a mixture of experts model layer according to an exemplary embodiment of the present disclosure;

[0053] Figure 4 Schematically shows a scenario example diagram of features of an original audio data and the corresponding original melody skeleton features according to an exemplary embodiment of the present disclosure;

[0054] Figure 5 Schematically shows a scenario example diagram of the melody trend features corresponding to the original melody skeleton features according to an exemplary embodiment of the present disclosure;

[0055] Figure 6 Schematically shows a scenario example diagram of the target melody skeleton features corresponding to the original melody skeleton features according to an exemplary embodiment of the present disclosure;

[0056] Figure 7 Schematically shows a scenario example diagram of evolving a near-related rhythm type based on the original rhythm type according to an exemplary embodiment of the present disclosure;

[0057] Figure 8Schematically shows a melody graph of a target rhythm type obtained based on the target melody skeleton features according to an exemplary embodiment of the present disclosure;

[0058] Figure 9 Schematically shows an example graph of a target melody pitch obtained based on the target melody skeleton features and melody trend features according to an exemplary embodiment of the present disclosure;

[0059] Figure 10 Schematically shows a block diagram of a generating device for audio data according to an exemplary embodiment of the present disclosure;

[0060] Figure 11 Schematically shows a computer-readable storage medium for storing a method for generating audio data according to an exemplary embodiment of the present disclosure;

[0061] Figure 12 Schematically shows an electronic device for implementing a method for generating audio data according to an exemplary embodiment of the present disclosure.

[0062] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Embodiments

[0063] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present disclosure, rather than to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0064] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0065] According to the embodiments of the present disclosure, a method for generating audio data, a generating device for audio data, a computer-readable storage medium, and an electronic device are provided.

[0066] In this document, the number of any element in the drawings is used for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0067] Next, with reference to several representative embodiments of the present disclosure, the principles and spirit of the present disclosure will be elaborated in detail. Summary of the Invention

[0069] The applicant has found that in the process of music creation using artificial intelligence, many technologies can achieve intelligent music creation; for example, Generative Adversarial Networks (GAN) and Long Short-Term Memory (LSTM) models can be used. In the actual application process, whether it is a generative adversarial network or a long short-term memory network, they can learn the characteristics of a large amount of music data by themselves, and then can create music works without the input of music information. Further, in the process of generating music works using models such as LSTM, a large amount of music MIDI (Musical Instrument Digital Interface) data is first obtained, and the MIDI data contains music note pitch information, etc.; then through data preprocessing, it is output to the model for model training; finally, after the training is completed, the output result is the generated music work. At the same time, as the dataset increases, the music features that the machine learning model can learn will be more and more, and the music it generates will be more in line with people's aesthetics; however, the music generated by the machine learning model will be blind; for example, when users have more specific requirements for the melody trend of the generated song, the uncontrollability of machine learning will be reflected, so it is difficult to meet higher customization requirements; such as song style transfer, etc., which further makes the accuracy of the obtained song lower.

[0070] In addition, there is also a style imitation melody generation algorithm based on music theory rules; specifically, this algorithm generates a melody that conforms to the original song structure under a specified chord sequence by analyzing the chord tones and end tones of a song set, so as to generate an imitation song melody with a refined style; however, the style of the song obtained based on this method is relatively single.

[0071] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure will be specifically introduced below.

[0072] Exemplary Method

[0073] The exemplary embodiment of the present disclosure first provides a method for generating audio data, which can run on a terminal device, a server, a cloud server, or a server cluster, etc.; of course, those skilled in the art can also run the method of the present disclosure on other platforms according to needs, and no special limitation is made in this exemplary embodiment. Specifically, referring to Figure 1 as shown, the method for generating audio data may include the following steps:

[0074] Step S110. Extract features from the original audio data to obtain the original melody skeleton features;

[0075] Step S120. Determine the melody trend feature of the original audio data according to the original melody skeleton feature;

[0076] Step S130. Perform feature evolution on the original melody skeleton feature to obtain the target melody skeleton feature;

[0077] Step S140. Generate the target audio data according to the melody trend feature and the target melody skeleton feature.

[0078] In the above-described method for generating audio data, the original audio data can be subjected to feature extraction to obtain the original melody skeleton feature; and according to the original melody skeleton feature, the melody trend feature of the original audio data can be determined; then the original melody skeleton feature is subjected to feature evolution to obtain the target melody skeleton feature; finally, the target audio data is generated according to the melody trend feature and the target melody skeleton feature, without the need to generate audio data in a machine learning manner, thereby significantly reducing the problem of low accuracy of the audio data obtained by using the machine learning method, and reducing the problem of single style of the audio data obtained by imitating the melody in the style of music theory rules, bringing a better experience to the user.

[0079] Hereinafter, the method for generating audio data described in the exemplary embodiments of the present disclosure will be further explained and described with reference to the accompanying drawings.

[0080] First, the proper nouns involved in the exemplary embodiments of the present disclosure will be explained and described.

[0081] Music theory rules: Music theory rules are the fundamental part of music theory, mainly including aspects such as music notation, intervals, chords, rhythm, meter, etc. Mastering music theory rules is of great significance for composition, arrangement, conducting, and performance. Among them, basic music theory concepts can include but are not limited to timbre, notes, rhythm, mode, chords, scales, key signatures, and meter, etc. Among them, the timbre recorded here refers to the perceptual characteristics of sound, which can be determined by the vibration mode of the sounding body and the combination of overtones; in the actual application process, human voice timbre can be divided into high, medium, and low pitches, while instrumental timbre varies according to different instrument types. The notes recorded here are the basic elements in music, which can include whole notes, half notes, quarter notes, eighth notes, etc.; among them, an eighth note consists of a note head, a stem, and a tail, and the durations of two eighth notes added together equal the duration of a quarter note. The rhythm recorded here refers to the combination of the length and intensity of notes in music and is one of the basic structures of music; in the actual application process, mastering a sense of rhythm is crucial for playing and creating music. The mode recorded here can be a scale system formed by arranging a series of notes according to certain rules; among them, common modes include major and minor, which have different scale structures and melodic characteristics. The chords recorded here are the sound effects formed by superimposing three or more notes according to certain interval relationships; among them, the types and functions of chords play an important role in music, and common chords include major chords, minor chords, etc. The scale recorded here refers to a series of pitches arranged in a specific order; among them, the basic structure of a major scale can be "whole - whole - half - whole - whole - whole - half", and the minor scale can be "whole - half - whole - whole - whole - half - whole". The key signature recorded here can be used to mark the key of a piece of music. For example, 1 = C means singing 1 as high as the C in the first leger line above middle C on the keyboard. The meter recorded here refers to the periodic repetition of accents and non - accents in music, and common meters include 2 / 4 time and 4 / 4 time, etc.

[0082] Secondly, the technical implementation principles of the exemplary embodiments of the present disclosure are explained and described. Specifically, the method for generating audio data recorded in the exemplary embodiments of the present disclosure mainly focuses on how to use the melody skeleton to achieve melody creation for imitating song styles; in the actual application process, first, the melody of the song needs to be extracted to obtain a rough-grained melody curve, then the evolution of the melody skeleton is realized on the premise of maintaining the same song style, and finally, based on the new melody skeleton, the melody is reversely restored according to the trajectory of the melody skeleton extraction to achieve melody generation. Among them, melody skeleton extraction can be used to abstractly extract the melody line composed of complex and variable melody notes to obtain a low-resolution melody curve that maintains the original style; melody trend calculation can be used to calculate the melody trend presented after skeleton extraction to facilitate style restoration for "melody generation based on melody skeleton"; melody skeleton evolution can be used to evolve the melody skeleton according to the summarized music theory rules without affecting the song style to achieve the generation of a new melody skeleton; melody generation based on melody skeleton and melody trend can be used to reversely restore the melody skeleton according to the calculated melody trend to achieve melody generation.

[0083] Further, the preset segment segmentation model involved in the exemplary embodiments of the present disclosure is explained and described. Specifically, referring to Figure 2 as shown, the preset segment segmentation model may include an input layer 201, an embedding mapping layer 202, an encoding layer 203, a mixture of experts model layer 204, and an output layer 205; among them, the mixture of experts model layer described here includes a gating network model and multiple expert neural network models, and the specific structure can be referred to Figure 3 as shown. Further, the functions of each model layer in the segment segmentation process will be described in detail later, and will not be elaborated further here.

[0084] Next, in combination with Figure 2 and Figure 3 for Figure 1 the method for generating audio data shown in is further explained and described. Specifically:

[0085] In step S110, the original audio data is feature-extracted to obtain the original melody skeleton feature.

[0086] Specifically, the original audio data recorded herein refers to data that can identify beat information, phrase information, chord information, and melody information; the original audio data recorded herein can include multiple different data formats, such as, but not limited to, MP3, WAV, FLAC, AAC, AIFF, WMA, etc., and this example does not make special restrictions on this; the original melody skeleton feature recorded herein can also be referred to as the original melody skeleton tone, which refers to the feature composed of the most core notes and / or the most stable notes in the original audio data; that is, the original melody skeleton tone refers to the most core and most stable notes in the original audio data; the most core and most stable notes recorded herein can include, but not limited to, main melody notes, harmony support notes, key rhythm points, and the starting and ending notes of the target phrase, etc.; among them, the original audio data recorded herein can refer to Figure 4 shown in 401 of Figure 4 shown in 402 of . On this premise, if it is necessary to extract the original melody skeleton feature from the original audio data, it can be achieved based on the following method: segment the original audio data based on a preset segment segmentation model to obtain one or more phrase segments; segment the phrase segments based on a preset beat segmentation model to obtain one or more current phrase beats; extract the sub-skeleton notes in the current phrase beats, and generate the original melody skeleton feature according to the sub-skeleton notes.

[0087] In an exemplary embodiment, segmenting the original audio data based on a preset segment segmentation model to obtain one or more musical phrase segments can be achieved in the following manner: obtaining audio attribute information of the original audio data; wherein the audio attribute information includes at least one of the name information of the original audio data, the audio creator information, and the lyric information corresponding to the original audio data; generating basic information to be predicted according to the original note information of the original audio data, and generating context information to be predicted according to the audio attribute information and preset parameter prompt information; performing embedding mapping processing on the basic information to be predicted based on an embedding mapping layer to obtain a first audio feature, and performing embedding mapping processing on the context information to be predicted based on the embedding mapping layer to obtain a first context flag sequence; performing encoding processing on the first audio feature and the first context flag sequence based on the encoding layer to obtain a first context overall representation, and performing segment segmentation on the first context flag sequence and the first context overall representation based on the mixture of experts model layer to obtain one or more musical phrase segments. That is, in the process of actual application, in order to obtain the original melody skeleton feature, first, the melody (i.e., the original audio data) needs to be divided into multiple musical phrase segments in units of musical phrases; further, the above-mentioned embedding mapping layer may include an Embedding embedding mapping layer and a Bert embedding mapping layer. In the process of actual application, the embedding mapping processing can be performed on the basic information to be predicted based on the Embedding embedding mapping layer to obtain a first audio feature, and the embedding mapping processing can be performed on the context information to be predicted based on the Bert embedding mapping layer to obtain a first context flag sequence.

[0088] In an exemplary embodiment of the present disclosure, based on the mixture-of-experts model layer, segmenting the first context flag sequence and the first overall context representation to obtain one or more musical phrase segments, including: determining, by the gated network model according to the first context flag sequence, a first model weight of the expert neural network model in the note structure dimension and a second model weight in the musical style dimension; determining, based on the first model weight and the second model weight, a first target neural network model required to perform the segmenting task in the note structure dimension and a second target neural network model required to perform the segmenting task in the musical style dimension from the multiple expert neural network models; inputting the first context flag sequence and the first overall context representation into the first target neural network model and the second target neural network model respectively to obtain a first segmenting result in the note structure dimension and a second segmenting result in the musical style dimension, and generating the one or more musical phrase segments according to the first segmenting result and the second segmenting result. Specifically, in the actual application process, the division of musical phrase segments can depend on the structural features of notes and the overall musical style. Therefore, the division of musical phrase segments is limited from two different dimensions here; at the same time, the structural features recorded here can include, but are not limited to, the regularity of rhythm, the change of harmony, and the trend of melody, etc. Further, in the process of generating one or more musical phrase segments according to the first segmenting result and the second segmenting result, the musical phrase segments can be directly obtained by combining the first segmenting result and the second segmenting result, or the musical phrase segments can be obtained by weighted summing the first segmenting result and the second segmenting result. This example does not make special restrictions on this.

[0089] In an exemplary embodiment, the specific model structure of the preset beat segmentation model described above is similar to the model structure of the preset segmenting model, and no further limitation is made here; on this premise, based on the preset beat segmentation model, performing beat segmentation on the musical phrase segments to obtain one or more current musical phrase beats can be achieved by the following method: generating to-be-predicted segment information according to the musical phrase segments and the preset beat segmentation parameter hint information to generate to-be-predicted beat context information; performing embedding mapping processing on the to-be-predicted segment information by the embedding mapping layer to obtain a second audio feature, and performing embedding mapping processing on the to-be-predicted beat context information by the embedding mapping layer to obtain a second context flag sequence; performing encoding processing on the second audio feature and the second context flag sequence by the encoding layer to obtain a second overall context representation, and performing beat segmentation on the second context flag sequence and the second overall context representation by the mixture-of-experts model layer to obtain one or more current musical phrase beats.

[0090] In an exemplary embodiment, extracting sub-skeleton notes from the beat of a musical phrase can be achieved in the following manner: constructing a first musical phrase beat set according to the current musical phrase beat; traversing the first musical phrase beat set, selecting any current musical phrase beat as the target musical phrase beat, and extracting the current note included in the target musical phrase beat to obtain a first note extraction result; when it is determined that the first note extraction result is not empty, extracting the in-chord notes of the target musical phrase beat from the first note extraction result to obtain a second note extraction result, and determining the sub-skeleton notes of the target musical phrase beat according to the second note extraction result; sequentially repeating the extraction process of the sub-skeleton notes of the target musical phrase beat to obtain the sub-skeleton notes of other current musical phrase beats in the first musical phrase beat set except the target musical phrase beat.

[0091] In an exemplary embodiment, determining the sub-skeleton notes of the target musical phrase beat according to the second note extraction result can be achieved in the following manner: when it is determined that the second note extraction result is empty, taking the first note in the first note extraction result as the sub-skeleton note of the target musical phrase beat; when it is determined that the second note extraction result is not empty, determining the number of notes of the in-chord notes in the second note extraction result, and when it is determined that the number of notes is one, taking this in-chord note as the sub-skeleton note of the target musical phrase beat; when it is determined that the number of notes is multiple, determining the target in-chord note from the in-chord notes according to the volume of the in-chord notes, and taking this target in-chord note as the sub-skeleton note of the target musical phrase beat.

[0092] Hereinafter, the specific extraction process of the sub-skeleton notes in the obtained musical phrase beats will be further explained and described. Specifically, first, taking a single musical phrase beat as a unit, identifying the melodic skeleton notes beat by beat (that is, determining the sub-skeleton notes beat by beat); secondly, judging whether there are notes in this beat; if there are no notes in this beat, then there is no melodic skeleton note in this beat; otherwise, identifying the in-chord notes included in the current beat (based on the chord names marked in the original audio data; and, if only ordinary chords are marked, it is considered a triad); further, if there are multiple in-chord notes in this beat, selecting the highest note among the in-chord notes as the skeleton note; if there are no in-chord notes, taking the first note as the skeleton note.

[0093] In step S120, according to the original melodic skeleton feature, determine the melodic trend feature of the original audio data.

[0094] Specifically, the specific determination process of the melody trend can be achieved through the following method: classify the original notes in the original audio data to obtain the in-key notes of the first key and the out-of-key notes of the first key, and classify the original characteristic notes in the original melody skeleton features to obtain the in-key notes of the second key and the out-of-key notes of the second key; determine the in-key pitch of the in-key notes of the first key and the out-of-key pitch of the out-of-key notes of the first key, and determine the in-key pitch of the in-key notes of the second key and the out-of-key pitch of the out-of-key notes of the second key; determine the in-key trend feature according to the in-key pitch of the first key and the in-key pitch of the second key, and determine the out-of-key trend feature according to the out-of-key pitch of the first key and the out-of-key pitch of the second key; determine the melody trend feature of the original audio data according to the in-key trend feature and the out-of-key trend feature. Among them, the in-key notes recorded here refer to the notes within the scale in a specific key; for example, in C major, C, D, E, F, G, A, B are the natural notes in the natural scale and are all in-key notes of C major. These notes occupy fixed positions in the scale and form specific interval relationships with other notes; the out-of-key notes recorded here refer to the notes that do not belong to the current key scale in a musical work; specifically, the out-of-key notes refer to those notes that are not in the current key scale. These notes are outside the scale and are therefore called out-of-key notes; for example, in C major, B is not a constituent note of C major and is therefore considered an out-of-key note; further, in the process of classifying the in-key notes and the out-of-key notes, since the input original audio data includes the corresponding pitch information, the corresponding in-key notes and out-of-key notes can be directly determined; moreover, the in-key pitch of the in-key notes and the out-of-key pitch of the out-of-key notes can also be directly determined based on the original audio data.

[0095] In an exemplary embodiment, determining the in-key trend feature according to the in-key pitch of the first key and the in-key pitch of the second key can be achieved through the following method: determine the first pitch difference between the in-key pitch of the first key and the in-key pitch of the second key, and classify the first pitch difference to obtain the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference; identify the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference with preset characters to obtain the difference identification result, and determine the in-key trend feature according to the difference identification result.

[0096] In an exemplary embodiment of the present disclosure, the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first height difference are identified with preset characters, including: extracting a first positive second-degree difference and a first negative second-degree difference from the first second-degree difference, and extracting a first positive second-degree adjacent difference and a first negative second-degree adjacent difference from the first second-degree adjacent difference; extracting a first positive height difference and a first negative height difference from the first height difference, and respectively configuring character identifiers for the first negative height difference, the first negative second-degree adjacent difference, the first negative second-degree difference, the first zero-degree difference, the first positive second-degree difference, the first positive second-degree adjacent difference, and the first positive height difference to obtain a first difference identification result, …, a seventh difference identification result; splicing the difference identification results corresponding to the first in-tune notes in sequence according to the order of the first in-tune notes in the original audio data to obtain the in-tune trend feature.

[0097] Hereinafter, the specific determination process of the in-tune trend feature will be further explained and described. Specifically, in the actual application process, the corresponding melody trend feature can be determined according to the difference between the original melody pitch and the melody skeleton pitch; further, since notes can be divided into in-tune notes and out-of-tune notes, in the process of calculating the melody trend feature, it is first necessary to divide the original notes in the original audio data and the original feature notes in the original melody skeleton feature into in-tune notes and out-of-tune notes, and then calculate the in-tune trend feature and the out-of-tune trend feature respectively.

[0098] In an exemplary embodiment, the specific calculation process of the in-key trend feature is as follows: First, define a set of unknowns for the pitch differences of in-key tones, from low to high in sequence: "a (first negative height difference)", "b (first negative second-degree difference)", "-1 (first negative second-degree adjacent difference)", "0 (first zero-degree difference)", "1 (first positive second-degree adjacent difference)", "x (first positive second-degree difference)", "y (first positive height difference)"; where. The "0" recorded here is the pitch of the melodic skeleton tone, that is, the pitch with a first pitch difference of 0 between the first in-key pitch and the second in-key pitch; "1" and "-1" are the pitches that are 1 second higher or lower than the melodic skeleton tone (that is, the pitches with a first pitch difference of one second between the first in-key pitch and the second in-key pitch); among them, whether it is a major second or a minor second specifically depends on the specific situation of the scale; for example, in the major scale "do re mi fa so la xi", if the skeleton tone is "re", then "1" refers to "mi" and "-1" refers to "do"; if the skeleton tone is "mi", then "1" refers to "fa" and "-1" refers to "re"; and the first-level unknowns "b" and "x" refer to the in-key tones greater than a second that are closest to the second-degree tone within the skeleton unit; further, the second-level unknowns "a" and "y" respectively refer to all in-key tones within the skeleton unit that are lower than "b" and higher than "x". For example, within a certain skeleton unit, the melodic skeleton tone is "so", and the melodic notes are "fa so la xi", then the melodic skeleton extraction trajectory will be calculated as "-1 0 1x"; if the melodic notes are "re mi fa xi", then the melodic skeleton extraction trajectory will be calculated as "a b -1 x".

[0099] In an exemplary embodiment, the out-of-key trend feature can be determined according to the first out-of-key pitch and the second out-of-key pitch, which can be achieved in the following way: Determine the second pitch difference between the first out-of-key pitch and the second out-of-key pitch, and classify the second pitch difference to obtain a second zero-degree difference, a second second-degree difference, and a second height difference; identify the second zero-degree difference, the second second-degree difference, and the second height difference with preset characters to obtain an out-of-key difference identification result, and determine the out-of-key trend feature according to the out-of-key difference identification result.

[0100] In an exemplary embodiment of the present disclosure, the second zero-degree difference, the second second-degree difference, and the second height difference are identified with preset characters to obtain an out-of-tune difference identification result, and an out-of-tune trend feature is determined according to the out-of-tune difference identification result, including: extracting a second positive second-degree difference (χ) and a second negative second-degree difference (β) from the second second-degree difference, and extracting a second positive height difference (ψ) and a second negative height difference (α) from the second height difference, and respectively configuring character identifiers for the second negative height difference, the second negative second-degree difference, the second zero-degree difference (0), the second positive second-degree difference, and the second positive height difference to obtain a first out-of-tune difference identification result, …, a fifth out-of-tune difference identification result; splicing the difference identification results corresponding to the first out-of-tune notes in sequence according to the order of the first out-of-tune notes in the original audio data to obtain an out-of-tune trend feature.

[0101] Hereinafter, the specific generation process of the out-of-tune trend feature will be further explained and described. Specifically, in the actual application process, if an out-of-tune note appears within a skeleton unit, an out-of-tune note pitch difference unknown set is defined, which from low to high is: "α", "β", "0", "χ", "ψ". Among them, "0" is still the pitch of the melody skeleton note, that is, the pitch with a second pitch difference of 0 between the first out-of-tune pitch and the second out-of-tune pitch; "β" and "χ" refer to the out-of-tune note closest to the skeleton note that appears within this skeleton unit (that is, the pitch with a second pitch difference of a second degree between the first out-of-tune pitch and the second out-of-tune pitch), that is, the first-level unknowns. "α" and "ψ" respectively refer to all out-of-tune notes that are lower than "β" and higher than "χ" that appear within this skeleton unit, that is, the second-level unknowns. For example, in the major scale "do re mi fa so la xi", if within a certain skeleton unit, the melody skeleton note is "so" and the melody notes are "#re #fa so #la", then the melody skeleton extraction trajectory will be calculated as "αβ0χ". For example, according to Figure 4 the original melody skeleton features extracted by 402 in Figure 5 the melody trend at the position shown in 501 in

[0102] In step S130, the original melody skeleton feature is subjected to feature evolution to obtain a target melody skeleton feature.

[0103] Specifically, the specific generation process of the target melody skeleton features can be implemented in the following manner: Based on the skeleton similarity value between the original audio data and the target audio data, extract the original beats to be evolved from the current phrase beats of the original audio data, and construct a second phrase beat set according to the extracted original beats to be evolved; traverse the second phrase beat set, extract any beat from the original beats to be evolved as the target beat to be evolved, and determine the original characteristic notes corresponding to the target beat to be evolved from the original melody skeleton features; extract the chord notes included in the target beat to be evolved, and determine the evolved characteristic notes of the target beat to be evolved according to the chord notes and the original characteristic notes; sequentially repeat the evolution process of the evolved characteristic notes of the target beat to be evolved to obtain the evolved characteristic notes of the other beats to be evolved in the second phrase beat set except the target beat to be evolved, and generate the target melody skeleton features according to the evolved characteristic notes. Specifically, the skeleton similarity recorded here can be determined according to actual needs; for example, if the skeleton similarity value is 100%, no evolution is performed; if the skeleton similarity value is 0%, all melody skeleton notes are evolved; if the skeleton similarity value is greater than 0 and less than 1 (that is, between 100% - 0%), the extracted melody skeleton notes are evolved by probability extraction.

[0104] In one exemplary embodiment, determining the evolved characteristic notes of the target beat to be evolved according to the chord notes and the original characteristic notes can be implemented in the following manner: Construct an original evolved note set according to the chord notes and the original characteristic notes, and calculate the interval relationship between the notes in the original evolved note set and the original characteristic notes of the current phrase beat adjacent to the target beat to be evolved; filter the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set, and extract any note from the target evolved note set as the evolved characteristic note of the target beat to be evolved. Among them, the interval relationship (Interval) recorded here refers to the mutual relationship in pitch between the notes in the original evolved note set and the original characteristic notes, that is, the high - low distance between the pitches of two notes, and its unit is "degree"; for example, from C to D is a second - degree interval, and from C to G is a fifth - degree interval.

[0105] In an exemplary embodiment, filtering the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set can be achieved in the following manner: If the interval between a note in the original evolved note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is greater than or equal to a preset interval threshold, then delete the note; if the interval between a note in the original evolved note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is less than the preset interval threshold, then retain the note, and generate the target evolved note set based on the retained notes.

[0106] Hereinafter, the specific determination process of the target melody skeleton feature will be further explained and described. Specifically, before evolving the original melody skeleton feature, it is necessary to first confirm the skeleton similarity value (i.e., the skeleton similarity value) between this evolved creation and the original piece; among them, the skeleton similarity value recorded here can be input by the user manually or by voice; further, if the input value is 100%, no evolution is performed; if the input value is 0%, all melody skeleton notes are evolved; if the input value is greater than 0 and less than 1, the extracted melody skeleton notes are evolved by probability extraction; furthermore, in the actual specific determination process of the target melody skeleton feature, first, the original beat to be evolved needs to be extracted from the current phrase beat according to the skeleton similarity value and the target beat to be evolved is determined; then, the original characteristic note corresponding to the target beat to be evolved and the chord notes included in the target beat to be evolved are determined; finally, the evolved characteristic note of the target beat to be evolved is determined based on the chord notes and the original characteristic note; specifically, the specific determination process of the evolved characteristic note is: taking the melody skeleton note (i.e.,) as an element of the "variable skeleton note set" (abbreviation "set S"), identifying all chord note increments within this skeleton unit as elements of set S (i.e., constructing set S according to the original characteristic note and the chord note); then, in set S, the interval relationship between each adjacent front and rear skeleton note (if there is no adjacent skeleton note, there is no need to judge) is judged one by one (i.e., calculating the interval relationship between the notes in set S and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved), and the skeleton note with an interval greater than a fifth (i.e., the preset interval threshold, this value can be adjusted according to the specific business scenario, the purpose is to reduce the occurrence probability of large skip intervals) is removed from set S to obtain the target evolved note set S'; further, if the target evolved note set S' is an empty set, then this melody skeleton note does not evolve; otherwise, in the target evolved note set S', any note element is randomly selected with equal probability as the evolved melody skeleton note; at the same time, if the target evolved note set S' only includes one note element, then directly use this note element as the evolved characteristic note. Among them, according toFigure 4 The original melody skeleton features extracted by 402 in Figure 6 can be used to calculate the target melody skeleton features shown in 601 in

[0107] In step S140, according to the melody trend features and the target melody skeleton features, target audio data is generated.

[0108] Specifically, the specific determination process of the target audio data can be achieved in the following way: determine the original rhythm type of the original audio data, and perform rhythm restoration on the target melody skeleton features according to the original rhythm type to obtain the target rhythm type; perform pitch restoration on the target melody skeleton features according to the melody trend features to obtain the target melody pitch, and generate the target audio data according to the target rhythm type, target melody pitch and target melody skeleton features. Specifically, in the actual application process, perform rhythm restoration according to the rhythm pattern of the original song (that is, the original rhythm type of the original audio data); at the same time, in the process of rhythm restoration, it can also be appropriately evolved into a closely related rhythm type corresponding to the original rhythm type; among them, the closely related rhythm types recorded here can be as shown in 7; in Figure 7 the skeleton evolution chain shown, the elements adjacent to the elements in the chain can be considered as closely related rhythm patterns; that is to say, in the process of rhythm restoration, the target melody skeleton features can be directly restored according to the original rhythm type to obtain the target rhythm type, or the notes adjacent to the target melody skeleton features can be restored according to the original rhythm type to obtain the target rhythm type. This example does not make special restrictions on this; among them, the melody curve of the target rhythm type obtained by performing rhythm restoration on the target melody skeleton features shown in 601 can be as shown in Figure 8 801 in Figure 9as shown by 901 in

[0109] So far, the method for generating audio data described in the exemplary embodiments of the present disclosure has been fully implemented. Based on the foregoing content, it can be known that the method for generating audio data described in the exemplary embodiments of the present disclosure can realize the function of imitating melody creation with refined melody trends from the dimension of the melody skeleton, and thus the obtained target song can be more accurately targeted for creation, providing a new ability basis for the style imitation creation of songs and improving the user experience.

[0110] Exemplary Apparatus

[0111] After introducing the method for generating audio data of the exemplary embodiments of the present disclosure, next, reference is made to Figure 10 to explain and illustrate the apparatus for generating audio data of the exemplary embodiments of the present disclosure. Specifically, referring to Figure 10 as shown, the apparatus for generating audio data may include an audio feature extraction module 1010, a trend feature determination module 1020, a feature evolution module 1030, and an audio data generation module 1040. Among them:

[0112] The audio feature extraction module 1010 can be used to extract features from the original audio data to obtain the original melody skeleton features; the trend feature determination module 1020 can be used to determine the melody trend features of the original audio data according to the original melody skeleton features; the feature evolution module 1030 can be used to perform feature evolution on the original melody skeleton features to obtain the target melody skeleton features; the audio data generation module 1040 can be used to generate target audio data according to the melody trend features and the target melody skeleton features.

[0113] In an exemplary embodiment of the present disclosure, the original melody skeleton features are features composed of the most core notes and / or the most stable notes in the original audio data; the notes included in the original melody skeleton features include at least one of the following: main melody notes, harmony support notes, key rhythm points, and the starting note and ending note of the target musical phrase.

[0114] In an exemplary embodiment of the present disclosure, feature extraction is performed on the original audio data to obtain original melody skeleton features, including: segmenting the original audio data based on a preset segment segmentation model to obtain one or more musical phrase segments; segmenting the musical phrase segments based on a preset beat segmentation model to obtain one or more current musical phrase beats; extracting sub-skeleton notes from the current musical phrase beats, and generating original melody skeleton features according to the sub-skeleton notes.

[0115] In an exemplary embodiment of the present disclosure, extracting sub-skeleton notes from the musical phrase beats includes: constructing a first musical phrase beat set according to the current musical phrase beats; traversing the first musical phrase beat set, selecting any current musical phrase beat as the target musical phrase beat, and extracting the current notes included in the target musical phrase beat to obtain a first note extraction result; when it is determined that the first note extraction result is not empty, extracting the current chord tone notes of the target musical phrase beat from the first note extraction result to obtain a second note extraction result, and determining the sub-skeleton notes of the target musical phrase beat according to the second note extraction result; sequentially repeating the extraction process of the sub-skeleton notes of the target musical phrase beat to obtain the sub-skeleton notes of other current musical phrase beats in the first musical phrase beat set except the target musical phrase beat.

[0116] In an exemplary embodiment of the present disclosure, determining the sub-skeleton notes of the target musical phrase beat according to the second note extraction result includes: when it is determined that the second note extraction result is empty, taking the first note in the first note extraction result as the sub-skeleton note of the target musical phrase beat; when it is determined that the second note extraction result is not empty, determining the number of current chord tone notes in the second note extraction result, and when it is determined that the number is one, taking this current chord tone note as the sub-skeleton note of the target musical phrase beat; when it is determined that the number is multiple, determining a target chord tone note from the current chord tone notes according to the volume of the current chord tone notes, and taking this target chord tone note as the sub-skeleton note of the target musical phrase beat.

[0117] In an exemplary embodiment of the present disclosure, the preset segment segmentation model includes an embedding mapping layer, an encoding layer, and a mixture-of-experts model layer; wherein, segmenting the original audio data based on the preset segment segmentation model to obtain one or more musical phrase segments includes: obtaining audio attribute information of the original audio data; wherein, the audio attribute information includes at least one of the name information of the original audio data, the audio creator information, and the lyric information corresponding to the original audio data; generating basic information to be predicted according to the original note information of the original audio data, and generating context information to be predicted according to the audio attribute information and preset parameter hint information; performing embedding mapping processing on the basic information to be predicted based on the embedding mapping layer to obtain a first audio feature, and performing embedding mapping processing on the context information to be predicted based on the embedding mapping layer to obtain a first context flag sequence; performing encoding processing on the first audio feature and the first context flag sequence based on the encoding layer to obtain a first context overall representation, and performing segment segmentation on the first context flag sequence and the first context overall representation based on the mixture-of-experts model layer to obtain one or more musical phrase segments.

[0118] In an exemplary embodiment of the present disclosure, the mixture-of-experts model layer includes a gated network model and multiple expert neural network models; wherein, segmenting the first context flag sequence and the first context overall representation based on the mixture-of-experts model layer to obtain one or more musical phrase segments includes: determining, based on the gated network model according to the first context flag sequence, a first model weight of the expert neural network model in the note structure dimension and a second model weight in the music style dimension; determining, based on the first model weight and the second model weight, a first target neural network model required to perform the segment segmentation task in the note structure dimension and a second target neural network model required to perform the segment segmentation task in the music style dimension from the multiple expert neural network models; inputting the first context flag sequence and the first context overall representation into the first target neural network model and the second target neural network model respectively to obtain a first segment segmentation result in the note structure dimension and a second segment segmentation result in the music style dimension, and generating one or more musical phrase segments according to the first segment segmentation result and the second segment segmentation result.

[0119] In an exemplary embodiment of the present disclosure, determining the melody trend feature of the original audio data according to the original melody skeleton feature includes: classifying the original notes in the original audio data to obtain first in-key notes and first out-of-key notes, and classifying the original characteristic notes in the original melody skeleton feature to obtain second in-key notes and second out-of-key notes; determining the first in-key pitch of the first in-key notes and the first out-of-key pitch of the first out-of-key notes, and determining the second in-key pitch of the second in-key notes and the second out-of-key pitch of the second out-of-key notes; determining the in-key trend feature according to the first in-key pitch and the second in-key pitch, and determining the out-of-key trend feature according to the first out-of-key pitch and the second out-of-key pitch; determining the melody trend feature of the original audio data according to the in-key trend feature and the out-of-key trend feature.

[0120] In an exemplary embodiment of the present disclosure, determining the in-key trend feature according to the first in-key pitch and the second in-key pitch includes: determining a first pitch difference between the first in-key pitch and the second in-key pitch, and classifying the first pitch difference to obtain a first zero-degree difference, a first second-degree difference, a first second-degree adjacent difference, and a first high-degree difference; identifying the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference with preset characters to obtain a difference identification result, and determining the in-key trend feature according to the difference identification result.

[0121] In an exemplary embodiment of the present disclosure, identifying the first zero-degree difference, the first second-degree difference, the first second-degree adjacent difference, and the first high-degree difference with preset characters includes: extracting a first positive second-degree difference and a first negative second-degree difference from the first second-degree difference, and extracting a first positive second-degree adjacent difference and a first negative second-degree adjacent difference from the first second-degree adjacent difference; extracting a first positive high-degree difference and a first negative high-degree difference from the first high-degree difference, and respectively configuring character identifications for the first negative high-degree difference, the first negative second-degree adjacent difference, the first negative second-degree difference, the first zero-degree difference, the first positive second-degree difference, the first positive second-degree adjacent difference, and the first positive high-degree difference to obtain a first difference identification result, …, a seventh difference identification result; splicing the difference identification results corresponding to the first in-key notes in sequence according to the order of the first in-key notes in the original audio data to obtain the in-key trend feature.

[0122] In an exemplary embodiment of the present disclosure, evolving the original melody skeleton features to obtain target melody skeleton features includes: extracting original beats to be evolved from the current phrase beats of the original audio data according to the skeleton similarity value between the original audio data and the target audio data, and constructing a second phrase beat set based on the extracted original beats to be evolved; traversing the second phrase beat set, extracting any beat from the original beats to be evolved as the target beat to be evolved, and determining the original characteristic note corresponding to the target beat to be evolved from the original melody skeleton features; extracting the chord notes included in the target beat to be evolved, and determining the evolved characteristic note of the target beat to be evolved according to the chord notes and the original characteristic note; sequentially repeating the evolution process of the evolved characteristic note of the target beat to be evolved to obtain the evolved characteristic notes of the other beats to be evolved in the second phrase beat set except the target beat to be evolved, and generating the target melody skeleton features according to the evolved characteristic notes.

[0123] In an exemplary embodiment of the present disclosure, determining the evolved characteristic note of the target beat to be evolved according to the chord notes and the original characteristic note includes: constructing an original evolved note set according to the chord notes and the original characteristic note, and calculating the interval relationship between the notes in the original evolved note set and the original characteristic notes of the current phrase beats adjacent to the target beat to be evolved; filtering the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set, and extracting any note from the target evolved note set as the evolved characteristic note of the target beat to be evolved.

[0124] In an exemplary embodiment of the present disclosure, filtering the notes in the original evolved note set according to the interval relationship to obtain a target evolved note set includes: if the interval relationship is that the interval between the note in the original evolved note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is greater than or equal to the preset interval threshold, deleting the note; if the interval relationship is that the interval between the note in the original evolved note set and the original characteristic note of the current phrase beat adjacent to the target beat to be evolved is less than the preset interval threshold, retaining the note, and generating the target evolved note set according to the retained notes.

[0125] In an exemplary embodiment of the present disclosure, generating target audio data according to the melody trend feature and the target melody skeleton feature includes: determining the original rhythm type of the original audio data, and performing rhythm restoration on the target melody skeleton feature according to the original rhythm type to obtain a target rhythm type; performing pitch restoration on the target melody skeleton feature according to the melody trend feature to obtain a target melody pitch, and generating the target audio data according to the target rhythm type, the target melody pitch, and the target melody skeleton feature.

[0126] Exemplary Storage Medium

[0127] After introducing the method for generating audio data and the apparatus for generating audio data according to the exemplary embodiments of the present disclosure, next, reference is made to Figure 11 describe the storage medium according to the exemplary embodiments of the present disclosure.

[0128] Reference is made to Figure 11 As shown, a program product 1100 for implementing the above method according to an embodiment of the present disclosure is described, which may be a portable compact disc read-only memory (CD-ROM) and includes program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.

[0129] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0130] The computer-readable signal medium may include a data signal propagated in a baseband or as a part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium.

[0131] Program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN).

[0132] Exemplary Electronic Device

[0133] After introducing the storage medium of the exemplary embodiments of the present disclosure, next, reference is made to Figure 12 to describe the electronic device of the exemplary embodiments of the present disclosure.

[0134] Figure 12 The displayed electronic device 1200 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0135] As Figure 12 shown, the electronic device 1200 is presented in the form of a general-purpose computing device. The components of the electronic device 1200 may include, but are not limited to: at least one of the above-mentioned processing units 1210, at least one of the above-mentioned storage units 1220, a bus 1230 connecting different system components (including the storage unit 1220 and the processing unit 1210), and a display unit 1240.

[0136] Among them, the storage unit 1220 stores program code, and the program code can be executed by the processing unit 1210, so that the processing unit 1210 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 1210 can execute steps S112 - S140 as shown in Figure 1 .

[0137] The storage unit 1220 may include a volatile storage unit, such as a random access storage unit (RAM) 12201 and / or a cache storage unit 12202, and may further include a read-only storage unit (ROM) 12203.

[0138] The storage unit 1220 may also include a program / utility 12204 having a set (at least one) of program modules 12205. Such program modules 12205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0139] The bus 1230 may include a data bus, an address bus, and a control bus.

[0140] The electronic device 1200 may also communicate with one or more external devices 1300 (such as a keyboard, a pointing device, a Bluetooth device, etc.) through an input / output (I / O) interface 1250. Moreover, the electronic device 1200 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1260. As shown in the figure, the network adapter 1260 communicates with other modules of the electronic device 1200 through the bus 1230. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0141] It should be noted that although several modules or sub-modules of the pop-up window processing device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units / modules may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.

[0142] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0143] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for generating audio data, characterized in that, Including: Performing feature extraction on the original audio data to obtain original melody skeleton features; Determining the melody trend features of the original audio data according to the original melody skeleton features; Performing feature evolution on the original melody skeleton features to obtain target melody skeleton features; Generating target audio data according to the melody trend features and the target melody skeleton features.

2. The method for generating audio data according to claim 1, wherein The original melody skeleton features are features composed of the most core notes and / or the most stable notes in the original audio data; The notes included in the original melody skeleton features include at least one of the following: main melody notes, harmonic support notes, key rhythm points, and the starting note and ending note of the target musical phrase.

3. The method for generating audio data according to claim 1, wherein Performing feature extraction on the original audio data to obtain original melody skeleton features, including: Performing segment segmentation on the original audio data based on a preset segment segmentation model to obtain one or more musical phrase segments; Performing beat segmentation on the musical phrase segments based on a preset beat segmentation model to obtain one or more current musical phrase beats; Extracting sub-skeleton notes in the current musical phrase beats and generating original melody skeleton features according to the sub-skeleton notes.

4. The method for generating audio data according to claim 3, wherein, Extracting sub-skeleton notes in the musical phrase beats, including: Constructing a first musical phrase beat set according to the current musical phrase beats; Traversing the first musical phrase beat set, selecting any current musical phrase beat as the target musical phrase beat, and extracting the current notes included in the target musical phrase beat to obtain a first note extraction result; When it is determined that the first note extraction result is not empty, extracting the current chord tone notes in the target musical phrase beat from the first note extraction result to obtain a second note extraction result, and determining the sub-skeleton notes of the target musical phrase beat according to the second note extraction result; Repeating the extraction process of the sub-skeleton notes of the target musical phrase beat in sequence to obtain the sub-skeleton notes of other current musical phrase beats in the first musical phrase beat set except the target musical phrase beat.

5. The method for generating audio data according to claim 4, wherein Determining the sub-skeleton notes of the target musical phrase beat according to the second note extraction result, including: When it is determined that the second note extraction result is empty, taking the first note in the first note extraction result as the sub-skeleton note of the target musical phrase beat; When it is determined that the second note extraction result is not empty, determining the number of current chord tone notes in the second note extraction result, and when it is determined that the number of notes is one, taking the current chord tone note as the sub-skeleton note of the target musical phrase beat; When it is determined that the number of notes is multiple, determining the target chord tone note from the current chord tone notes according to the volume of the current chord tone notes, and taking the target chord tone note as the sub-skeleton note of the target musical phrase beat.

6. The method for generating audio data according to claim 3, wherein The preset segment segmentation model includes an embedding mapping layer, an encoding layer, and a mixture of experts model layer; Among them, performing segment segmentation on the original audio data based on a preset segment segmentation model to obtain one or more musical phrase segments, including: Obtain the audio attribute information of the original audio data; wherein, the audio attribute information includes at least one of the name information of the original audio data, the audio creator information, and the lyrics information corresponding to the original audio data; Generate the basic information to be predicted according to the original note information of the original audio data, and generate the context information to be predicted according to the audio attribute information and the preset parameter prompt information; Perform an embedding mapping process on the basic information to be predicted based on the embedding mapping layer to obtain the first audio feature, and perform an embedding mapping process on the context information to be predicted based on the embedding mapping layer to obtain the first context flag sequence; Perform an encoding process on the first audio feature and the first context flag sequence based on the encoding layer to obtain the first context overall representation, and perform a segment segmentation on the first context flag sequence and the first context overall representation based on the mixture of experts model layer to obtain one or more phrase segments.

7. The method for generating audio data according to claim 6, wherein The mixture of experts model layer includes a gated network model and multiple expert neural network models; Wherein, performing a segment segmentation on the first context flag sequence and the first context overall representation based on the mixture of experts model layer to obtain one or more phrase segments includes: Based on the gated network model, determine the first model weight in the note structure dimension and the second model weight in the music style dimension of the expert neural network model according to the first context flag sequence; Based on the first model weight and the second model weight, determine the first target neural network model required to perform the segment segmentation task in the note structure dimension and the second target neural network model required to perform the segment segmentation task in the music style dimension from the multiple expert neural network models; Input the first context flag sequence and the first context overall representation into the first target neural network model and the second target neural network model respectively to obtain the first segment segmentation result in the note structure dimension and the second segment segmentation result in the music style dimension, and generate the one or more phrase segments according to the first segment segmentation result and the second segment segmentation result.

8. An apparatus for generating audio data, characterized in that, Including: An audio feature extraction module, configured to extract features from the original audio data to obtain the original melody skeleton feature; A trend feature determination module, configured to determine the melody trend feature of the original audio data according to the original melody skeleton feature; A feature evolution module, configured to perform feature evolution on the original melody skeleton feature to obtain the target melody skeleton feature; An audio data generation module, configured to generate target audio data according to the melody trend feature and the target melody skeleton feature.

9. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for generating audio data according to any one of claims 1-7 is implemented.

10. An electronic device, including: A processor; And A memory, configured to store the executable instructions of the processor; Wherein, the processor is configured to execute the method for generating audio data according to any one of claims 1-7 by executing the executable instructions.