A method, apparatus and storage medium for generating music
By acquiring the note information and context information of the melody audio, and combining it with the playing technique rules of the target instrument, matching sound sources are selected from the sound source database to generate natural music that matches the characteristics of the instrument, thus solving the problems of stiff music and high training costs in existing technologies.
Patent Information
- Application Number
- CN202411327201.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-09-23
AI Technical Summary
In existing technologies, music generated by instrument timbre synthesis methods based on digital signal processing or machine learning is stiff and unnatural, and timbre modeling requires a lot of time and model training costs.
By acquiring the note information and context information of the melody audio, and combining it with the playing technique rules of the target instrument, matching sound sources are selected from the sound source database to generate music that conforms to the actual playing characteristics of the instrument and is natural.
The generated music is more natural, matches the actual playing characteristics of the instrument, and can quickly generate audio files of instrument performances without requiring a lot of training.
Smart Images

Figure CN119252218B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus and storage medium for generating music. Background Technology
[0002] To enrich musical content, especially for pieces played on a specific instrument, music generation methods are often used to automatically generate music performed on that instrument. Current technologies employ digital signal processing or machine learning-based approaches for instrument timbre synthesis. Specifically, this involves acquiring the fundamental frequency sequence of the desired melody, modeling the instrument's timbre to extract its features, and finally synthesizing the acquired fundamental frequency sequence and timbre features to obtain the audio file of the instrument's performance. However, the synthesized music sounds stiff and unnatural, and parameter-based timbre modeling requires significant time and training costs. Summary of the Invention
[0003] In view of this, the purpose of this invention is to provide a music generation method, device, and storage medium that enables the final synthesized music to be more natural and conform to the actual performance characteristics of musical instruments. The specific solution is as follows:
[0004] In a first aspect, this application discloses a music generation method, including:
[0005] Acquire melody audio and extract musical information from the melody audio; the musical information includes note information, note context information, and melody information of different notes in the melody audio;
[0006] Each note in the melody audio is traversed sequentially, and based on the note information of the current note, the first sound source that matches the pitch of the current note is selected from the sound source database of the target type of instrument.
[0007] Based on the note information and note context information of the current note, a second sound source is selected from the first sound source according to the playing technique rules corresponding to the target type of instrument; the playing technique rules include the correspondence between playing fingering and notes;
[0008] The matching sound source corresponding to the current note is determined based on the second sound source;
[0009] Based on the matching sound source corresponding to each note in the melody audio and the melody information, music for the target type instrument corresponding to the melody audio is generated.
[0010] Optionally, the performance technique rules include a first technique rule that represents the correspondence between the target position of a note in a musical phrase or beat and the fingering, and a second technique rule that represents the correspondence between dynamics and string signs and fingering.
[0011] The step of selecting a second sound source from the first sound source based on the note information and note context information of the current note, according to the playing technique rules corresponding to the target type of instrument, includes:
[0012] Based on the note context information of the current note, and in accordance with the first technique rule, a target first sound source that matches the fingering of the current note is selected from the first sound source;
[0013] Based on the note information of the current note, and in accordance with the second technique rule, a second sound source matching the fingering of the current note is selected from the target first sound source.
[0014] Optionally, the target position includes the first note of a musical phrase, the last note of a musical phrase, and notes located on a downbeat;
[0015] Before selecting a target first sound source matching the fingering of the current note from the first sound source according to the first technique rule based on the note context information of the current note, the method further includes:
[0016] The position information of the current note is determined based on the note context information of the current note;
[0017] Based on the position information, determine whether the current note belongs to the target position. If it belongs to the target position, then execute the step of selecting a target first sound source that matches the fingering of the current note from the first sound source according to the first technique rule based on the note context information of the current note.
[0018] If it does not belong to the target position, then according to the note information of the current note, and in accordance with the second technique rule, a second sound source that matches the fingering of the current note is selected from the first sound source.
[0019] Optionally, the position information of the current note is determined based on the note context information of the current note, including:
[0020] Based on the note context information of the current note, determine the first time difference between the current note and the previous note, and the second time difference between the current note and the next note;
[0021] If the first time difference is greater than the time of one beat, then the current note is determined to be the first note of the musical phrase;
[0022] If the second time difference is greater than the time of one beat, then the current note is determined to be the last note of the musical phrase.
[0023] Optionally, the position information of the current note is determined based on the note context information of the current note, including:
[0024] If the time difference between the current note and the start time of the stressed beat in the current time signature is less than a preset duration threshold, then the current note is determined to be on a stressed beat.
[0025] The start time of the repetition within the time signature is an element in the repetition start time sequence; the repetition start time sequence is generated by combining the time signature and music tempo with a linearly increasing sequence.
[0026] Optionally, the target type of musical instrument is the guqin;
[0027] Before sequentially traversing each note in the melody audio, the following is also included:
[0028] Determine whether the range of the melody information is within the range of the guqin; if it is not within the range of the guqin, then convert the range of the melody audio to the range of the guqin by raising or lowering an octave.
[0029] Optionally, the process of generating the sound source database for the target type of musical instrument includes:
[0030] Obtain the sound source data of the target type of musical instrument;
[0031] The sound source data is cleaned; the sound source cleaning includes pitch value checking and sound source start point checking.
[0032] The cleaned audio source data is expanded to obtain an audio source database.
[0033] Optionally, after generating the music for the target type instrument corresponding to the melody audio, the method further includes:
[0034] Query whether there exist two consecutive notes on the same string and the note interval is less than the interval threshold in the music of the target type instrument corresponding to the melody audio;
[0035] If they exist, the first note of the two connected notes will be gradually weakened.
[0036] Secondly, this application discloses an electronic device, including:
[0037] Memory, used to store computer programs;
[0038] A processor is used to execute the computer program to implement the aforementioned music generation method.
[0039] Thirdly, this application discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned music generation method.
[0040] Fourthly, this application discloses a computer program product, including a computer program that, when executed by a processor, implements the aforementioned music generation method.
[0041] In this application, a melody audio is acquired, and its musical information is extracted. The musical information includes note information, note context information, and melody information for different notes in the melody audio. Each note in the melody audio is sequentially traversed, and based on the note information of the current note, a first sound source matching the pitch of the current note is selected from the sound source database of the target type instrument. Based on the note information and note context information of the current note, a second sound source is selected from the first sound source according to the playing technique rules corresponding to the target type instrument. The playing technique rules include the correspondence between playing fingering and notes. The matching sound source corresponding to the current note is determined based on the second sound source. Based on the matching sound source corresponding to each note in the melody audio and the melody information, music for the target type instrument corresponding to the melody audio is generated. By first selecting a subset of sound sources from a sound source database based on note pitch, and then selecting a second sound source from the first sound source based on note information and note context information, and according to the correspondence between playing fingering and notes; by combining note information and note context information, and according to the correspondence between notes and playing fingering, the sound source with matching fingering is determined for the notes, making the final synthesized music more natural and in line with the actual playing characteristics of the instrument; and, users only need to provide the melody they want to synthesize, and an audio file of the instrument playing that melody can be quickly generated. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0043] Figure 1 A flowchart of a music generation method provided in this application;
[0044] Figure 2 A schematic diagram of the system framework applicable to the music generation scheme provided in this application;
[0045] Figure 3 This application provides a structural diagram of an electronic device. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] In existing technologies, instrument timbre synthesis employs methods based on digital signal processing or machine learning. Specifically, this involves acquiring the fundamental frequency sequence of the desired melody, modeling the instrument timbre to extract its features, and finally synthesizing the acquired fundamental frequency sequence and timbre features to obtain the instrument's audio file. However, the synthesized music sounds stiff and unnatural, and parameter-based timbre modeling requires significant time for model training. To overcome these technical problems, this application proposes a music generation method that produces more natural synthesized music that matches the actual performance characteristics of the instrument, without requiring substantial training costs.
[0048] This application discloses a music generation method, see [link to relevant documentation]. Figure 1 As shown, the method may include the following steps:
[0049] Step S11: Obtain the melody audio and extract the music information of the melody audio; the music information includes note information of different notes in the melody audio, note context information, and melody information.
[0050] First, the melody audio is acquired to generate performance music for the target type of instrument based on the melody. Musical information is extracted from the melody audio, including but not limited to note information for different notes, contextual information between notes, and melody information.
[0051] This involves preprocessing the input melody audio. The preprocessing includes: extracting note information from the melody audio and storing each note as a token; traversing the token sequence and recording the context information of each token; extracting the tonality, range, tempo (BPM, Beat Per Minute), and time signature of the note information; and adjusting the range of the input melody. The note information includes, but is not limited to, pitch, starting point, duration, and dynamics; the melody information includes, but is not limited to, tonality, range, tempo, and time signature; and the target instrument type can be a guqin or guzheng, etc.
[0052] Specifically, for note information, the melody audio is input as a MIDI (Musical Instrument Digital Interface) file. The MIDI file is parsed, and all notes are traversed sequentially, recording the pitch, onset, duration, and velocity of each note. In essence, the melody audio is segmented according to notes, resulting in multiple audio segments. The onset indicates the point in time within the segment where the note begins playing. These four attributes are stored in a token: [pitch, onset, duration, velocity]. This process is repeated. The note token sequence is traversed, recording the context information of each token. Specifically, for each note, its preceding note, its following note, and its own token need to be recorded; this context information can be used to determine certain special techniques.
[0053] For melody information, the tonality, range, tempo (BPM), and time signature of the input melody are extracted. The tonality is determined by calculating the Pearson product-moment correlation coefficient between the cumulative total duration of notes at different pitches in the melody audio and the weights of different tonality features; the tonality corresponding to the Pearson product-moment correlation coefficient with the largest absolute value is taken as the tonality of the melody audio. For example, the Krumhansl-Schmuckler algorithm is used to extract the tonality of the input melody; its core idea is to calculate the Pearson product-moment correlation coefficients between the cumulative total duration of notes at 12 different pitches in the melody audio and the weights of 24 different tonality features, and the tonality corresponding to the largest absolute value among the 24 values is the tonality of the melody. The range of the input melody is the interval between the maximum and minimum pitches of the notes in the melody audio. Tempo and time signature can be directly parsed from the MIDI file.
[0054] If the target instrument is the guqin, before sequentially traversing each note in the melody audio, the process may further include: determining whether the range of the melody information is within the guqin's range; if not, converting the melody audio's range to the guqin's range by raising or lowering an octave. It is understandable that, since the guqin's range is only in the C1-D5 interval, and to ensure the final generated guqin music is more natural, the input melody's range needs to be converted to the guqin's range. Converting it to the guqin's range by raising or lowering an octave specifically includes: determining whether the input melody's range is within the guqin's range; if so, no processing is performed; otherwise, the MIDIpitch value of all note tokens is reduced by 12 (i.e., one octave), processed, and then the determination is repeated until the condition is met.
[0055] Step S12: Iterate through each note in the melody audio in sequence, and select the first sound source that matches the pitch of the current note from the sound source database of the target type of instrument based on the note information of the current note.
[0056] In this embodiment, the notes extracted from the melody audio are matched sequentially with the cleaned sound source database. First, sound sources with the same pitch as the note are selected as the first batch of candidate sound sources, resulting in the aforementioned first sound source, which is also the sound source data with the same string.
[0057] In some embodiments, the process of generating the sound source database for the target type of musical instrument may include: acquiring sound source data of the target type of musical instrument; cleaning the sound source data; the sound source cleaning includes pitch value checking and sound source start point checking; and expanding the cleaned sound source data to obtain the sound source database. It is understood that the acquired sound source data may be inaccurate, and direct use may lead to poor synthesized music quality; therefore, cleaning is performed first. Furthermore, the acquired sound source data may not be comprehensive; therefore, expansion is used to increase the richness of the sound source data to meet the needs of music synthesis and adapt to melody inputs of different keys.
[0058] The aforementioned sound source cleaning process can include: detecting the fundamental frequency value of the sound source data using a fundamental frequency detection method; determining whether the pitch of the current sound source data is standard by comparing the fundamental frequency value with the pitch value in the annotation information corresponding to the current sound source data; and updating the pitch value based on the fundamental frequency value for non-standard sound source data. It can be understood that sound source database cleaning involves checking whether the labeled pitch value of each sound source is accurate.
[0059] Specifically, the fundamental frequency value predicted by the fundamental frequency detection algorithm is compared with the pitch value labeled in the audio data. If the difference between the extracted fundamental frequency value and the labeled value is less than a minor second, the pitch value labeled in the audio source is considered accurate; otherwise, it is considered inaccurate. The fundamental frequency detection algorithm can use methods such as the PYIN algorithm, autocorrelation, or average amplitude difference to extract the fundamental frequency. Secondly, the fundamental frequency value and pitch value need to be converted to the same MIDI pitch for comparison, that is, both are unified to the same unit for comparison. The formula for converting fundamental frequency to MIDI pitch is:
[0060] ;
[0061] The fundamental frequency value is extracted by the fundamental frequency detection algorithm, and the unit is Hz. The process of converting the pitch value to MIDI pitch is a simple mapping, as shown in Table 1 below:
[0062] Table 1
[0063] pitch value MIDI pitch C4 48 C#4 / Db4 49 D4 50 D#4 / Eb4 51
[0064] If the difference between the extracted fundamental frequency value and the labeled pitch value is less than a minor second, that is, the absolute value of the MIDI pitch converted from the fundamental frequency value minus the absolute value converted from the labeled pitch value is less than 1, then the pitch is considered standard; otherwise, it is not standard. For non-standard sound source data, the pitch label needs to be removed or changed.
[0065] The aforementioned audio source cleaning of the audio source data may further include: determining the time point of the first local peak of the audio source data by peak extraction based on the energy of the audio source data in each frame; aligning the time point of the first local peak with a preset time point. That is, audio source database cleaning includes checking whether the starting point of each audio source is aligned. Starting point detection is performed on each audio source, specifically by horizontally shifting the audio along the time axis, aligning the detected audio source starting point with 20ms (or other time points), meaning that all audio data begins to produce sound after 20ms. The starting point detection process includes: first calculating the energy envelope of the audio, dividing the audio into frames, and then calculating the energy and timestamp of each frame x(n); wherein, the energy calculation is as follows: Then, a peak picking algorithm is used to find all the local peaks in the energy envelope, and the time point of the first local peak is the starting point of the sound source. Aligning the starting point ensures that the effect of all sound source data is consistent when played, and avoids notes being out of sync.
[0066] The step of expanding the cleaned audio source data to obtain an audio source database includes: acquiring audio source information for each audio source; the audio source information is pre-labeled and includes audio source ID, player, fingering, string number, pitch, velocity, and file path; determining the target audio source to be acquired by traversing the audio source information, and determining the most similar original audio source data corresponding to the target audio source based on the audio source information; performing a Fourier transform on the original audio source data to obtain the corresponding frequency domain data; performing frequency expansion on the frequency domain data using the frequency expansion method to obtain expanded data, and performing an inverse Fourier transform on the expanded data to obtain the target audio.
[0067] First, the filename of each sound source is parsed to extract the following information: sound source ID, which hand is playing (left hand (L) or right hand (R)), playing technique, string number (flag position), pitch, velocity, and file path. For example, if a sound source is named "L_MO_NEI_Str_5_E4_T0_f_TK1.wav", the extracted information is: [0001, L, smudge, fifth string, 52, f, "data / L_HARM_NEI_Str_5_E4_T0_f_TK1.wav"]. Here, MO is the smudge technique, 52 is the MIDI pitch of E4, and f is the abbreviation for "strong" velocity, indicating that the volume of this note needs to be relatively high.
[0068] In this embodiment, considering that typical guqin sound source databases only collect notes from the pentatonic scale within a fixed key, the sound source database is expanded to include uncollected notes in order to generate guqin music in different keys. The sound source expansion process includes: traversing note information and filtering out uncollected notes based on the guqin's range C1-D5. For example, if the note F4 played on the fifth string using the mo technique and p dynamics is not collected, it is recorded as the target sound source. Then, the sound source with the smallest pitch difference from F4 on the same string, using the same technique and dynamics, is found from the sound source database and recorded as the original sound source. The target sound source is then generated based on the original sound source using a frequency expansion algorithm. The frequency spreading algorithm not only shifts the fundamental frequency but also shifts all harmonics by the same proportion, thus preserving the characteristics of the original timbre. Specifically, it involves: first, using a Fast Fourier Transform (FFT) to convert the original sound source signal from the time domain to the frequency domain; then, in the frequency domain, shifting all frequency components by the same proportion, for example, shifting E4 (approximately 329.63 Hz) to F4 (approximately 349.23 Hz) requires multiplying all frequency components by 349.23 / 329.63, which is the shift ratio; finally, using an Inverse Fast Fourier Transform (IFFT) to convert the modified spectrum back from the frequency domain to the time domain, obtaining the target sound source.
[0069] Step S13: Based on the note information and note context information of the current note, and according to the playing technique rules corresponding to the target type of instrument, select the second sound source from the first sound source; the playing technique rules include the correspondence between playing fingering and notes.
[0070] That is, the fingering for playing notes is also related to the note information and the context in which the note is located. Therefore, after filtering based on pitch, the first sound source is further filtered based on the note information and the note context information, according to the playing technique rules of the target type of instrument, to select the second sound source.
[0071] That is, for each note in the melodic audio, a corresponding matching sound source is selected from the sound source database. The matching includes pitch matching, loudness matching, and appropriate fingering. The above-mentioned sound source database contains a large amount of sound source data of the target type of musical instrument, and the sound source data distinguishes fingering, that is, the techniques used for playing. Taking the guqin as an example, the fingering includes but is not limited to: wiping, picking, hooking, flicking, striking, plucking, splitting, supporting, hitting, harmonic sounds, etc. The sound source database contains audio data of different strings using different fingering and audio data of different strings with different loudness. Using the sound source database constructed therefrom to generate music can form guqin music with rich playing techniques.
[0072] In this embodiment, the playing technique rules include a first technique rule representing the corresponding relationship between the target position of a note in a musical phrase or beat and the fingering, and a second technique rule representing the corresponding relationship between the intensity and the string number and the fingering. The screening of the second sound source from the first sound source according to the note information and note context information of the current note according to the playing technique rules corresponding to the target type of musical instrument may include: screening out a target first sound source whose fingering matches the current note from the first sound source according to the note context information of the current note according to the first technique rule; and screening out a second sound source whose fingering matches the current note from the target first sound source according to the note information of the current note according to the second technique rule. That is, after obtaining the first sound source, according to the context information of the note, the target first sound source is screened out from the first sound source by using the first technique rule.
[0073] It can be understood that, taking the guqin as an example, the playing of the guqin generally involves taking sounds by pressing the strings with the left hand and producing sounds by plucking the strings with the right hand. In terms of the pressing position, there is a difference in timbre between pressing with the flesh and pressing with half flesh and nail. In terms of the plucking of the right hand, there are also differences between fingers and differences between plucking with the back of the nail and plucking with the flesh of the finger surface. The amplitude of its strength change is quite different. Therefore, it is modeled through the context information of the note. The context information of the note is divided into three cases: 1. The current note is the first note of the musical phrase where it is located; 2. The current note is the last note of the musical phrase where it is located; 3. The current note appears on a strong beat (such as the first beat of a 4 / 4 beat). For example, when a note is the first note in a musical phrase, fingering such as "releasing and combining", "lifting like one", "plucking like one", or "pulling" can be randomly adopted; when a note is the last note in a musical phrase, fingering such as "single hitting", "harmonic sound (inside)", "harmonic sound (outside)", "harmonic sound (supporting)" can be randomly adopted; when a note appears on a strong beat (the strong beat is the first beat of each measure), fingering such as "releasing and combining", "lifting like one", "plucking like one", or "pulling" can be randomly adopted. In addition, it should be noted that when it is determined that the technique of the note is "harmonic sound", a sound source one octave higher than the pitch of the note needs to be matched in the sound source library.
[0074] The aforementioned target positions include the first note of a musical phrase, the last note of a musical phrase, and notes located on a downbeat. Before selecting a target first sound source matching the fingering of the current note from the first sound source based on the note context information of the current note and according to the first technique rule, the method further includes: determining the position information of the current note based on the note context information of the current note; determining whether the current note belongs to the target position based on the position information; if it belongs to the target position, then performing the step of selecting a target first sound source matching the fingering of the current note from the first sound source based on the note context information of the current note and according to the first technique rule; if it does not belong to the target position, then selecting a second sound source matching the fingering of the current note from the first sound source based on the note information of the current note and according to the second technique rule. First, determine if it is the target position. If it is, use the fingering corresponding to the position as the filtering condition to further filter the first sound source to obtain the target first sound source. Then, based on the target first sound source, filter out the second sound source according to the second technique rules. If it is not the target position, directly filter out the second sound source based on the first sound source according to the second technique rules.
[0075] Specifically, determining the position information of the current note based on its context information includes: determining a first time difference between the current note and the previous note, and a second time difference between the current note and the next note, based on the context information. If the first time difference is greater than one beat, the current note is determined to be the first note of the phrase; if the second time difference is greater than one beat, the current note is determined to be the last note of the phrase. For example, using the context information, the time differences between the current note and the previous and next notes are calculated. If the starting time difference between the current note and the previous note is greater than or equal to the time taken for one beat, the current note is the first note of the phrase; if the starting time difference between the current note and the next note is greater than or equal to the time taken for one beat, the current note is the last note of the phrase. The time taken for one beat is denoted as t, and the calculation formula is:
[0076] .
[0077] Specifically, determining the position information of the current note based on the note context information includes: if the time difference between the current note and the start time of the stressed beat within the current time signature is less than a preset duration threshold, then the current note is located on a stressed beat; the start time of the stressed beat within the time signature is an element in the stressed beat start time sequence; the stressed beat start time sequence is generated based on the time signature, musical tempo, and a linearly increasing sequence.
[0078] For example, if the time difference between the starting point of the current note and the starting point of the downbeat of the current time signature is less than the set threshold, then the current note is considered to appear on the downbeat; the threshold can be set to 20 ms because this value is the minimum time difference that humans can distinguish between two sounds. The process of determining whether it is a downbeat is as follows: First, generate a time series of downbeat starting points, denoted as downbeat(n), and the calculation formula is:
[0079] ;
[0080] where is a linearly increasing sequence representing the number of bars, which is [0, 1, 2, 3, 4...]; is the upper number of the time signature. Subtract the elements in the downbeat(n) sequence from the starting point of the current note, denoted as , which is the time difference between the starting point of the current note and the starting point of the downbeat of the current time signature. The formula is as follows:
[0081] ; where is the starting point of the current note. It can be seen that through the above three methods, it is possible to determine whether the note is the note at the target position, and then, according to the fingering usually used for the note at the target position by the instrument, the fluency and authenticity of the finally generated music can be improved.
[0082] Furthermore, different fingering methods are used for different strings at different volumes, and the same string may also use different fingering methods at different volumes. For example, for a soft note played on the first or second string, "strike" or "pick" is used, and for a strong note, "hook" or "flick" is used; for a soft note played on the third, fourth, or fifth string, "wipe" or "hook" is used, and for a strong note, "hook" or "flick" is used. Therefore, the second sound source is selected from the target first sound source according to the intensity of the note.
[0083] It can be seen that according to the correspondence between the note position and the fingering, the target first sound source with matching fingering is further selected from the first sound source, and then, according to the correspondence between the intensity, string number and fingering, the second sound source with matching fingering is further selected from the target first sound source; it can be seen that considering that different note positions, intensities and string numbers will use different fingering methods, the sound source database contains sound source data with different fingering methods. By combining the fingering factors to generate instrument music, the finally synthesized audio is more natural and conforms to the actual playing characteristics of the instrument; the user only needs to provide the melody to be synthesized, and the audio file of the instrument playing this melody can be quickly generated.
[0084] Step S14: Determine the matching sound source corresponding to the current note according to the second sound source.
[0085] If multiple second sound sources exist, one is randomly selected as the matching sound source for the current note; this process continues until all notes are matched. If, after filtering, the current note has only one second sound source, that second sound source is directly used as the matching sound source for the current note.
[0086] Step S15: Generate music for the target type instrument corresponding to the melody audio based on the matching sound source corresponding to each note in the melody audio and the melody information.
[0087] Finally, based on the matching sound sources corresponding to each note, music for the target type of instrument corresponding to the melody audio is generated according to the melody information of the melody audio. Thus, by first extracting the information of each note from the user-input melody audio, matching the corresponding sound source in the sound source database based on pitch, note information, and note context information, combined with performance technique rules, and then superimposing the matched sound sources in the time domain. Finally, after post-processing, including but not limited to reverb, delay, and compression, an audio file of the instrument being played is generated, achieving automated music synthesis based on performance technique rules; making the final synthesized audio more natural and more reminiscent of the guqin.
[0088] In some embodiments, after generating the music of the target type instrument corresponding to the melody audio, the process may further include: querying whether there are two connected notes on the same string and the note interval is less than an interval threshold in the music of the target type instrument corresponding to the melody audio; if so, then the first note of the two connected notes is gradually weakened to update the music of the target type instrument corresponding to the melody audio.
[0089] Understandably, to ensure the generated audio conforms to the physical principles of musical instrument sound production and allows each note to sound naturally, it's necessary to handle situations where two notes that are too close in time are played on the same string simultaneously. The processing logic involves weakening the signal before the second note begins to sound. Specifically, this weakening operation can be achieved by multiplying the signal 20ms before the second note by a non-linear amplitude envelope (denoted as fadeout(t)), as shown in the following formula:
[0090] ;
[0091] Where T represents time. After post-processing such as reverb, delay, and compression, the final music is obtained, achieving a realistic effect.
[0092] As can be seen from the above, in this embodiment, melody audio is acquired, and musical information of the melody audio is extracted. The musical information includes note information, note context information, and melody information for different notes in the melody audio. Each note in the melody audio is sequentially traversed, and based on the note information of the current note, a first sound source matching the pitch of the current note is selected from the sound source database of the target type instrument. Based on the note information and note context information of the current note, a second sound source is selected from the first sound source according to the playing technique rules corresponding to the target type instrument. The playing technique rules include the correspondence between playing fingering and notes. The matching sound source corresponding to the current note is determined based on the second sound source. Based on the matching sound source corresponding to each note in the melody audio and the melody information, music for the target type instrument corresponding to the melody audio is generated. Based on the matching sound source corresponding to each note in the melody audio and the melody information, music for the target type instrument corresponding to the melody audio is generated. By first selecting a subset of sound sources from a sound source database based on note pitch, and then selecting a second sound source from the first sound source based on note information and note context information, and according to the correspondence between playing fingering and notes; by combining note information and note context information, and according to the correspondence between notes and playing fingering, the sound source with matching fingering is determined for the notes, making the final synthesized music more natural and in line with the actual playing characteristics of the instrument; and, users only need to provide the melody they want to synthesize, and an audio file of the instrument playing that melody can be quickly generated.
[0093] The system framework used in the music generation scheme of this application can be found in [reference needed]. Figure 2 As shown, this may specifically include: a backend server and a number of user terminals that establish communication connections with the backend server. The user terminals include, but are not limited to, tablets, laptops, smartphones, and personal computers (PCs).
[0094] In this application, the user terminal sends melody audio to the backend server, and the backend server executes the following steps of the music generation method: acquiring the melody audio and extracting its music information; the music information includes note information, note context information, and melody information for different notes in the melody audio; sequentially traversing each note in the melody audio, and based on the note information of the current note, selecting a first sound source matching the pitch of the current note from the sound source database of the target type instrument; based on the note information and note context information of the current note, selecting a second sound source from the first sound source according to the playing technique rules corresponding to the target type instrument; the playing technique rules include the correspondence between playing fingering and notes; determining the matching sound source corresponding to the current note based on the second sound source; and generating music for the target type instrument corresponding to the melody audio based on the matching sound source corresponding to each note in the melody audio and the melody information. The backend server then pushes the generated music to the user terminal.
[0095] The following explanation uses a music app as an example to illustrate the technical solution of this application. Assume a user has installed this music app on their device. When they want to listen to a specific melody played by instrument A, they enter the app and upload the melody audio and instrument type to the server. The melody audio can be a song or instrumental music, etc.
[0096] After the server obtains the melody audio, it extracts the music information of the melody audio; the music information includes note information of different notes in the melody audio, note context information, and melody information; note information includes pitch, starting point, duration, and dynamics; melody information includes tonality, range, tempo, and time signature.
[0097] The server iterates through each note in the melody audio, selecting 10 sound sources with the same pitch as the current note from the sound source database of instrument A to obtain the first sound source. Based on the note context information and the first technique rule corresponding to the target instrument type (which determines the correspondence between the target position of the note in a musical phrase or beat and the fingering), it selects 5 sound sources from the first sound source that match the fingering of the current note to obtain the target first sound source. Based on the note information and the second technique rule corresponding to the target instrument type, it selects 2 sound sources from the target first sound source that match the fingering of the current note to obtain the second sound source, and selects one sound source from the second sound source as the matching sound source for the current note. The second technique rule determines the correspondence between dynamics and string number and fingering. Finally, based on the matching sound sources corresponding to each note in the melody audio, it generates music for instrument A corresponding to the melody audio.
[0098] Accordingly, this application also discloses a music generation device, which includes:
[0099] The audio acquisition module 11 is used to acquire melody audio and extract the music information of the melody audio; the music information includes note information of different notes in the melody audio, note context information, and melody information;
[0100] The first sound source determination module 12 is used to sequentially traverse each note in the melody audio and, based on the note information of the current note, filter out the first sound source that matches the pitch of the current note from the sound source database of the target type of instrument.
[0101] The second sound source determination module 13 is used to select a second sound source from the first sound source based on the note information and note context information of the current note, and according to the playing technique rules corresponding to the target type of instrument; the playing technique rules include the correspondence between playing fingering and notes;
[0102] The matching sound source determination module 14 is used to determine the matching sound source corresponding to the current note based on the second sound source;
[0103] The music generation module 15 is used to generate music for a target type instrument corresponding to the melody audio based on the matching sound source corresponding to each note in the melody audio and the melody information.
[0104] As can be seen from the above, in this embodiment, melody audio is acquired, and musical information of the melody audio is extracted. The musical information includes note information, note context information, and melody information of different notes in the melody audio. Each note in the melody audio is traversed sequentially, and a first sound source matching the pitch of the current note is selected from the sound source database of the target type instrument based on the note information of the current note. Based on the note information and note context information of the current note, a second sound source is selected from the first sound source according to the playing technique rules corresponding to the target type instrument. The playing technique rules include the correspondence between playing fingering and notes. The matching sound source corresponding to the current note is determined based on the second sound source. Music for the target type instrument corresponding to the melody audio is generated based on the matching sound source corresponding to each note in the melody audio and the melody information. By first selecting a subset of sound sources from a sound source database based on note pitch, and then selecting a second sound source from the first sound source based on note information and note context information, and according to the correspondence between playing fingering and notes; by combining note information and note context information, and according to the correspondence between notes and playing fingering, the sound source with matching fingering is determined for the notes, making the final synthesized music more natural and in line with the actual playing characteristics of the instrument; and, users only need to provide the melody they want to synthesize, and an audio file of the instrument playing that melody can be quickly generated.
[0105] In some specific embodiments, the performance technique rules include a first technique rule representing the correspondence between the target position of a note in a musical phrase or beat and the fingering, and a second technique rule representing the correspondence between dynamics and string signs and fingering; the second sound source determination module 13 may specifically include:
[0106] The target first sound source filtering unit is used to filter out the target first sound source that matches the fingering of the current note from the first sound source according to the note context information of the current note and the first technique rule.
[0107] The second sound source filtering unit is used to filter out a second sound source that matches the fingering of the current note from the target first sound source according to the note information of the current note and the second technique rule.
[0108] In some specific embodiments, the target position may specifically include the first note of a musical phrase, the last note of a musical phrase, or a note located on a stressed beat;
[0109] The music generation device may further include:
[0110] The position determination unit is used to determine the position information of the current note based on the note context information of the current note before selecting a target first sound source that matches the fingering of the current note from the first sound source according to the first technique rule based on the note context information of the current note.
[0111] The first execution unit is used to determine whether the current note belongs to the target position based on the position information. If it belongs to the target position, the unit executes the step of selecting a target first sound source that matches the fingering of the current note from the first sound source according to the note context information of the current note and the first technique rule.
[0112] The second execution unit is used to, if it does not belong to the target position, select a second sound source from the first sound source that matches the fingering of the current note according to the second technique rule, based on the note information of the current note.
[0113] In some specific embodiments, the position determination unit may specifically include:
[0114] The time difference determination unit is used to, if it does not belong to the target position, select a second sound source from the first sound source that matches the fingering of the current note according to the second technique rule based on the note information of the current note;
[0115] The phrase first note determination unit is used to determine the current note as the phrase first note if the first time difference is greater than one beat.
[0116] The phrase last note determination unit is used to determine the current note as the phrase last note if the second time difference is greater than one beat.
[0117] In some specific embodiments, the position determination unit may specifically include:
[0118] The accent determination unit is used to determine that the current note is on an accent if the time difference between the current note and the start time of the accent within the current time signature is less than a preset duration threshold; the start time of the accent within the time signature is an element in the accent start time sequence; the accent start time sequence is generated by combining the time signature and the tempo with a linearly increasing sequence.
[0119] In some specific embodiments, the target type of musical instrument may specifically be a guqin;
[0120] The music generation device may further include:
[0121] The pitch range conversion unit is used to determine whether the pitch range in the melody information is within the range of the guqin pitch range; if it is not within the range of the guqin pitch range, the pitch range of the melody audio is converted to the range of the guqin pitch range by raising or lowering an octave.
[0122] In some specific embodiments, the music generation device may specifically include:
[0123] An interval query unit is used to query, after generating music for the target type instrument corresponding to the melody audio, whether there are two connected notes on the same string and the note interval is less than an interval threshold.
[0124] A decay processing unit is used to decay the first note of the two connected notes if present.
[0125] Furthermore, this application also discloses an electronic device, see [link to relevant documentation]. Figure 3 As shown, the content in the figure should not be considered as any limitation on the scope of use of this application.
[0126] Figure 3 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the music generation method disclosed in any of the foregoing embodiments.
[0127] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0128] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include operating system 221, computer program 222 and data 223 including melody audio, etc. The storage method can be temporary storage or permanent storage.
[0129] The operating system 221 manages and controls the various hardware devices on the electronic device 20 and the computer program 222 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the music generation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0130] Furthermore, this application also discloses a computer storage medium storing computer-executable instructions. When the computer-executable instructions are loaded and executed by a processor, they implement the music generation method steps disclosed in any of the foregoing embodiments.
[0131] Furthermore, this application also discloses a computer program product, including a computer program that, when executed by a processor, implements the music generation method steps disclosed in any of the foregoing embodiments.
[0132] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0133] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0134] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0135] The above provides a detailed description of the music generation method, device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for generating music, characterized in that, include: Acquire the melody audio and extract the musical information from the melody audio; The music information includes note information, note context information, and melody information for different notes in the melody audio. Each note in the melody audio is traversed sequentially, and based on the note information of the current note, the first sound source that matches the pitch of the current note is selected from the sound source database of the target type of instrument. Based on the note information and note context information of the current note, a second sound source is selected from the first sound source according to the playing technique rules corresponding to the target type of instrument; the playing technique rules include the correspondence between playing fingering and notes; The matching sound source corresponding to the current note is determined based on the second sound source; Based on the matching sound source corresponding to each note in the melody audio and the melody information, music for the target type instrument corresponding to the melody audio is generated.
2. The music generation method according to claim 1, characterized in that, The performance technique rules include a first technique rule that represents the correspondence between the target position of a note in a musical phrase or beat and the fingering, and a second technique rule that represents the correspondence between dynamics and string signs and fingering. The step of selecting a second sound source from the first sound source based on the note information and note context information of the current note, according to the playing technique rules corresponding to the target type of instrument, includes: Based on the note context information of the current note, and in accordance with the first technique rule, a target first sound source that matches the fingering of the current note is selected from the first sound source; Based on the note information of the current note, and in accordance with the second technique rule, a second sound source matching the fingering of the current note is selected from the target first sound source.
3. The music generation method according to claim 2, characterized in that, The target positions include the first note of a musical phrase, the last note of a musical phrase, and notes located on a stressed beat; Before selecting a target first sound source matching the fingering of the current note from the first sound source according to the first technique rule based on the note context information of the current note, the method further includes: The position information of the current note is determined based on the note context information of the current note; Based on the position information, determine whether the current note belongs to the target position. If it belongs to the target position, then execute the step of selecting a target first sound source that matches the fingering of the current note from the first sound source according to the first technique rule based on the note context information of the current note. If it does not belong to the target position, then according to the note information of the current note, and in accordance with the second technique rule, a second sound source that matches the fingering of the current note is selected from the first sound source.
4. The music generation method according to claim 3, characterized in that, Based on the note context information of the current note, the position information of the current note is determined, including: Based on the note context information of the current note, determine the first time difference between the current note and the previous note, and the second time difference between the current note and the next note; If the first time difference is greater than the time of one beat, then the current note is determined to be the first note of the musical phrase; If the second time difference is greater than the time of one beat, then the current note is determined to be the last note of the musical phrase.
5. The music generation method according to claim 3, characterized in that, Based on the note context information of the current note, the position information of the current note is determined, including: If the time difference between the current note and the start time of the stressed beat in the current time signature is less than a preset duration threshold, then the current note is determined to be on a stressed beat. The start time of the repetition within the time signature is an element in the repetition start time sequence; the repetition start time sequence is generated by combining the time signature and music tempo with a linearly increasing sequence.
6. The music generation method according to claim 1, characterized in that, The target type of musical instrument is the guqin; Before sequentially traversing each note in the melody audio, the following is also included: Determine whether the range of the melody information is within the range of the guqin. If the melody is not within the range of the guqin, then the range of the melody audio is converted to the range of the guqin by raising or lowering the octave.
7. The music generation method according to claim 1, characterized in that, The process of generating the sound source database for the target type of musical instrument includes: Obtain the sound source data of the target type of musical instrument; The sound source data is cleaned; the sound source cleaning includes pitch value checking and sound source start point checking. The cleaned audio source data is expanded to obtain an audio source database.
8. The music generation method according to any one of claims 1 to 7, characterized in that, After generating music for the target type instrument corresponding to the melody audio, the process further includes: Query whether there exist two consecutive notes on the same string and the note interval is less than the interval threshold in the music of the target type instrument corresponding to the melody audio; If they exist, the first note of the two connected notes will be gradually weakened.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the music generation method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein the computer programs, when executed by a processor, implement the music generation method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Music generation method and device, electronic equipment and storage medium
CN112820254A
Music generation method and device, equipment and storage medium
CN117153132A