A music generation method and related apparatus

By extracting the fundamental frequency and segmenting the notes from the humming audio, a reference note sequence is generated, and melody continuation and arrangement are performed. This solves the problem that existing technologies cannot create music based on humming melodies, and achieves high efficiency and accuracy in automatic music production.

CN116994544BActive Publication Date: 2026-08-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-09-30
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing music generation methods cannot create music based on users' hummed melodies, thus failing to meet users' actual needs.

Method used

By extracting the fundamental frequency and segmenting the notes in the humming audio to be processed, a basic note sequence is generated. The reference beat point is determined within the time interval, and the note timing information is adjusted to generate the reference note sequence. The melody is then continued, arranged, and loaded with timbres to generate the target audio.

Benefits of technology

It enables automatic music production based on humming audio, improving the efficiency and accuracy of music production, lowering the threshold for music production, and preserving the rhythm and characteristics of humming melodies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994544B_ABST
    Figure CN116994544B_ABST
Patent Text Reader

Abstract

The application discloses a music generation method and related device, and relates to the fields of artificial intelligence, machine learning and the like. A fundamental note sequence of first notes is obtained by performing fundamental frequency extraction and note segmentation on to-be-processed humming audio, the first notes are calibrated to obtain a reference note sequence, the melody embodied by the obtained reference note sequence has better rhythm, meanwhile, note time information of multiple first notes is adjusted, which can improve the adjustment efficiency of the first notes, and then the reference note sequence can be subjected to melody continuation, orchestration and loading of tone color to obtain target audio, automatic music production based on to-be-processed humming audio is realized, the demand of music production is targetedly met, and the threshold of music production is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a method and apparatus for generating music. Background Technology

[0002] With the development of computer and network technologies, digital technologies can increasingly provide music services, enriching people's entertainment lives and promoting the development of music creation. Music enthusiasts often spontaneously hum melodies, which can reflect their mood and inspiration at the time. Using these hummed melodies as the basis for music creation can preserve these feelings and inspiration. However, current music generation methods cannot be based on existing hummed melodies for music creation, thus failing to meet practical needs. Summary of the Invention

[0003] To address the aforementioned technical problems, this application provides a music generation method and related apparatus, which enables the automatic production of music based on the humming audio to be processed, specifically meeting the needs of music production and lowering the threshold for music production.

[0004] The embodiments of this application disclose the following technical solutions:

[0005] On the one hand, this application provides a music generation method, the method comprising:

[0006] The fundamental frequency of the humming audio to be processed is extracted and the note is segmented to obtain a basic note sequence including N first notes. The basic note sequence is used to indicate the note timing information of the N first notes, where N is an integer greater than 1.

[0007] Within the time interval corresponding to the basic note sequence, a beat point sequence including multiple reference beat points is determined, wherein the time length between two adjacent reference beat points is a preset multiple of the time length of a fractional note.

[0008] The N first notes are calibrated sequentially according to time to obtain a reference note sequence; when calibrating the i-th first note among the N first notes, the note timing information of the i-th first note is synchronously adjusted according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note; where i is a positive integer less than or equal to N;

[0009] The melody is continued by the reference note sequence to obtain a complete note sequence including the reference note sequence;

[0010] Arrange the complete note sequence to obtain an arranged note sequence;

[0011] The target audio is obtained by loading the target timbre into the arranged note sequence.

[0012] On the other hand, this application provides a music generation apparatus, the apparatus comprising:

[0013] The transcription unit is used to extract the fundamental frequency and segment the notes of the humming audio to be processed, so as to obtain a basic note sequence including N first notes. The basic note sequence is used to indicate the note timing information of the N first notes, where N is an integer greater than 1.

[0014] A beat point sequence determination unit is used to determine a beat point sequence including multiple benchmark beat points within the time interval corresponding to the basic note sequence, wherein the time length between two adjacent benchmark beat points is a preset multiple of the time length of a note.

[0015] A calibration unit is used to calibrate the N first notes sequentially in chronological order to obtain a reference note sequence. When calibrating the i-th first note among the N first notes, the unit synchronously adjusts the note timing information of the i-th to N-th first notes according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note; where i is a positive integer less than or equal to N.

[0016] A melody continuation unit is used to continue the melody of the reference note sequence to obtain a complete note sequence including the reference note sequence.

[0017] The arrangement unit is used to arrange the complete note sequence to obtain an arranged note sequence;

[0018] The loading unit is used to load the target timbre into the arranged note sequence to obtain the target audio.

[0019] On the other hand, this application provides a computer device, the device including a processor and a memory:

[0020] The memory is used to store computer programs and to transfer the computer programs to the processor;

[0021] The processor is configured to execute the music generation method described above according to instructions in the computer program.

[0022] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program for executing the music generation method described above.

[0023] On the other hand, embodiments of this application provide a computer program product including a computer program, which, when run on a computer device, causes the computer device to execute the music generation method.

[0024] As can be seen from the above technical solution, by extracting the fundamental frequency and segmenting the notes of the humming audio to be processed, a basic note sequence including N first notes can be obtained. The basic note sequence can represent the melody of the humming audio to be processed. A beat point sequence including multiple reference beat points is determined within the time interval corresponding to the basic note sequence. The time length between two adjacent reference beat points is a multiple of the note fraction. The reference beat points are used to represent the beat information within the time interval corresponding to the basic note sequence. Then, the N first notes are calibrated sequentially according to time to obtain the reference note sequence. When calibrating the i-th first note among the N first notes, the note timing information of the i-th first note is adjusted synchronously according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note. Calibrating the first notes can make the melody of the obtained reference note sequence have a better rhythm than the basic note sequence. Adjusting multiple first notes at the same time can improve the adjustment efficiency of the first notes, and while optimizing the rhythm, the time interval between the first notes can be preserved to the greatest extent, thus preserving the melodic characteristics of the humming audio to be processed.

[0025] The melody can then be continued from the reference note sequence to obtain a complete note sequence that includes the reference note sequence. This complete note sequence is a complete melody based on the humming audio to be processed, and since the reference note sequence has a good rhythm, the resulting complete note sequence can better capture the rhythm of the reference note sequence. Based on this, a highly relevant continuation can be performed, resulting in higher accuracy in melody continuation. Next, the complete note sequence is arranged to obtain an arranged note sequence. A target timbre is then loaded onto the arranged note sequence to obtain the target audio. The target audio is generated based on the humming audio to be processed, has a related melody, and contains richer content than the humming audio to be processed. This enables the automatic production of music based on the humming audio to be processed, specifically meeting the needs of music production and lowering the barrier to entry for music production. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram illustrating an application scenario of a music generation method provided in an embodiment of this application;

[0028] Figure 2 A flowchart illustrating a music generation method provided in this application embodiment;

[0029] Figure 3 A schematic diagram of a basic note sequence provided in an embodiment of this application;

[0030] Figure 4 A schematic diagram illustrating the calibration process for N first notes provided in an embodiment of this application;

[0031] Figure 5 A schematic diagram of a beat point sequence after one calibration, provided in an embodiment of this application;

[0032] Figure 6 A schematic diagram of the beat point sequence after two calibrations provided in an embodiment of this application;

[0033] Figure 7 A schematic diagram of the beat point sequence after three calibrations provided in an embodiment of this application;

[0034] Figure 8 A schematic diagram of a note sequence after adjusting the note length, provided as an embodiment of this application;

[0035] Figure 9 A schematic diagram illustrating the calibration process for the i-th first note provided in an embodiment of this application;

[0036] Figure 10 A schematic diagram illustrating another calibration process for the i-th first note provided in an embodiment of this application;

[0037] Figure 11 A schematic diagram of a two-dimensional piano roller blind provided in an embodiment of this application;

[0038] Figure 12 This is a schematic diagram illustrating the working process of a melody generation model provided in an embodiment of this application;

[0039] Figure 13 This is a schematic diagram illustrating the operation of another melody generation model provided in an embodiment of this application;

[0040] Figure 14 A structural block diagram of a music generation device provided in an embodiment of this application;

[0041] Figure 15 A structural diagram of a terminal device provided in an embodiment of this application;

[0042] Figure 16 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation

[0043] The embodiments of this application will now be described with reference to the accompanying drawings.

[0044] Currently, the melodies that users hum can reflect their mood and inspiration at the time. If the hummed melody is used as the basis for music creation, the user's mood and inspiration can be preserved. However, the current music generation method cannot create music based on existing hummed melodies, which cannot meet the actual needs.

[0045] To address the aforementioned technical problems, this application provides a music generation method and related apparatus. The method involves extracting the fundamental frequency and segmenting the notes of the humming audio to be processed to obtain a basic note sequence. The first notes are then calibrated to obtain a reference note sequence, resulting in a melody with better rhythm. Simultaneously, the note timing information of multiple first notes is adjusted to improve the adjustment efficiency. Subsequently, the reference note sequence can be used for melody continuation, arrangement, and timbre loading to obtain the target audio. This achieves automatic music production based on the humming audio to be processed, specifically meeting the needs of music production and lowering the barrier to entry for music production.

[0046] The music generation method provided in this application is based on Artificial Intelligence (AI). AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0047] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0048] In the embodiments of this application, the main artificial intelligence software technologies involved include the aforementioned machine learning / deep learning directions. For example, it may involve deep learning in machine learning (ML), including various artificial neural networks (ANNs).

[0049] The music generation method provided in this application can be implemented using a computer device, which can be a terminal device or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminal devices include, but are not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, and this application does not impose any limitations on this connection.

[0050] This computer device, equipped with data processing capabilities, possesses machine learning capabilities. Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0051] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, and intelligent transportation. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0052] To facilitate understanding of the technical solutions provided in this application, the following section will introduce a music generation method provided in an embodiment of this application, in conjunction with a practical application scenario.

[0053] See Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of a music generation method provided in an embodiment of this application. Figure 1The application scenario shown includes a server 10 and a terminal device 20. The terminal device 20 has an application installed for music generation. The server 10 and the terminal device 20 can interact via a network. The server 10 generates target audio based on the audio to be processed, and the terminal device 20 interacts with the user, obtains the humming audio to be processed, and displays and plays the generated target audio.

[0054] Server 10 can perform transcription operations on the humming audio to be processed. The transcription operations include fundamental frequency extraction and note segmentation to obtain a basic note sequence including N first notes 201. The basic note sequence can represent the melody of the humming audio to be processed. Within the time interval corresponding to the basic note sequence, a sequence of beat points including multiple reference beat points (represented by grid lines 202) is determined. The time length between two adjacent reference beat points is a multiple of a preset note. The reference beat points are used to represent the beat information within the time interval corresponding to the basic note sequence.

[0055] Afterwards, server 10 can calibrate the N first notes sequentially according to time to obtain a reference note sequence. When calibrating the i-th first note among the N first notes, the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note are adjusted synchronously to make the i-th first note aligned with the reference beat point corresponding to the i-th first note. Calibrating the first notes can make the melody reflected by the obtained reference note sequence have a better rhythm than the basic note sequence. Adjusting multiple first notes at the same time can improve the adjustment efficiency of the first notes, and while optimizing the rhythm, the time interval between the first notes can be preserved to the greatest extent, thus preserving the melodic characteristics of the humming audio to be processed.

[0056] Afterwards, server 10 can perform melody continuation on the reference note sequence to obtain a complete note sequence including the reference note sequence. In this way, the complete note sequence is a complete melody based on the humming audio to be processed, and the reference note sequence has a good rhythm. The obtained complete note sequence can also better capture the rhythm of the reference note sequence, and perform continuation with good relevance based on this. Therefore, the melody continuation has higher accuracy.

[0057] Afterwards, server 10 arranges the complete note sequence to obtain the arranged note sequence, loads the target timbre into the arranged note sequence to obtain the target audio. The target audio is generated based on the humming audio to be processed, has a related melody to the humming audio to be processed, and has richer content than the humming audio to be processed. This realizes the automatic production of music based on the humming audio to be processed, specifically meets the needs of music production, and lowers the threshold of music production.

[0058] Next, a music generation method provided by an embodiment of this application will be described in conjunction with the accompanying drawings. See also: Figure 2 , Figure 2 A flowchart of a music generation method provided in this application embodiment, the method including:

[0059] S101, perform fundamental frequency extraction and note segmentation on the humming audio to be processed, to obtain a basic note sequence including N first notes.

[0060] In this embodiment, the humming audio to be processed is obtained by a user humming, and its vocal characteristics correspond to vocal notes in a song. These vocal notes represent the pitch and duration of the humming audio. A continuous vocal segment of a certain duration in the humming audio corresponds to a single vocal note. The number of bars in the humming audio is within a certain range, for example, greater than or equal to 1 and less than or equal to a preset number. When the preset number is 3, the humming audio can include 1 bar, 2 bars, or 3 bars. The humming audio to be processed can be extracted from an initial audio file, which can have a longer duration. The humming audio to be processed can be in formats such as CompactDisc (CD), Windows Media Audio (WAV), or Moving Picture Experts Group Audio Layer III (MP3).

[0061] The humming audio to be processed can be transcribed, converting it into a sequence of basic notes. This transcription process identifies the vocal notes in the audio, extracting their fundamental tones as the first notes. The basic note sequence can include multiple first notes, each corresponding to a different vocal note in the humming audio, thus conveying the melody of the audio. The number of first notes can be denoted as N, where N is an integer greater than 1. A first note, as a musical symbol, is used to record notes of varying lengths. Each first note has pitch and timing information. The pitch information represents the frequency of the first note, while the timing information reflects the time interval corresponding to the first note. The timing information can include at least two of the following: the start time, the end time, and the actual duration of the note.

[0062] The transcription operation performed on the humming audio is a process of converting audio into musical notation, which can include two steps: fundamental frequency extraction and note segmentation. Specifically, fundamental frequency extraction can extract the fundamental frequency trajectory from the humming audio, which reflects the frequency changes in the audio. Note segmentation can divide the fundamental frequency trajectory into multiple independent first notes, which are arranged sequentially in time to form a basic note sequence. The fundamental frequency refers to the frequency of the fundamental tone in a polyphonic voice. The human voice, as a polyphonic voice, is composed of several notes. Among these notes, the fundamental tone has the lowest frequency and the highest intensity. The pitch of the polyphony is determined by the fundamental frequency. Most music is polyphonic, which makes the music richer and more rounded, rather than sharp and harsh.

[0063] Specifically, the fundamental frequency (F0) of the humming audio to be processed can be estimated frame by frame. This allows for the precise output of an estimated value for each note in the frame, thereby extracting the fundamental frequency of the humming audio and obtaining a continuous fundamental frequency trajectory within the playback time of the humming audio. Frame-by-frame fundamental frequency estimation of the humming audio to be processed can be achieved, for example, using the YIN algorithm for fundamental frequency (F0) estimation.

[0064] Specifically, the fundamental frequency (FFM) of the humming audio to be processed can be extracted to obtain the FFM trajectory within the playback time interval of the audio. Within the target time interval, there are multiple candidate FFMs, each with a predicted probability. Based on the predicted probabilities of these candidate FFMs and the FFM trajectories of other time intervals adjacent to the target time interval, the first FFM is selected as the target FFM for that target time interval. This method determines the target FFM for each time interval, and the updated FFM trajectory is obtained from the target FFM. This method of determining the target FFM using predicted probabilities and the FFM trajectories of other adjacent time intervals provides a reliable alternative explanation for the FFM determination process within the target time interval, eliminating short-term errors and resulting in a smoother updated FFM trajectory. The FFM trajectory can be obtained using the pYIN algorithm, an improved version of the YIN algorithm. The predicted probabilities naturally arise from the prior distribution of the threshold parameters of YIN, resulting in lower algorithm complexity. The pYIN algorithm can be implemented in C++. In practice, these predicted probabilities can be used as observations in the hidden Markov model in the pYIN algorithm, which is then decoded by Viterbi to generate the updated fundamental frequency trajectory.

[0065] After obtaining the fundamental frequency trajectory or the updated fundamental frequency trajectory using the aforementioned fundamental frequency extraction method, the fundamental frequency trajectory or the updated fundamental frequency trajectory can be segmented into notes to obtain first notes corresponding to multiple time periods. Then, pitch information and note timing information can be determined for the first notes in multiple time periods. Specifically, for the j-th first note out of N first notes, the pitch information of the j-th first note can be determined based on the target fundamental frequency corresponding to the time period to which the j-th first note belongs, and the note timing information of the j-th first note can be determined based on the time period to which the j-th first note belongs, where j is a positive integer less than or equal to N. Thus, after determining the pitch information and note timing information for each first note, multiple first notes with both pitch and note timing information can be obtained.

[0066] After acquiring multiple first notes with pitch and note timing information, a basic note sequence can be constructed based on these first notes. The first notes in the basic note sequence are ordered chronologically. This basic note sequence indicates the note timing information of the N first notes, thus reflecting the rhythmic information of the humming audio to be processed. It also indicates the pitch information of the N first notes, thus reflecting the intonation information of the humming audio to be processed. The basic note sequence can exist in Musical Instrument Digital Interface (MIDI) format. MIDI was proposed in the early 1980s to solve the communication problem between electroacoustic instruments and is the most widely used music standard format in the music production industry. It can be called "computer-understandable sheet music," and it can record music using digital control signals for notes.

[0067] refer to Figure 3 The diagram shown is a schematic representation of a basic note sequence provided in an embodiment of this application. The basic note sequence is visually represented by a piano roll. The horizontal axis of the piano roll represents time, and the vertical axis represents pitch. The basic note sequence includes four first notes 201, denoted as N1, N2, N3, and N4. The position of any first note 201 in the piano roll is determined by its pitch information and note timing information. The pitch information determines the vertical coordinate (position height) of the first note 201 in the piano roll. For example, the pitch of the first note N1 is higher than the pitches of the first notes N2, N3, and N4, and the pitches of the first notes N2, N3, and N4 decrease sequentially. The note timing information determines the horizontal length of the first note 201 in the piano roll. For example, the duration of the first notes N1 and N2 is greater than the duration of the first note N3, but less than the duration of the first note N4. Of course, the piano roll is only used to visually illustrate the relationship between the first notes. In actual operation, it is not necessary to draw the first notes in the form of a piano roll.

[0068] S102, determine a beat point sequence including multiple reference beat points within the time interval corresponding to the basic note sequence.

[0069] After obtaining the basic note sequence, a beat point sequence including multiple reference beat points can be determined within the time interval corresponding to the basic note sequence. The time length between two adjacent reference beat points is a preset multiple of the time length of a fractional note. That is, the beat point sequence including multiple reference beat points is determined based on a preset multiple of fractional notes. In this way, there is a fixed time interval between the reference beat points. Each reference beat point is used to represent the beat information within the time interval corresponding to the basic note sequence, and the beat point sequence is used to represent the reference rhythm. In addition, the larger the value of the aforementioned preset multiple, the smaller the time length between two adjacent reference beat points, indicating that the time interval corresponding to the basic note sequence is divided into smaller granularities, and the subsequent adjustment accuracy of the note timing information is also higher.

[0070] The preset multiples of note values ​​can be eighth notes, sixteenth notes, thirty-second notes, etc. Different note values ​​represent different durations. The durations of eighth notes, sixteenth notes, and thirty-second notes are 1 / 8, 1 / 16, and 1 / 32 of a whole note, respectively. For example, in 4 / 4 time, each measure consists of 4 beats, with a quarter note being one beat, meaning each measure contains 4 quarter notes. In 2 / 4 time, each measure consists of 2 beats, with a quarter note being one beat, meaning each measure contains 2 quarter notes.

[0071] Specifically, the time length between two adjacent reference beat points can be preset. In practice, the time length between two adjacent reference beat points can be determined based on the rhythmic information required in the subsequent melody continuation process. For example, if the subsequent melody continuation process requires rhythmic information based on sixteenth notes, the time length between two adjacent reference beat points can be set to the time length of a sixteenth note, that is, the preset multiple of the note is a sixteenth note.

[0072] Specifically, the time length between two adjacent benchmark beats can also be determined based on the rhythmic information of the basic note sequence. In practice, the time length between two adjacent benchmark beats can be determined based on the minimum note length of the first note in the basic note sequence. For example, a note with the same time length as the minimum note length of the first note can be used as a note with a preset multiple, or a note with a time length closest to the minimum note length of the first note can be used as a note with a preset multiple. Alternatively, it can be determined based on the minimum note length of the first note in the basic note sequence and the number of first notes with the minimum note length. For example, when the number of first notes with the minimum note length reaches a preset number, a note with the same time length as the minimum note length of the first note can be used as a note with a preset multiple, or a note with a time length closest to the minimum note length of the first note can be used as a note with a preset multiple. When the number of first notes with the minimum note length does not reach a preset number, a note with a smaller multiple than the aforementioned note with a time length closest to the minimum note length of the first note can be used as a note with a preset multiple. In this way, the time length between two adjacent reference beat points can be equal to or close to the minimum note length of the first note, which can ensure that each first note can have good adjustment accuracy.

[0073] For example, if the minimum note length of the first note in the basic note sequence is the duration of a thirty-second note, then the duration between two adjacent reference beat points can be set to the duration of a thirty-second note, i.e., a preset multiple of the note length is a thirty-second note. Or, if the minimum note length of the first note in the basic note sequence is the duration of a thirty-second note, and there are multiple first notes with the minimum note length, then the duration between two adjacent reference beat points can be set to the duration of a thirty-second note, i.e., a preset multiple of the note length is a thirty-second note. Or, if the minimum note length of the first note in the basic note sequence is the duration of a thirty-second note, and there are only a few first notes with the minimum note length (less than the preset number, e.g., only one), then the first note with the minimum note length can be deleted, and the duration between two adjacent reference beat points can be set to the duration of a sixteenth note, i.e., a sixteenth note that is a multiple of the thirty-second note can be set as a preset multiple of the note length.

[0074] Multiple benchmark beat points have their own timing information, including the beat point time, which indicates the time of the benchmark beat point within the time interval of the basic note sequence. (Reference) Figure 2As shown, a measure consists of four beats, as indicated by the vertical lines P1-P3 (P0 is not included as it is the start of the measure). P4 exists after P3 but is not shown in the diagram. A measure also includes four beat intervals (the interval between any two adjacent vertical lines), corresponding to the time interval of that measure. When the preset multiple of the note value is a sixteenth note, the time length between two adjacent reference beat points is a sixteenth note. The four beat intervals corresponding to a measure are divided into sixteen parts, that is, each beat interval is divided into four parts, as shown by the dotted lines in the diagram.

[0075] Figure 2 In the diagram, vertical lines P0-P4 and the dashed lines between them serve as grid lines dividing time intervals. These lines characterize a reference beat point 202, and their positions indicate the beat time of that reference beat point. All sixteen reference beat points 202 are arranged sequentially to form the aforementioned beat point sequence. The reference beat points are numbered starting from 0. The reference beat point with number 0 is represented by the vertical line shown in P0, the reference beat point with number 4 is represented by the vertical line shown in P1, the reference beat point with number 8 is represented by the vertical line shown in P2, the reference beat point with number 12 is represented by the vertical line shown in P3, and the reference beat points 14, 15, and 16 are not shown. In fact, the reference beat point with number 16 is represented by the vertical line shown in P4.

[0076] from Figure 2 It can be seen that the start and end times of the first note obtained through transcription represent its absolute time in the humming audio to be processed, and are not aligned with the grid lines, that is, not aligned with the reference beat points indicated by the grid lines. This will cause the basic note sequence to sound lacking in rhythm, and will also lead to a decrease in the effect of subsequent melody continuation. This is because the method of melody continuation is usually implemented for a rhythmic intro.

[0077] S103, calibrate the N first notes in chronological order to obtain the reference note sequence.

[0078] After determining the basic note sequence including the first note and the beat point sequence including multiple reference beat points, the basic note sequence can be adjusted based on the multiple reference beat points. Specifically, the reference note sequence can be obtained by calibrating N first notes in chronological order, so that the reference note sequence has a better rhythm than the basic note sequence, which improves the sense of rhythm and has a certain sound "beautification" effect. It also helps to improve the accuracy of subsequent melody continuation.

[0079] The N first notes are calibrated sequentially according to time. This can be done either from front to back or from back to front. The calibration operation performed on the i-th first note is denoted as the i-th calibration operation. The following explanation uses the calibration of the i-th first note as an example, where i is a positive integer less than or equal to N.

[0080] When calibrating the i-th first note out of N first notes, the note timing information of the i-th first note is adjusted synchronously with the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note. This allows for the simultaneous adjustment of multiple first notes, improving the adjustment efficiency of the first notes. Furthermore, while optimizing the rhythm, the time interval between the first notes can be preserved to the greatest extent, thus preserving the melodic characteristics of the humming audio to be processed.

[0081] In determining the reference beat point for the i-th first note, if i is 1, then the reference beat point for the i-th first note can be the reference beat point located at the initial position in the beat point sequence. Thus, by calibrating the first first note, it can be aligned to the reference beat point located at the initial position, which can be the 0th reference beat point. During the calibration of the first first note, the note timing information of the first to Nth first notes can be adjusted simultaneously to align the first first note to the 0th reference beat point.

[0082] In determining the reference beat point for the i-th first note, if i is greater than 1, after calibrating the first i-1 first notes, the reference beat point for the i-th first note can be determined based on the synchronized adjustment of the note timing information of the i-th first note and the timing information of the reference beat point. For example, after calibrating the 1st first note, the note timing information of the 2nd to Nth first notes has been synchronized once. At this time, the reference beat point for the 2nd first note can be determined based on the synchronized adjustment of the note timing of the 2nd first note and the timing information of the reference beat point. Then, the note timings of the 2nd to Nth first notes can be synchronized to align the 2nd first note with its corresponding reference beat point.

[0083] refer to Figure 4The diagram illustrates the calibration process for N first notes according to an embodiment of this application. This process may include: 1031. Determining the reference beat point corresponding to the first first note; 1032. Adjusting the note timing information of the first to Nth first notes; 1033. Determining the reference beat point corresponding to the second first note based on the synchronized adjusted note timing information of the second first note; 1034. Adjusting the note timing information of the second to Nth first notes; and so on; 1035. Determining the reference beat point corresponding to the Nth first note based on the synchronized adjusted note timing information of the Nth first note; 1036. Adjusting the note timing information of the Nth first note. This achieves the calibration of N first notes. Of course, if the note timing information of a certain first note, after being synchronized, aligns the first note to its corresponding reference beat point, then calibration of that first note is unnecessary. Therefore, the synchronous adjustment of multiple first notes can improve adjustment efficiency.

[0084] When i is greater than 1, the corresponding reference beat point can be determined based on the start time of the i-th first note. Aligning the i-th first note to the corresponding reference beat point can be specifically achieved by adjusting the start time of the i-th first note to the beat point time of its corresponding reference beat point. Specifically, based on the synchronized adjustment of the i-th first note's note timing information, the synchronized adjustment start time of the i-th first note can be determined. Then, the reference beat point whose beat point time is closest to the start time of the i-th first note can be selected as the reference beat point corresponding to the i-th first note. Thus, the synchronous adjustment of the note timing information of the i-th to N-th first notes can be specifically achieved by adjusting the note timing information of the i-th to N-th first notes according to the time interval and relative order between the synchronized start time of the i-th first note and the beat point time of the reference beat point corresponding to the i-th first note. This ensures that the synchronized start time of the i-th first note is adjusted to match the beat point time of the reference beat point corresponding to the i-th first note, thus aligning the i-th first note with the corresponding reference beat point. Synchronous adjustment refers to adjustments based on the same adjustment direction and the same adjustment step size. From the perspective of the first notes displayed in the piano roll, the synchronized adjustment of the i-th to N-th first notes can mean either being synchronized forward by the same distance or synchronized backward by the same distance.

[0085] For example, the time interval and relative order between the first note and the 0th reference beat can be determined. Based on this, the calibration strategy for the first note is determined to be an adjustment forward by t1. Then, the first to Nth notes can be moved forward by t1, with the adjustment direction referenced. Figure 3 The arrows below the first notes N1, N2, N3, and N4 indicate this. (See reference.) Figure 5The diagram shows the beat point sequence after one calibration in an embodiment of this application. After moving the first notes N1, N2, N3, and N4 forward by t1, the first note N1 is aligned with the 0th reference beat point. After one calibration, the time interval and relative order between the second first note and its corresponding reference beat point can be determined. Based on this, the calibration strategy for the first first note is determined to be moving it forward by t2. Then, the second to Nth first notes can be moved forward by t2, and so on, until the Nth first notes are calibrated.

[0086] refer to Figure 5 As shown, after adjusting the first to Nth first notes forward by t1, the second first note is aligned with the second reference beat point. Therefore, the calibration of the second first note N2 can be skipped, i.e., t2 is 0, and the calibration of the third first note can continue. During the calibration of the third first note, the time interval and relative order between the third first note and the fifth reference beat point can be determined. Based on this, the calibration strategy for the third first note is determined to be moving it forward by t3. Thus, the third to Nth first notes can be moved forward by t3. The adjustment directions of the first notes N3 and N4 are shown by the arrows below. (Reference) Figure 6 The diagram shown is a schematic of the beat point sequence after two calibrations in an embodiment of this application. After moving the first note N3 and N4 forward by t3, the first note N3 is aligned with the 5th reference beat point.

[0087] Based on music theory experience, the beat type of the benchmark beat point is divided into strong beat and weak beat. If the actual length of the first note is relatively long, then the benchmark beat point with a strong beat type will sound better. Therefore, in the process of determining the benchmark beat point corresponding to the i-th first note, if the actual length of the i-th first note is greater than the time length of a preset multiple of the fractional notes, that is, greater than the preset benchmark beat point interval, then the i-th first note is considered to be relatively long. The benchmark beat point whose beat point time is closest to the start time of the i-th first note and whose beat point type is strong can be determined as the benchmark beat point corresponding to the i-th first note. Specifically, when the preset multiple of the note is a sixteenth note, the even-numbered reference beat points (e.g., the 0th, 2nd, 4th, 6th, 8th, 10th, 12th, 14th, and 16th reference beat points) can be designated as strong beats, and the odd-numbered reference beat points (e.g., the 1st, 3rd, 5th, 7th, 9th, 11th, 13th, and 15th reference beat points) can be designated as unstressed beats; when the preset multiple of the note is another note, the strong and unstressed beats can be determined according to the actual situation.

[0088] refer to Figure 6As shown, after moving the first notes N3 and N4 forward by t3, the starting position of the first note N4 is closest to the 7th reference beat point. However, the actual length of the first note N4 is relatively long, so it needs to be aligned to a reference beat point of type accent, such as the 6th or 8th reference beat point. Therefore, among the 6th and 8th reference beat points, the starting position of the first note N4 is closest to the 8th reference beat point, and the 8th reference beat point is taken as the reference beat point corresponding to the first note N4. Next, the time interval and relative order between the 4th first note and the 8th reference beat point can be determined. Based on this, the calibration strategy for the 4th first note N4 is determined to be moving backward by t4. The adjustment direction of the first note N4 is shown by the arrow below. (Reference) Figure 7 The diagram shown is a schematic of the beat point sequence after three calibrations in an embodiment of this application. After moving the first note N4 backward by t4, the first note N4 is aligned with the 8th reference beat point.

[0089] In this embodiment, the note length of the first note can also be adjusted. This adjustment can be made before or at the initial adjustment point. Specifically, based on the synchronized adjustment time information of the i-th first note, the actual note length and the synchronized termination time of the i-th first note can be determined. Then, a target note length can be determined for the i-th first note. The target note length is the note length that the i-th first note needs to be adjusted to. The target note length is typically an integer multiple of the duration of a predefined note. Using the predefined multiple of the note duration as a reference length, the target note length can be determined as an integer number of reference lengths closest to the actual note length. Then, the synchronized termination time of the i-th first note can be adjusted to the target time, thus adjusting the note length of the i-th first note from its actual length to its target length, further optimizing the rhythm of the note sequence.

[0090] For example, after adjusting the first to Nth first notes forward by t1, the ending time of the first first note can be adjusted so that the note length of the first first note is an integer number of reference lengths; after adjusting the second to Nth first notes forward by t2, the ending time of the second first note can be adjusted so that the note length of the second first note is an integer number of reference lengths, and so on.

[0091] refer to Figure 7As shown, after moving the fourth first note N4 backward by t4, the first note N4 aligns with the eighth reference beat point. Its start time is the same as the beat point of the eighth reference beat point, but its end time is different from the beat points of all reference beat points. Therefore, its note length is not an integer number of reference lengths. At this point, it can be determined that the actual note length of the fourth first note is close to 4 reference lengths. Thus, the target note length of the fourth first note N4 can be determined to be 4 reference lengths. By adjusting the end time of the fourth first note N4 to the beat point of the 12th reference beat point, the note length of the fourth first note N4 can be adjusted to 4 reference lengths. (Reference) Figure 8 The diagram shown is a schematic of a note sequence after adjusting the note length according to an embodiment of this application. After adjusting the termination time of the fourth first note N4, the note length of the fourth first note N4 is adjusted to 4 reference lengths.

[0092] Since adjusting the ending time of the i-th first note will change the note length of the i-th first note from the actual note length to the previously determined target note length, if the note length of the i-th first note is increased, the adjusted note length may make the i-th first note meet the condition of aligning to the downbeat. Therefore, after adjusting the note timing information of the i-th to N-th first notes, the note timing information of the i-th to N-th first notes can be adjusted again to make the i-th first note align to the reference beat point with the nearest beat point type being downbeat.

[0093] refer to Figure 9The diagram illustrates a calibration process for the i-th first note according to an embodiment of this application. This process may include: 1031' Determining the reference beat point corresponding to the i-th first note based on its start time and the beat point time of the reference beat point. Specifically, the reference beat point whose beat point time is closest to the start time of the i-th first note can be used as the reference beat point corresponding to the first first note; 1032' Adjusting the note timing information of the i-th to N-th first notes based on the reference beat point corresponding to the i-th first note. Specifically, the i-th to N-th first notes can be moved synchronously so that the start time of the i-th first note is adjusted to its corresponding reference beat point time. 1033' Adjust the ending time of the i-th first note to give it a target note length; 1034' Determine the updated target note length based on the target note length. Specifically, if the target note length is greater than or equal to a preset target length, then the target note length corresponding to the i-th first note is an accented target note, which is then used as the updated target note length for the i-th first note; 1035' Adjust the note timing information of the i-th to N-th first notes again based on the updated target note length corresponding to the i-th first note.

[0094] Of course, the target note length can be determined before determining the reference beat point corresponding to the i-th first note. This way, after determining the target note length, the reference beat point corresponding to the i-th first note can be determined based on the target note length, reducing the number of adjustments to the note timing information of the i-th first note. Specifically, if the target note length is long, it means the i-th first note will be adjusted to a longer note length. Therefore, the reference beat point with an accented beat type corresponding to the i-th first note can also be used. Thus, if the target note length of the i-th first note is greater than or equal to a preset reference length, the reference beat point whose beat point timing is closest to the start time of the i-th first note and whose beat point type is accented can be determined as the reference beat point corresponding to the i-th first note.

[0095] refer to Figure 10The diagram illustrates another calibration process for the i-th first note provided in this application embodiment. This process may include: 1031”, determining the reference beat point corresponding to the i-th first note based on the start time of the i-th first note, the beat point time of the reference beat point, and the target note length. Specifically, the reference beat point whose beat point time is closest to the start time of the i-th first note can be used as the reference beat point corresponding to the first first note. If the target note length is greater than a preset reference length, the beat point time closest to the start time of the i-th first note can be determined. 1032”, Based on the benchmark point of the i-th first note, adjust the note timing information of the i-th to N-th first notes. Specifically, the i-th to N-th first notes can be moved synchronously so that the start time of the i-th first note is adjusted to the beat point time of its corresponding benchmark point, thus aligning the i-th first note with its corresponding benchmark point; 1033”, Adjust the end time of the i-th first note so that the i-th first note has the target note length.

[0096] S104, continue the melody of the reference note sequence to obtain a complete note sequence including the reference note sequence.

[0097] After calibrating the first note of N in the basic note sequence to obtain the reference note sequence, the reference note sequence, as a note sequence with good rhythm and able to reflect the user's humming characteristics, can be used as the prelude (Prime) of the score. By continuing the melody of the reference note sequence, a complete note sequence including the reference note sequence can be obtained. In this way, the complete note sequence is a complete melody based on the humming audio to be processed, and the reference note sequence has good rhythm, which improves the quality of melody continuation. The obtained complete note sequence can also better capture the rhythm of the reference note sequence, and based on this, a continuation with good correlation can be performed. Therefore, the melody continuation has higher accuracy.

[0098] Melody continuation from a baseline note sequence can be achieved using a melody generation model. This model can be a machine learning model, such as a deep learning model. The input to this model is the baseline note sequence, and the output is the complete note sequence. The melody generation model can be an end-to-end music melody generation model, such as the Music Transformer and CP Transformer models, or a multi-stage melody generation model, such as the CMT model.

[0099] Specifically, the melody generation model can construct a one-dimensional event sequence based on a baseline note sequence. This sequence represents the start and end times, pitch, and dynamics of each note in the baseline note sequence. Subsequent notes are then constructed based on this one-dimensional event sequence to obtain a complete note sequence. This working mode is suitable for end-to-end music melody generation models.

[0100] The basic idea of ​​the Music Transformer model is to model MIDI event sequences (i.e., one-dimensional event sequences). The Transformer model is a classic sequence modeling model with an attention mechanism as its core module. Broadly speaking, it belongs to the autoregressive generative model category, primarily using self-attention and sinusoidal positional information. Each layer consists of a self-attention layer and a feedforward sublayer. Since the Transformer model has been widely used in text sequence modeling, and music is essentially a sequence of notes, modeling note sequences can also be achieved using the Transformer. Specifically, a two-dimensional piano roll can be represented as a one-dimensional event sequence, thus allowing direct use of the Transformer model to learn and continue the melody.

[0101] refer to Figure 11 The diagram shown is a schematic of a two-dimensional piano roll shutter provided in an embodiment of this application. The horizontal axis represents time (in seconds), and the vertical axis represents pitch. The note sequence in this diagram can be represented by the following one-dimensional event sequence: Initially, the playing intensity (SET_VELOCITY) is set to 80, and the note at pitch 60 is started (NOTE_ON). 500ms later (TIME_SHIFT), the note at pitch 64 is started (NOTE_ON). 500ms later (TIME_SHIFT), the note at pitch 67 is started (NOTE_ON). The note at pitch 60 is switched forward (TIME_SHIFT) for 1000ms, the note at pitch 64 is switched forward (NOTE_OFF), the note at pitch 67 is switched forward (TIME_SHIFT) for 500ms, the playing velocity is set to 100, the note at pitch 65 is started (NOTE_ON), ​​the note at pitch 65 is switched forward (TIME_SHIFT) for 500ms, and the note at pitch 65 is switched forward (NOTE_OFF).

[0102] The basic idea of ​​CP Transformer and Music Transformer is the same, except that CP Transformer incorporates more information at each time step. For example, it can incorporate at least one of the following: dynamics, duration, pitch, chord, beat, tempo, and measure. This incorporation not only allows the model to learn more information (such as note dynamics) but also shortens the sequence, reducing the difficulty of modeling long sequences. (Reference) Figure 12 The diagram shown is a schematic representation of the working process of a melody generation model provided in an embodiment of this application. Figure 12 A represents a schematic diagram of the information fused at each time step, where the horizontal axis represents different times. Information corresponding to the same time step can be fused. Taking six time steps as an example, T = 6. Figure 12 B represents the working principle of melody continuation, where w t-1,1 to w t-1,K Let K represent the K pieces of information to be fused at time step t-1. These pieces of information are concatenated (concave) and then input into a linear layer for processing. The result of the linear layer processing is combined with the positional encoding result to obtain the fused result. The fused result is processed using a Transformer Causal self-attention layer. The hidden feature h is obtained through processing. t Hidden feature h t As the fusion feature corresponding to time step t, it can be divided into multiple pieces of information to be fused corresponding to time step t through a linear layer. to Thus, based on the information corresponding to the previous time step, the information corresponding to the next time step is predicted, thereby realizing the continuation of the melody.

[0103] Specifically, the melody generation model decodes a reference note sequence to obtain the note timing and pitch information of the first note. Based on this information, the model can then continue the melody of the reference note sequence. This working mode is suitable for multi-stage melody generation models. The training dataset for the melody generation model can include training data from multiple musical styles, giving the model the ability to continue in various styles. It can adaptively continue the melody of reference note sequences of different styles, generating complete note sequences of different styles. The musical style of the complete note sequences is at least one of the multiple musical styles in the training data. This allows it to adapt to user input with varying musical styles or humming audio involving multiple musical styles. Furthermore, the amount of training data can be large, such as using millions of training data points, making the melody generation model's continuation ability considerable and resulting in high-quality melody continuation results. Multiple musical styles can represent multiple genres or different emotions.

[0104] In practice, the melody generation model is a CMT model, which includes a rhythm decoder and a pitch decoder. Rhythm and pitch are decoded step-by-step, with the rhythm information determined based on the note timing information of the reference note sequence. (Reference) Figure 13 The diagram shown is a schematic representation of another melody generation model provided in this application embodiment. Figure 13 A is a schematic diagram of the melody continuation process. This model also includes a chord encoder, used to obtain the existing chord sequence c. 1:T The length of the chord sequence is the same as the length of the complete note sequence, denoted by T. T can be 8 measures, each measure has 4 beats, and each beat includes 4 sixteenth notes. The chord sequence can include 128 note groups corresponding to different times. The notes in the same note group correspond to the same time. When a note group includes multiple notes, the note group is a chord.

[0105] The working process of the rhythm decoder is represented by dashed lines. The chord encoder works based on the chord sequence c. 1:T Encoding yields a hidden representation of the note information for the first t-1 note groups of the chord sequence. The input to the rhythm decoder also includes the rhythm information r of the t-1 notes that have already been written. 1:t-1 The hidden representation of the note information of the first t-1 note groups The rhythmic information r of the t-1 notes that have already been written. 1:t-1 It is possible to predict the rhythmic information of the t-th note. The rhythm information of the t-th note is output through the Rhythm Output Layer. Prediction Probalities

[0106] The solid lines represent the working process of the pitch decoder, while the chord encoder uses the chord sequence c. 1:T Encoding yields a hidden representation of the note information for the first t-1 note groups of the chord sequence. And input the pitch decoder, and input the existing chord sequence c 1:T Hidden representation The input to the rhythm decoder also includes the rhythm information r of the existing chord sequence. 1:T The rhythm decoder can obtain the hidden representation of the rhythm information of the completed t-1 notes. The pitch decoder uses a hidden representation of the note information from the first t-1 note groups of the input chord sequence. Hidden representation of the rhythmic information of the completed t-1 notes The pitch information p of the t-1 notes that have been completed and written. 1:t-1 The pitch information of the i-th note is predicted as follows: The pitch information of the t-th note is output through the Pitch Output Layer. Prediction Probalities

[0107] refer to Figure 13 B, any one of the rhythm decoder, pitch decoder, and chord encoder, can include N processing layers from input to output. Each processing layer can include a multi-head attention layer, a residual and normalization layer (Add&Norm), a feed forward network, and a residual and normalization layer (Add&Norm) in sequence.

[0108] In this embodiment, multiple note sequences can be obtained by continuing the melody of a reference note sequence, denoted as first note sequences. By scoring each first note sequence, the aforementioned complete note sequence can be determined from the first note sequences. Specifically, the reference note sequence can be selectively continued to obtain multiple first note sequences including the reference note sequence. A comprehensive score for the first note sequence is determined based on at least one of its rhythm score and similarity score. Based on the comprehensive score, the complete note sequence is determined from the first note sequences. The rhythm score of the first note sequence reflects whether its rhythm is good; a first note sequence with good rhythm will have a higher rhythm score. The similarity score reflects whether the repetition of the first note sequence is good; repetition at specific positions can constitute a memorable point, and the corresponding similarity score can be higher.

[0109] Specifically, the rhythm score can be determined based on the matching relationship between multiple benchmark beat points in the first note sequence and the scoring beat point sequence. This scoring beat point sequence is determined based on the note timing information of the first note sequence within the time interval corresponding to the first note sequence. Unlike the beat point sequence corresponding to the basic note sequence in S102, the scoring beat point sequence corresponds to the first note sequence. The method for determining the scoring beat point sequence can refer to the method for determining the beat point sequence in S102, and will not be elaborated here. The similarity score is determined based on the similarity between the first preset measure and the second preset measure in the first note sequence. The first preset measure can be the first measure, and the second preset measure can be the fifth measure. Since a section includes four measures, the first measure is the first measure in the first section, and the fifth measure is the first measure in the second section. The similarity between the two makes the first section and the second section have a high degree of repetition. In addition, the first measure originates from the benchmark note sequence, while the fifth measure is often obtained by continuation. Therefore, the similarity between the first measure and the fifth measure can reflect the correlation between the continuation part and the intro part.

[0110] The higher the overall score of the first note sequence, the better the listening experience. Therefore, the top M notes with the highest overall scores in the first note sequence can be used as the complete note sequence to improve the quality of melody continuation. M can be 1 or greater than 1, so the number of complete note sequences can be one or more.

[0111] In this embodiment, before determining the comprehensive score of the first note sequence, the generated note sequence can be filtered to remove note sequences that do not conform to music theory, thus obtaining the first note sequence. This ensures that the complete note sequence will not be a result with obvious music theory errors and reduces invalid calculations in the scoring process of the first note sequence. Specifically, by continuing the melody based on the benchmark note sequence, multiple second note sequences including the benchmark note sequence can be obtained. From these multiple second note sequences, note sequences that meet the music theory filtering conditions are removed to obtain multiple first note sequences. The note sequences that meet the music theory filtering conditions do not conform to common music theory knowledge. The music theory filtering conditions can include at least one of the following conditions: having a minor second interval, having a non-key note, and having at least one note in at least one measure. Here, an interval refers to the relationship between two notes in pitch; a minor second interval is a semitone difference between two adjacent notes; a minor second interval is a dissonant interval; a non-key note is a note that does not belong to a key.

[0112] In addition, in S101, the first note in the basic note sequence can be adjusted based on music theory filtering conditions. For example, when the basic note sequence has a minor second interval or an out-of-key note, the pitch information of the first note can be adjusted. Furthermore, in S101, if there are at least one or more measures in the basic note sequence with a number of notes less than 1, the measure can be deleted from the basic note sequence.

[0113] S105, arranges the complete note sequence to obtain the arranged note sequence.

[0114] After obtaining the complete note sequence, it can be arranged to produce an arranged note sequence. Specifically, target chord notes can be matched to target notes in the complete note sequence. Target chord notes include multiple notes corresponding to the same time period, which can be obtained by adding notes of other frequencies to the target notes. The arranged note sequence can then be obtained based on the target chord notes, or, based on the note timing information of the complete note sequence, accompaniment track notes can be matched from the accompaniment library. The arranged note sequence can then be obtained based on the target chord notes and the target accompaniment track notes. Accompaniment track notes can include drum accompaniment notes, bass accompaniment notes, string accompaniment notes, etc. Matching accompaniment track notes to the complete note sequence can be achieved through pre-set rules. For example, the rhythm information of the complete note sequence can be determined by the note timing information, and accompaniment track notes can be matched based on the rhythm information. This allows different accompaniment track notes to be matched for different complete note sequences.

[0115] An arrangement of notes can include one target chord note or multiple target chord notes. Multiple consecutive target chord notes constitute a chord progression, which is also called a chord progression. It is a sequence of multiple target chord notes, such as the canon chord progression C, G, Am, Em, F, C, F, G.

[0116] In an arrangement of notes, a single target chord note can be a block chord or an arpeggio. An arrangement can contain only block chords, only arpeggios, or both. In a block chord, all notes occur simultaneously, meaning each note has the same start and end time. When a block chord is repeated rhythmically, it appears as a series of blocks on the staff. An arpeggio, on the other hand, refers to notes played sequentially, meaning each note has a slightly different start and end time. It is one of the decorative or stylized methods for chord progressions. Compared to block chords, arpeggios sound more fluid, less mechanical, and more like a genuine human composition when played.

[0117] To match the target chord notes in a complete note sequence with the corresponding target chord notes, a chord generation model can be used. When the target chord notes in the arrangement note sequence include arpeggios, the complete note sequence can be input into the chord generation model, which will output an arrangement note sequence including arpeggios. The chord generation model can adopt a seq2seq architecture and can be trained using a public dataset (such as the POP909 dataset).

[0118] S106 loads the target timbre into the arranged note sequence to obtain the target audio.

[0119] After arranging a complete note sequence to obtain an arranged note sequence, a target timbre can be loaded onto the arranged note sequence to obtain the target audio. By loading the target timbre, the arranged note sequence can be converted into a playable target audio. For example, the arranged note sequence may be in MIDI format, while the target audio may be in MP3, CD, WAV, or other formats. The target audio is generated based on the humming audio to be processed. It has a related melody to the humming audio and contains richer content than the humming audio to be processed. This enables the automatic production of music based on the humming audio to be processed, specifically meeting the needs of music production and lowering the barrier to entry for music production.

[0120] In this embodiment, mixing can also be added to the complete note sequence. This allows loading the target timbre and adding mixing to the complete note sequence to obtain the target audio. Specifically, a SoundFont format timbre library can be automatically loaded into the complete note sequence. Adding mixing to the complete note sequence can be achieved through convolution. Adding mixing can achieve a reverb effect to enhance the spatial sense of the music, making the target audio closer to the live music performance.

[0121] In this embodiment, music-assisted creation can be achieved. Users only need to hum a melody into the terminal device to obtain a complete musical work based on that melody, realizing automatic music production, specifically meeting the needs of music production, and lowering the threshold for music production.

[0122] Based on the music generation method provided in the embodiments of this application, the embodiments of this application also provide a music generation apparatus, see reference. Figure 14 The diagram shown is a structural block diagram of a music generation device provided in an embodiment of this application. The music generation device 1300 includes:

[0123] Transcription unit 1301 is used to extract the fundamental frequency and segment the note in the humming audio to be processed, to obtain a basic note sequence including N first notes, wherein the basic note sequence is used to indicate the note timing information of the N first notes, and N is an integer greater than 1;

[0124] The beat point sequence determination unit 1302 is used to determine a beat point sequence including multiple benchmark beat points within the time interval corresponding to the basic note sequence, wherein the time length between two adjacent benchmark beat points is a preset multiple of the time length of a note.

[0125] The calibration unit 1303 is used to calibrate the N first notes sequentially according to time to obtain a reference note sequence; when calibrating the i-th first note among the N first notes, the calibrator adjusts the note timing information of the i-th to N-th first notes synchronously according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note; where i is a positive integer less than or equal to N;

[0126] Melody continuation unit 1304 is used to perform melody continuation on the reference note sequence to obtain a complete note sequence including the reference note sequence;

[0127] Arrangement unit 1305 is used to arrange the complete note sequence to obtain an arranged note sequence;

[0128] Loading unit 1306 is used to load the target timbre into the arranged note sequence to obtain the target audio.

[0129] Optionally, the device further includes:

[0130] The reference beat point determination unit is used to determine the corresponding reference beat point for the i-th first note after calibrating the first i-1 first notes among the plurality of first notes, based on the synchronized note timing information of the i-th first note and the timing information of the reference beat point, if i is greater than 1.

[0131] Optionally, the reference beat point determination unit includes:

[0132] The start time determination unit is used to determine the start time of the synchronized adjustment of the i-th first note based on the synchronized note time information of the i-th first note after calibrating the first i-1 first notes among the plurality of first notes if i is greater than 1.

[0133] The reference beat point determination subunit is used to determine the reference beat point whose beat point time is closest to the start time of the i-th first note, and use it as the reference beat point corresponding to the i-th first note;

[0134] The calibration unit 1303 is specifically used for:

[0135] When calibrating the i-th first note among the N first notes, the note timing information of the i-th to N-th first notes is adjusted according to the time interval and relative order between the synchronized start time of the i-th first note and the beat point time of the reference beat point corresponding to the i-th first note, so that the synchronized start time of the i-th first note is adjusted to the beat point time of the reference beat point corresponding to the i-th first note.

[0136] Optionally, the square device further includes:

[0137] The termination time determination unit is used to determine the actual length of the i-th first note and the termination time after synchronization adjustment based on the synchronized adjustment note time information of the i-th first note during the calibration process of the i-th first note among the N first notes.

[0138] The target length determination unit is used to determine the target length of the note for the i-th first note, wherein the target length of the note is an integer number of reference lengths that are closest to the actual length of the note, and the reference length is the time length of the sub-note of the preset multiple;

[0139] The termination time adjustment unit is used to adjust the termination time of the i-th first note after the synchronized start time of the i-th first note is adjusted to the beat point time of the reference beat point corresponding to the i-th first note, and then adjusts the synchronized termination time of the i-th first note to the target time.

[0140] Optionally, the reference beat point determination subunit is specifically used for:

[0141] If the actual length of the note is greater than or equal to a preset number of the reference lengths, or the target length of the note is greater than or equal to a preset number of the reference lengths, then the reference beat point whose beat point time is closest to the start time of the i-th first note and whose beat point type is an accent is determined as the reference beat point corresponding to the i-th first note.

[0142] Optionally, the basic note sequence is also used to indicate the pitch information of the N first notes, and the transcription unit 1301 includes:

[0143] The fundamental frequency prediction unit is used to extract the fundamental frequency of the humming audio to be processed, and obtain the fundamental frequency trajectory within the playback time interval of the humming audio to be processed. There are multiple candidate fundamental frequencies within the target time period of the playback time interval, and the candidate fundamental frequencies have prediction probabilities.

[0144] The base frequency determination unit is used to select a first base frequency from the multiple candidate base frequencies as the target base frequency in the target time period based on the predicted probability of the multiple candidate base frequencies and the base frequency trajectory in other time periods adjacent to the target time period.

[0145] The note segmentation unit is used to segment the updated fundamental frequency trajectory obtained based on the target fundamental frequency into notes, so as to obtain first notes corresponding to multiple time periods respectively.

[0146] The note information determination unit is used to determine the pitch information and note timing information of the first notes in the multiple time periods; wherein, for the j-th first note among the N first notes, the pitch information of the j-th first note is determined according to the target fundamental frequency corresponding to the time period to which the j-th first note belongs, and the note timing information of the j-th first note is determined according to the time period to which the j-th first note belongs, where j is a positive integer less than or equal to N.

[0147] Optionally, the melody continuation unit 1304 includes:

[0148] The continuation unit performs melody continuation on the reference note sequence to obtain multiple first note sequences including the reference note sequence;

[0149] A scoring unit is used to determine a comprehensive score for the first note sequence based on at least one of a rhythm score and a similarity score; the rhythm score is determined based on the matching relationship between the first note sequence and multiple reference beat points in a scoring beat point sequence, and the scoring beat point sequence is determined based on the note timing information of the first note sequence within the time interval corresponding to the first note sequence; the similarity score is determined based on the similarity between a first preset measure and a second preset measure in the first note sequence.

[0150] The selection unit is used to determine the complete note sequence from the first note sequence based on the comprehensive score of the first note sequence.

[0151] Optionally, the continuation unit includes:

[0152] A continuation subunit is used to continue the melody based on the reference note sequence, thereby obtaining multiple second note sequences including the reference note sequence;

[0153] A filtering unit is used to remove the note sequences to be filtered that meet the music theory filtering conditions from a plurality of second note sequences, thereby obtaining a plurality of first note sequences. The music theory filtering conditions include at least one of the following: having a minor second interval, having a non-key tone, and having at least one note in a measure with less than one note.

[0154] Optionally, the basic note sequence is also used to indicate the pitch information of the N first notes, and the continuation unit is specifically used for:

[0155] The melody generation model decodes the reference note sequence to obtain the note timing information and pitch information of the first note. The training dataset of the melody generation model includes training data of multiple music styles. Based on the note timing information and pitch information of the first note, the melody generation model continues the melody of the reference note sequence to obtain multiple first note sequences including the reference note sequence. The music style of the first note sequence is at least one of the multiple music styles.

[0156] Optionally, the arrangement unit 1305 includes:

[0157] The chord note acquisition unit is used to match the target chord note corresponding to the target note according to the target note in the complete note sequence, wherein at least one of the target chord notes is an arpeggiated chord note;

[0158] The arrangement subunit is used to obtain the arrangement note sequence based on the target chord notes.

[0159] Optionally, the arrangement unit 1305 includes:

[0160] A chord note determination unit is used to generate a target chord note corresponding to the target note based on at least one target note in the complete note sequence;

[0161] The accompaniment track note determination unit is used to match target accompaniment track notes from the accompaniment library according to the note timing information of the complete note sequence.

[0162] The arrangement subunit is used to obtain the arrangement note sequence based on the target chord notes and the target accompaniment track notes.

[0163] Optionally, the loading unit 1306 is specifically used for:

[0164] The target timbre is loaded and mixed to obtain the target audio by loading the complete note sequence.

[0165] Therefore, by extracting the fundamental frequency and segmenting the notes in the humming audio to be processed, a basic note sequence including N first notes can be obtained. The basic note sequence can represent the melody of the humming audio to be processed. Within the time interval corresponding to the basic note sequence, a beat point sequence including multiple reference beat points is determined. The time length between two adjacent reference beat points is a multiple of the note fraction. The reference beat points are used to represent the beat information within the time interval corresponding to the basic note sequence. Then, the N first notes are calibrated sequentially according to time to obtain the reference note sequence. When calibrating the i-th first note among the N first notes, the note timing information of the i-th first note is adjusted synchronously according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note. Calibrating the first notes can make the melody of the obtained reference note sequence have a better rhythm than the basic note sequence. Adjusting multiple first notes at the same time can improve the adjustment efficiency of the first notes, and while optimizing the rhythm, the time interval between the first notes can be preserved to the greatest extent, thus preserving the melodic characteristics of the humming audio to be processed.

[0166] The melody can then be continued from the reference note sequence to obtain a complete note sequence that includes the reference note sequence. This complete note sequence is a complete melody based on the humming audio to be processed, and since the reference note sequence has a good rhythm, the resulting complete note sequence can better capture the rhythm of the reference note sequence. Based on this, a highly relevant continuation can be performed, resulting in higher accuracy in melody continuation. Next, the complete note sequence is arranged to obtain an arranged note sequence. A target timbre is then loaded onto the arranged note sequence to obtain the target audio. The target audio is generated based on the humming audio to be processed, has a related melody, and contains richer content than the humming audio to be processed. This enables the automatic production of music based on the humming audio to be processed, specifically meeting the needs of music production and lowering the barrier to entry for music production.

[0167] This application also provides a computer device, which is the computer device described above, and may include a terminal device or a server. The aforementioned music generation device may be configured in this computer device. The computer device will now be described in conjunction with the accompanying drawings.

[0168] If the computer device is a terminal device, please refer to Figure 15 As shown, this application provides a terminal device, taking a mobile phone as an example:

[0169] Figure 15 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 15 The mobile phone includes components such as a radio frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a Wi-Fi module 1470, a processor 1480, and a power supply 1490. Those skilled in the art will understand that... Figure 15 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0170] The following is combined Figure 15 A detailed introduction to each component of a mobile phone:

[0171] The RF circuit 1410 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1480; in addition, it transmits uplink data to the base station.

[0172] The memory 1420 can be used to store software programs and modules. The processor 1480 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 1420. The memory 1420 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 1420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0173] The input unit 1430 can be used to receive input numeric or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.

[0174] The display unit 1440 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1440 may include a display panel 1441.

[0175] The mobile phone may also include at least one sensor 1450, such as a light sensor, a motion sensor, and other sensors.

[0176] Audio circuitry 1460, speaker 1461, and microphone 1462 provide an audio interface between the user and the mobile phone.

[0177] WiFi is a short-range wireless transmission technology. Through the WiFi module 1470, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access.

[0178] The processor 1480 is the control center of the mobile phone. It connects to various parts of the mobile phone through various interfaces and lines. It performs various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 1420 and calling data stored in the memory 1420.

[0179] The mobile phone also includes a power supply 1490 (such as a battery) that powers the various components.

[0180] In this embodiment, the processor 1480 included in the terminal device also has the following functions:

[0181] The fundamental frequency of the humming audio to be processed is extracted and the note is segmented to obtain a basic note sequence including N first notes. The basic note sequence is used to indicate the note timing information of the N first notes, where N is an integer greater than 1.

[0182] Within the time interval corresponding to the basic note sequence, a beat point sequence including multiple reference beat points is determined, wherein the time length between two adjacent reference beat points is a preset multiple of the time length of a fractional note.

[0183] The N first notes are calibrated sequentially according to time to obtain a reference note sequence; when calibrating the i-th first note among the N first notes, the note timing information of the i-th first note is synchronously adjusted according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note; where i is a positive integer less than or equal to N;

[0184] The melody is continued by the reference note sequence to obtain a complete note sequence including the reference note sequence;

[0185] Arrange the complete note sequence to obtain an arranged note sequence;

[0186] The target audio is obtained by loading the target timbre into the arranged note sequence.

[0187] If the computer device is a server, this application embodiment also provides a server; please refer to [link to relevant documentation]. Figure 16 As shown, Figure 16 This is a structural diagram of a server 1500 provided in an embodiment of this application. The server 1500 can vary significantly due to different configurations or performance. It may include one or more processors 1522, such as a Central Processing Unit (CPU), a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the processor 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations in the storage media 1530 on the server 1500.

[0188] Server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0189] The steps performed by the server in the above embodiments can be based on Figure 16 The server structure shown.

[0190] In addition, this application embodiment also provides a storage medium for storing a computer program for executing the method provided in the above embodiment.

[0191] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the methods provided in the above embodiments.

[0192] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disk, etc., and other media capable of storing program code.

[0193] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0194] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Moreover, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating music, characterized in that, The method includes: The fundamental frequency of the humming audio to be processed is extracted and the note is segmented to obtain a basic note sequence including N first notes. The basic note sequence is used to indicate the note timing information of the N first notes, where N is an integer greater than 1. Within the time interval corresponding to the basic note sequence, a beat point sequence including multiple reference beat points is determined, wherein the time length between two adjacent reference beat points is a preset multiple of the time length of a fractional note. The N first notes are calibrated sequentially according to time to obtain a reference note sequence. When calibrating the i-th first note among the N first notes, the note timing information of the i-th first note is synchronously adjusted according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note; where i is a positive integer less than or equal to N; the timing information of the reference beat point includes the beat point timing. If i is greater than 1, after calibrating the first i-1 first notes among the N first notes, based on the synchronized adjustment note timing information of the i-th first note, determine the synchronized adjustment start time of the i-th first note; determine the reference beat point whose beat point time is closest to the start time of the i-th first note, as the reference beat point corresponding to the i-th first note; the step of synchronously adjusting the note timing information of the i-th to N-th first notes according to the note timing information of the i-th first note and the time information of the reference beat point corresponding to the i-th first note includes: adjusting the note timing information of the i-th to N-th first notes according to the time interval and relative order between the synchronized adjustment start time of the i-th first note and the beat point time of the reference beat point corresponding to the i-th first note, so that the synchronized adjustment start time of the i-th first note is adjusted to the beat point time of the reference beat point corresponding to the i-th first note; The melody is continued by the reference note sequence to obtain a complete note sequence including the reference note sequence; Arrange the complete note sequence to obtain an arranged note sequence; The target audio is obtained by loading the target timbre into the arranged note sequence.

2. The method according to claim 1, characterized in that, In the process of calibrating the i-th first note among the N first notes, the method further includes: Based on the synchronized timing information of the i-th first note, determine the actual length of the i-th first note and the synchronized termination time. A target note length is determined for the i-th first note, wherein the target note length is an integer number of reference lengths that are closest to the actual note length, and the reference length is the time length of the sub-note of the preset multiple; After the start time of the synchronized adjustment of the i-th first note is adjusted to the beat point time of the reference beat point corresponding to the i-th first note, the end time of the synchronized adjustment of the i-th first note is adjusted to the target time.

3. The method according to claim 2, characterized in that, The determination of the reference beat point whose time is closest to the start time of the i-th first note, as the reference beat point corresponding to the i-th first note, includes: If the actual length of the note is greater than or equal to a preset number of the reference lengths, or the target length of the note is greater than or equal to a preset number of the reference lengths, then the reference beat point whose beat point time is closest to the start time of the i-th first note and whose beat point type is an accent is determined as the reference beat point corresponding to the i-th first note.

4. The method according to claim 1, characterized in that, The basic note sequence is also used to indicate the pitch information of the N first notes. The process of extracting the fundamental frequency and segmenting the note in the humming audio to be processed to obtain a basic note sequence including N first notes includes: The fundamental frequency of the humming audio to be processed is extracted to obtain the fundamental frequency trajectory within the playback time interval of the humming audio to be processed. There are multiple candidate fundamental frequencies within the target time period of the playback time interval, and the candidate fundamental frequencies have prediction probabilities. Based on the predicted probabilities of the multiple candidate fundamental frequencies and the fundamental frequency trajectories in other time periods adjacent to the target time period, a first fundamental frequency is selected from the multiple candidate fundamental frequencies as the target fundamental frequency for the target time period. The updated fundamental frequency trajectory obtained based on the target fundamental frequency is segmented into notes to obtain first notes corresponding to multiple time periods; The pitch information and note timing information of the first note in the multiple time periods are determined; wherein, for the j-th first note among the N first notes, the pitch information of the j-th first note is determined according to the target fundamental frequency corresponding to the time period to which the j-th first note belongs, and the note timing information of the j-th first note is determined according to the time period to which the j-th first note belongs, where j is a positive integer less than or equal to N.

5. The method according to any one of claims 1-4, characterized in that, The process of continuing the melody from the reference note sequence to obtain a complete note sequence including the reference note sequence includes: The melody is continued by the reference note sequence to obtain multiple first note sequences including the reference note sequence; A comprehensive score for the first note sequence is determined based on at least one of the rhythm score and the similarity score; the rhythm score is determined based on the matching relationship between the first note sequence and multiple reference beat points in the scoring beat point sequence, and the scoring beat point sequence is determined based on the note time information of the first note sequence within the time interval corresponding to the first note sequence; the similarity score is determined based on the similarity between the first preset measure and the second preset measure in the first note sequence. Based on the comprehensive score of the first note sequence, the complete note sequence is determined from the first note sequence.

6. The method according to claim 5, characterized in that, The process of continuing the melody from the reference note sequence to obtain multiple first note sequences including the reference note sequence includes: Based on the reference note sequence, a melody is continued to be written, resulting in multiple second note sequences including the reference note sequence; Remove the note sequences that meet the music theory filtering conditions from the multiple second note sequences to obtain multiple first note sequences. The music theory filtering conditions include at least one of the following: having a minor second interval, having a non-key tone, and having less than 1 note in at least one measure.

7. The method according to claim 5, characterized in that, The basic note sequence is also used to indicate the pitch information of the N first notes. The process of continuing the melody from the basic note sequence to obtain multiple first note sequences including the basic note sequence includes: The melody generation model decodes the reference note sequence to obtain the note timing information and pitch information of the first note. The training dataset of the melody generation model includes training data for multiple music styles. Using a melody generation model, the melody is continued from the reference note sequence based on the note timing information and pitch information of the first note, resulting in multiple first note sequences including the reference note sequence, wherein the musical style of the first note sequence is at least one of the multiple musical styles.

8. The method according to any one of claims 1-4, characterized in that, The process of arranging the complete note sequence to obtain an arranged note sequence includes: Based on the target note in the complete note sequence, match the target chord note corresponding to the target note, wherein at least one of the target chord notes is an arpeggiated chord note; The arrangement note sequence is obtained based on the target chord notes.

9. The method according to any one of claims 1-4, characterized in that, The process of arranging the complete note sequence to obtain an arranged note sequence includes: Generate the target chord note corresponding to the target note based on at least one target note in the complete note sequence; Based on the note timing information of the complete note sequence, match the target accompaniment track notes from the accompaniment library for the complete note sequence; The arrangement note sequence is obtained based on the target chord notes and the target accompaniment track notes.

10. The method according to any one of claims 1-4, characterized in that, The process of loading the target timbre onto the complete note sequence to obtain the target audio includes: The target timbre is loaded and mixed to obtain the target audio by loading the complete note sequence.

11. A music generation device, characterized in that, The device includes: The transcription unit is used to extract the fundamental frequency and segment the notes of the humming audio to be processed, so as to obtain a basic note sequence including N first notes. The basic note sequence is used to indicate the note timing information of the N first notes, where N is an integer greater than 1. A beat point sequence determination unit is used to determine a beat point sequence including multiple benchmark beat points within the time interval corresponding to the basic note sequence, wherein the time length between two adjacent benchmark beat points is a preset multiple of the time length of a note. A calibration unit is used to calibrate the N first notes sequentially in chronological order to obtain a reference note sequence. When calibrating the i-th first note among the N first notes, the unit synchronously adjusts the note timing information of the i-th to N-th first notes according to the note timing information of the i-th first note and the timing information of the reference beat point corresponding to the i-th first note, so that the i-th first note is aligned with the reference beat point corresponding to the i-th first note; where i is a positive integer less than or equal to N; and the timing information of the reference beat point includes the beat point timing. The start time determination unit is used to determine the start time of the synchronized adjustment of the i-th first note based on the synchronized note time information of the i-th first note after calibrating the first i-1 first notes among the N first notes if i is greater than 1. The reference beat point determination subunit is used to determine the reference beat point whose beat point time is closest to the start time of the i-th first note, and use it as the reference beat point corresponding to the i-th first note; The calibration unit is specifically used to: when calibrating the i-th first note among the N first notes, adjust the note timing information of the i-th to N-th first notes according to the time interval and relative order between the synchronized start time of the i-th first note and the beat point time of the reference beat point corresponding to the i-th first note, so that the synchronized start time of the i-th first note is adjusted to the beat point time of the reference beat point corresponding to the i-th first note; A melody continuation unit is used to continue the melody of the reference note sequence to obtain a complete note sequence including the reference note sequence. The arrangement unit is used to arrange the complete note sequence to obtain an arranged note sequence; The loading unit is used to load the target timbre into the arranged note sequence to obtain the target audio.

12. The apparatus according to claim 11, characterized in that, The device further includes: The termination time determination unit is used to determine the actual length of the i-th first note and the termination time after synchronization adjustment based on the synchronized adjustment note time information of the i-th first note during the calibration process of the i-th first note among the N first notes. The target length determination unit is used to determine the target length of the note for the i-th first note, wherein the target length of the note is an integer number of reference lengths that are closest to the actual length of the note, and the reference length is the time length of the sub-note of the preset multiple; The termination time adjustment unit is used to adjust the termination time of the i-th first note after the start time of the synchronized adjustment is adjusted to the beat point time of the reference beat point corresponding to the i-th first note, and then adjust the termination time of the synchronized adjustment to the target time.

13. The apparatus according to claim 12, characterized in that, The reference beat point determination subunit is specifically used for: If the actual length of the note is greater than or equal to a preset number of the reference lengths, or the target length of the note is greater than or equal to a preset number of the reference lengths, then the reference beat point whose beat point time is closest to the start time of the i-th first note and whose beat point type is an accent is determined as the reference beat point corresponding to the i-th first note.

14. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the music generation method according to any one of claims 1-10 according to instructions in the computer program.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the music generation method according to any one of claims 1-10.

16. A computer program product comprising a computer program, characterized in that, When it is run on a computer device, it causes the computer device to perform the music generation method according to any one of claims 1-10.