Methods for generating music audio, computer equipment, and computer storage media
By using neural network technology to generate chord progressions and musical structures, music is generated based on the user's listening scenario, solving the problems of fixed arrangement templates and complex technical terminology, and realizing personalized customization and efficient generation of music.
Patent Information
- Application Number
- CN202411185332.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-08-27
AI Technical Summary
Existing music generation technologies use fixed templates and complex technical terms, which forces users to spend time changing templates or learning technical terms, thus affecting the efficiency of music generation.
By generating chord progressions and musical structures based on the user's listening scenario, using neural networks to generate melodic fragments, determining motif melodic fragments, and adding interpolated melodic fragments to the melodic fragment sequence, an audio file matching the musical scenario is generated.
It enables flexibility and personalized customization of music arrangement, reduces generation time, and improves music generation efficiency and user listening experience.
Smart Images

Figure CN119169978B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing, specifically to a method for generating musical audio, a computer device, and a computer storage medium. Background Technology
[0002] The automatic music generation solutions provided by related technologies typically employ a rule-based approach with multiple pre-set arrangement templates. Users can input a chord progression and generate a complete melody that harmonizes with it, specifying the timbre or instruments used in the accompaniment. However, these solutions have two drawbacks. First, the arrangement templates are fixed and may not meet the user's actual musical needs. To ensure the generated music satisfies the user's listening requirements, additional time is needed to modify the arrangement templates and then use the modified templates to generate the music, leading to delays in music generation. Second, these arrangement templates mostly use industry jargon, which non-professional users cannot understand. Users need to spend time learning the terminology to understand the template's function, which also contributes to delays and low efficiency in music generation.
[0003] Therefore, in addition to the time required for generating music using arrangement templates, some users also need to spend more time modifying or learning the functions of arrangement templates, which leads to a relatively slow music generation process and affects the efficiency of music generation. Summary of the Invention
[0004] This application provides a method for generating music audio, a computer device, and a computer storage medium, which are used to intelligently generate music according to the user's listening scenarios and needs, thereby satisfying the user's personalized music customization.
[0005] The first aspect of this application provides a method for generating musical audio, the method comprising:
[0006] Generate chord progressions and musical structure based on the currently defined musical scene;
[0007] Multiple melody fragments are generated. Based on the musical characteristics represented by each melody fragment, a motivic melody fragment is determined from each melody fragment. The motivic melody fragment serves as a musical motif corresponding to the musical scene.
[0008] The motivic melody fragments are arranged according to the musical structure to reproduce the musical motivator, resulting in a sequence of melody fragments.
[0009] Based on the chord progression, at least one interpolated melody segment is added to the sequence of melody segments to develop the musical motif, resulting in an audio file that matches the musical scene.
[0010] A second aspect of this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method of the first aspect described above.
[0011] A third aspect of this application provides a computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.
[0012] A fourth aspect of this application provides a computer program product that, when run on a computer device, causes the computer device to perform the method described in the first aspect.
[0013] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0014] The computer device generates chord progressions and musical structures based on the currently set musical scene, producing multiple melodic fragments. Based on the musical characteristics represented by each melodic fragment, a motivic melodic fragment is identified from each fragment. This motivic melodic fragment serves as the musical motif corresponding to the musical scene. The motivic melodic fragments are arranged according to the musical structure to reproduce the motivic motif, resulting in a sequence of melodic fragments. At least one interpolated melodic fragment is added to this sequence based on the chord progression to develop the motivic motif, resulting in an audio file that matches the musical scene. Therefore, the arrangement of the music can be flexibly changed according to the user's listening scenario and needs, satisfying the user's personalized music customization. Users only need to set the musical scene; there is no need to change the music arrangement template or understand the professional terminology and music theory of music arrangement. This simplifies the process of generating audio files, reduces the time spent on music generation, improves the efficiency of music generation, and ultimately enhances the user's listening experience. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the network framework in an embodiment of this application;
[0016] Figure 2 This is a flowchart illustrating the method for generating music audio in an embodiment of this application;
[0017] Figure 3 This is a flowchart illustrating one possible method for generating interpolated melody fragments corresponding to interpolated chords in the embodiments of this application;
[0018] Figure 4 This is an exemplary schematic diagram illustrating how melodic fragments serve as compositional motifs and are repeated and reproduced based on musical structure in embodiments of this application.
[0019] Figure 5A schematic diagram illustrating an exemplary display style for the audio density setting interface;
[0020] Figure 6 This is a schematic diagram illustrating a scenario in which target audio events are filled into the timeline of a melody segment sequence, as described in an embodiment of this application.
[0021] Figure 7 for Figure 4 The example shown is a schematic diagram of each interpolated chord and its corresponding interpolated melodic fragment;
[0022] Figure 8 An exemplary schematic diagram illustrating the feature information describing the audio event corresponding to a note in a musical score in this application embodiment;
[0023] Figure 9 This is a schematic diagram illustrating the pitch distribution characteristics of two exemplary melody segments in embodiments of this application;
[0024] Figure 10 This is a schematic diagram of a computer device in an embodiment of this application. Detailed Implementation
[0025] This application provides a method for generating music audio, a computer device, and a computer storage medium, which are used to intelligently generate music according to the user's listening scenarios and needs, thereby satisfying the user's personalized music customization.
[0026] Please see Figure 1 The network framework in this embodiment includes:
[0027] The business server 100 and the terminal cluster; the terminal cluster may include: terminal devices 200a, terminal devices 200b, terminal devices 200c, ..., terminal devices 200n and other terminal devices.
[0028] The aforementioned business server 100 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices (including terminal devices 200a, 200b, 200c, ..., 200n) can be smartphones, tablets, laptops, desktop computers, PDAs, mobile internet devices (MIDs), wearable devices (such as smartwatches, smart bracelets, etc.), smart computers, smart in-vehicle systems, and other smart terminals.
[0029] The service server 100 can establish communication connections with each terminal device in the terminal cluster, and the terminal devices in the terminal cluster can also establish communication connections with each other. In other words, the service server 100 can establish communication connections with each terminal device among terminal devices 200a, 200b, 200c, ..., 200n. For example, terminal device 200a can establish a communication connection with the service server 100. Terminal devices 200a and 200b can establish a communication connection, and terminal devices 200a and 200c can also establish a communication connection. The communication connection method is not limited; it can be established directly or indirectly through wired communication or wireless communication, etc., depending on the actual application scenario. This application does not impose any restrictions on this.
[0030] It should be understood that, such as Figure 1 Each terminal device in the terminal cluster shown can have an application client installed. When the application client runs on each terminal device, it can interact with the business server 100, allowing the business server 100 to receive business data from each terminal device (such as user identity data uploaded by users through the terminal device). This application client can be a music player, karaoke app, browser, social networking app, instant messaging app, live streaming app, game app, short video app, video app, shopping app, novel app, payment app, or any other application client capable of displaying text, images, audio, and video data. The specific application client can be determined based on the actual application scenario requirements and is not limited here. This application client can be a standalone client or an embedded sub-client integrated into a client (such as a music app or karaoke app), depending on the actual application scenario and is not limited here.
[0031] The following will combine Figure 1 The network framework and application scenarios of the embodiments of this application are used to describe the music audio generation method in the embodiments of this application:
[0032] Please see Figure 2 One embodiment of the music audio generation method in this application includes:
[0033] 201. Generate chord progressions and musical structure based on the currently set musical scene;
[0034] The method in this embodiment can be applied to a computer device, which may be... Figure 1The network framework shown includes a service server 100 or various terminal devices. In some embodiments, the computer device can implement the music audio generation method provided in this embodiment by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can also be a native application (APP), i.e., a program that needs to be installed in the operating system to run; it can also be a small program, i.e., a program that only needs to be downloaded to a browser environment to run; or it can be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module, or plugin.
[0035] In one alternative approach, the computer device can provide a display interface for the user to select music scenes, such as learning, sleeping, meditation, working, exercising, and other scenes.
[0036] Chord progression (or "chord progression") refers to the order in which chords are played in a song / piece of music. Different chord progressions carry different emotional connotations, and chord progressions with different emotional connotations can be generated based on the user's selected music scene. For example, various chord progressions suitable for different music scenes or listening scenarios can be preset, allowing the computer device to determine the appropriate chord progression for the user's chosen music scene. Alternatively, a neural network model can be trained using a large set of labeled audio samples with chromatic and chordal colors to obtain a model for generating chord progressions. Inputting the user's selected music scene into this model will output the corresponding chord progression. Another approach is to use traditional machine learning Markov chains to generate chord progressions by listing Markov matrices based on common chord tendencies. This embodiment does not limit the method of generating chord progressions.
[0037] Similarly, a neural network model can be trained using labeled music structure samples and their corresponding music scenes, so that the model can output the corresponding music structure based on the music scene input by the user; or various music structures can be preset to suit different music scenes or listening scenarios, such as the preset ABAB music structure to suit sleep or meditation scenarios, or the preset AAAB music structure to suit study or focus scenarios, etc., so as to determine the music structure that suits the music scene selected by the user.
[0038] Among them, the ABAB and AAAB musical structures refer to musical pieces composed of two different melodic segments (section A and section B) that alternate. Specifically, in an ABAB musical structure, the alternation order of the melodies is ABAB, and in an AAAB musical structure, the alternation order is AAAB. This musical structure of alternating melodic segments is commonly found in narrative songs, using two different themes or melodies to tell a story.
[0039] 202. Generate multiple melody fragments, and determine motivic melody fragments from each melody fragment according to the musical characteristics represented by each melody fragment, wherein the motivic melody fragments serve as musical motifs corresponding to the musical scene;
[0040] When generating melody fragments, neural network models such as Recurrent Neural Networks (RNNs) or Long Short-Term Memory (LSTMs) can be used to generate melody fragments of specific lengths using MIDI (Musical Instrument Digital Interface) as the medium. These fragments contain musical features such as pitch, duration, and start time of notes. All output melody fragments are of uniform length in the temporal domain. These fragments can be parsed and stored as multiple audio events in the generation framework. Each audio event represents a feature of the pitch of a melody fragment, and each event corresponds to a pitch, including pitch name, duration, start time, timbre, and end time.
[0041] Pitch refers to the various levels of pitch of a musical note, that is, the height of the note, and one of the basic characteristics of a sound. The pitch name indicates the height of the note, and can be represented by the note names C, D, E, F, G, A, B, or by the solfège syllables 1, 2, 3, 4, 5, 6, 7.
[0042] After obtaining multiple melodic fragments, a motivic melodic fragment can be determined from the multiple melodic fragments based on the musical characteristics represented by each melodic fragment. For example, a motivic melodic fragment can be determined based on multiple audio events of each melodic fragment, and this motivic melodic fragment serves as a musical motif corresponding to a pre-set musical scene.
[0043] 203. Arrange the motivic melody fragments according to the musical structure to reproduce the musical motivator and obtain a sequence of melody fragments;
[0044] In music generation, the repetition of motifs helps maintain the coherence of the work. By repeatedly presenting the same or similar motifs, a stable musical image can be established in the listener's mind, enhancing the recognizability of the music. Therefore, in this embodiment, after obtaining the motif melody fragments, motif reproduction can be performed based on these fragments. That is, the motif melody fragments are arranged according to the determined musical structure to obtain a sequence of melody fragments, thus achieving motif reproduction based on the musical motif.
[0045] 204. Based on the chord progression, add at least one interpolated melody segment to the melody segment sequence to develop the motif for the musical motif, thereby obtaining a musical audio file that matches the musical scene;
[0046] After obtaining the melody fragment sequence and its multiple interpolated melody fragments, the melody fragment sequence and its multiple interpolated melody fragments can be output as a music audio file to play the corresponding music. For example, the melody fragment sequence and its multiple interpolated melody fragments can be rendered into an audio file using a synthesizer or related sample timbre files for output and playback, thereby enabling the playback of music suitable for the user's selected music scene and meeting the user's personalized listening needs.
[0047] After obtaining the melody fragment sequence, at least one interpolated melody fragment can be added to the sequence based on the determined chord progression. This method corresponds to the repetition of notes in the development of musical motifs, thereby realizing the development of musical motifs. The melody fragment sequence with added interpolated melody fragments can be output as an audio file. Here, an interpolated melody fragment refers to a new melody fragment added between adjacent melody fragments. Since both the interpolated melody fragments and the melody fragment sequence are generated according to the user-defined musical scene, the audio file output based on the interpolated melody fragments and the melody fragment sequence can match the user's defined musical scene, satisfying the user's personalized listening needs.
[0048] In this embodiment, the computer device generates chord progressions and musical structures based on the currently set music scene, generating multiple melodic fragments. Based on the musical characteristics represented by each melodic fragment, a motivic melodic fragment is determined from each melodic fragment. This motivic melodic fragment serves as the musical motif corresponding to the music scene. The motivic melodic fragments are arranged according to the musical structure to reproduce the musical motif, resulting in a melodic fragment sequence. At least one interpolated melodic fragment is added to the melodic fragment sequence based on the chord progression to develop the musical motif, resulting in a musical audio file that matches the music scene. Therefore, the arrangement of the music can be flexibly changed according to the user's listening scene and needs, satisfying the user's personalized music customization. Thus, the user only needs to set the music scene, without needing to change the music arrangement template or understand the professional terminology and music theory of music arrangement. This simplifies the process of generating the music audio, reduces the time spent on music generation, improves the efficiency of music generation, and ultimately enhances the user's listening experience.
[0049] based on Figure 2In a preferred embodiment of the illustrated example, when adding at least one interpolated melody fragment to a melody fragment sequence, at least one interpolated chord can be added to the melody fragment sequence based on a determined chord progression, and an interpolated melody fragment corresponding to the interpolated chord can be generated. The melody fragment sequence and its multiple interpolated melody fragments are then output as a music audio file that matches the music scene set by the user.
[0050] In this sequence of melodic fragments, to ensure a smoother and more harmonious transition between fragments, achieving a natural flow from one fragment to another and creating a gradual change in melody, an interpolation chord can be determined for each pair of adjacent melodic fragments within the defined chord progression. This interpolation chord can be a chord that theoretically matches the chord corresponding to the adjacent melodic fragment, thus achieving a comfortable listening experience. For example, based on the arrangement relationship between the tonic and dominant chords, a chord that matches the chord corresponding to the adjacent melodic fragment can be selected as the interpolation chord in the chord progression. If the chord corresponding to the adjacent melodic fragment is a C chord, a G chord can be selected to accompany it in the chord progression, resulting in a more pleasant listening experience.
[0051] After determining the interpolation chord, an interpolated melodic segment between two adjacent melodic segments can be generated based on the multiple audio events of each adjacent melodic segment and the interpolation chord.
[0052] When generating interpolated melody fragments between each pair of adjacent melody fragments, for each interpolated chord corresponding to the adjacent melody fragments, an interpolated melody fragment corresponding to that interpolated chord can be generated. Specifically, refer to... Figure 3 Generating the interpolated melody fragment corresponding to the interpolated chord involves the following steps:
[0053] 301. Based on the multiple audio events of the melody segment corresponding to the interpolated chord and the pitch of the interpolated chord, determine the multiple target audio events to be filled corresponding to the interpolated chord;
[0054] In this embodiment, the interpolated chord includes at least one interpolated chord corresponding to each melodic segment in the adjacent melodic segments; multiple interpolated chords are arranged according to their order in the chord progression; multiple pitches corresponding to multiple audio events are arranged sequentially to form a pitch array. It can be determined that the audio event with the same pitch as the interpolated chord among the multiple audio events of the melodic segment corresponding to the interpolated chord is the target audio event. For each audio event with a pitch different from the interpolated chord among the multiple audio events of the melodic segment corresponding to the interpolated chord, it is determined whether the pitch of the audio event has been replaced with the pitch of the interpolated chord. If so, the pitch of the audio event is replaced with the pitch of the interpolated chord, and the audio event with the replaced pitch is determined as the target audio event; otherwise, the audio event is determined as the target audio event.
[0055] For example, in Figure 4 In the scenario shown, both motive 1 and motive 2 correspond to C chords. Multiple chords can be selected during the chord progression as interpolation chords between motive 1 and motive 2, such as... Figure 4 The diagram shows the selection of the chords G, Am, Em, F, and G as interpolation chords. Specifically, G, Am, and Em chords are assigned as interpolation chords for motive 1, and F and G chords are assigned as interpolation chords for motive 2. Furthermore, if the number of interpolation chords n is even, then the first n / 2 chords are interpolation chords for motive 1, and the last n / 2 chords are interpolation chords for motive 2.
[0056] First, the audio event sequence of motif 1 is read, and its pitch information is filtered by the interpolation chords assigned to motif 1. That is, the audio events with the same pitch as the interpolation chord among the multiple audio events of motif 1 are identified as target audio events. For example, the pitch array corresponding to the multiple audio events of motif 1 is [1, 2, 3], while the pitches contained in the G chord are [5, 7, 2, 4]. The intersection of the two is the pitch 2. Therefore, the audio event with the pitch 2 among the multiple audio events of motif 1 is identified as the target audio event.
[0057] Secondly, for each audio event in motif 1 whose pitch differs from the pitch of the G chord, determine whether the pitch of that audio event has been replaced with the pitch of the G chord. If so, replace the pitch of that audio event with the pitch of the G chord and identify that audio event as the target audio event; otherwise, identify that audio event as the target audio event. For example, the audio event with pitch 1 in motif 1 may be replaced with any one of the pitches 5, 7, and 4 in the G chord, or it may be retained; the audio event with pitch 3 in motif 1 may be replaced with any one of the pitches 5, 7, and 4 in the G chord, or it may be retained.
[0058] Specifically, determining whether the pitch of an audio event in a melody segment should be replaced with the pitch of its corresponding interpolated chord can be achieved by: determining the replacement probability range corresponding to the audio event and generating a random probability; if the random probability is within the replacement probability range corresponding to the audio event, then the pitch of the audio event needs to be replaced with the pitch of the interpolated chord; if the random probability exceeds the replacement probability range corresponding to the audio event, then the pitch of the audio event is retained, i.e., there is no need to replace the pitch of the audio event.
[0059] Similarly, multiple target audio events can be obtained by filtering Motion 1 with Am and Em chords, and multiple target audio events can be obtained by filtering Motion 2 with F and G chords.
[0060] 302. Obtain the audio density corresponding to each time position between adjacent melody segments in the time axis where the melody segment sequence is located;
[0061] In this embodiment, the audio density can be a value within a preset range or a value set by the user, serving as a benchmark for determining whether the target audio event is filled at the corresponding time position on the timeline. When the user sets the audio density, the computer device can provide an interface for the user to set the audio density, receive the user's setting operation for the audio density, and obtain the audio density corresponding to each target audio event based on the setting operation.
[0062] For example Figure 5 As shown, the user can adjust the audio density at three levels within the interface. The computer device determines the corresponding audio density based on the user's input. Each level corresponds to a different audio density range: a high level corresponds to a higher audio density value, a medium level to a medium audio density value, and a low level to a lower audio density value. Alternatively, the user-input audio density can be obtained in other ways, such as by entering a numerical value in the input box or by selecting the desired audio density from multiple options.
[0063] 303. For each target audio event, determine the time position to fill the target audio event based on the audio density of each time position in the time axis, and fill the target audio event at the time position to fill the target audio event;
[0064] In this embodiment, audio density represents the distribution density of multiple target audio events along the time axis. Therefore, the higher the audio density, the denser the distribution of multiple target audio events; the lower the audio density, the more dispersed the distribution of multiple target audio events. For each target audio event, the time position used to fill the target audio event can be determined based on the audio density at each time position on the time axis, and the target audio event is filled at the time position used to fill the target audio event.
[0065] When all target audio events are filled at their respective time positions, all target audio events that have been filled at their respective time positions can constitute the interpolated melody fragment corresponding to the interpolated chord.
[0066] In a preferred embodiment, the timeline can be moved sequentially with the end time of the previous melody segment in an adjacent melody segment as the starting point, and each step is a preset duration. When moving forward, it is determined whether the random number generated based on the current time position is greater than the audio density corresponding to the current time position. If so, the target audio event to be filled is filled at the current time position, until all target audio events of the interpolated chord are filled at the corresponding time positions in the timeline, thus obtaining the interpolated melody segment corresponding to the interpolated chord.
[0067] For example, after determining the multiple target audio events corresponding to the current interpolation chord and the audio density corresponding to each target audio event, it can be determined whether to fill the target audio event at that time position according to the random number corresponding to the time position of each step and the audio density corresponding to that time position. Here, the random number is a value calculated based on the current time position according to a preset algorithm. For example, the random number can be determined based on the distance between the current time position and the two preceding and following melodic segments in the adjacent melodic segment. For instance, the functional relationship between the random number and this distance can be a linear function, such as a linear function like the random number y = ax + bx' + c, where a and b are the slopes, c is a constant, x is the distance between the current time position and the preceding melodic segment of the adjacent melodic segment, and x' is the distance between the current time position and the following melodic segment of the adjacent melodic segment.
[0068] Of course, the functional relationship between the random number and the distance is not limited to the above functional relationship. As long as the functional relationship can represent how the random number changes with the distance, it is acceptable. No limitation is made here.
[0069] Specifically, if the audio density is represented as a probability value, then the random number is specifically a random probability value. Of course, the audio density and the random number can also be other types of values, which are not limited here.
[0070] For example, in Figure 4In the example shown, after motif 1 is filtered by the G chord, assuming the identified target audio events are the three audio events corresponding to pitches 5, 2, and 7, these three audio events are the target audio events to be filled. Starting from the end time of motif 1, the steps are moved sequentially in 16th note increments, using a pointer to indicate the time position reached at each step, such as... Figure 6 As shown, when a certain step reaches a corresponding time position, if the random number corresponding to the current time position is greater than the audio density corresponding to the current time position, then an audio event with a pitch of 5 is filled. Continuing on, when a certain step reaches a corresponding time position, if the random number corresponding to the current time position is greater than the audio density corresponding to the current time position, then an audio event with a pitch of 2 is filled. Continuing on, similarly, an audio event with a pitch of 7 can be filled at a certain time position. At this point, all target audio events have been filled to their corresponding time positions, and all target audio events constitute the interpolated melodic fragment corresponding to the G chord between motif 1 and motif 2.
[0071] Similarly, interpolated melodic fragments corresponding to Am, Em, F, and G chords can be obtained, such as... Figure 7 The process shown yields interpolated melodic fragments for multiple chords, including those corresponding to the G chord, Am chord, and Em chord. The all-zero array is used to pre-allocate a portion of the storage space. Its purpose is to prevent the storage space from being completely occupied by other data, leaving no available storage space for the newly generated interpolated melodic fragments. Therefore, after generating the interpolated melodic fragments, simply replacing the all-zero array with the interpolated melodic fragments completes the storage of the interpolated melodic fragments.
[0072] If the audio density is a value within a preset range, then in this timeline, the closer the current time position is to one of the motifs, the smaller the corresponding audio density value; the farther away it is from the two motifs before and after, the larger the corresponding audio density value. For example, the audio density value range can be set to multiple probability value ranges such as [0%~50%].
[0073] Therefore, by adding interpolated melodic fragments to the sequence of melodic fragments based on audio density, it is possible to achieve the repetition of notes and motifs in the development of compositional motifs, so that the generated melody can maintain structural consistency, has a structure that is discernible to the human ear, and is more memorable and pleasant to listen to.
[0074] Based on the above embodiments and their preferred implementations, in another optional implementation, after all target audio events have been filled at their corresponding time positions on the time axis, the offset amount of each target audio event shifted forward or backward on the time axis, and the offset probability of each target audio event, can also be obtained. For each target audio event that has been filled at its corresponding time position on the time axis, it is determined whether to move the target audio event at its time position on the time axis based on the offset probability of the target audio event. If yes, the target audio event is moved at its time position on the time axis based on the offset amount of the target audio event; otherwise, the time position of the target audio event on the time axis is maintained.
[0075] For example, the offset could be one or more sixteenth notes or other note values. Suppose the offset represents moving forward by two sixteenth notes, if it is determined that a target audio event needs to be moved, then based on this offset, the target audio event can be moved forward by two sixteenth notes, thereby determining the new time position of the target audio event on the timeline.
[0076] When determining whether to shift the target audio event based on the offset probability, a random probability value can be calculated based on the time position of the target audio event on the time axis according to the preset algorithm. Then, it is determined whether the random probability value is greater than the offset probability corresponding to the target audio event. If so, it is determined that the target audio event needs to be moved; otherwise, the time position of the target audio event on the time axis is maintained.
[0077] The random probability value can be determined based on the distance between the current time position of the target audio event and the time positions of adjacent audio events. For example, the functional relationship between the random probability value and the distance can be a linear function.
[0078] In addition, both the offset and the offset probability can be customized by users, allowing them to freely adjust the rhythm and style of the melody according to their own listening preferences.
[0079] Therefore, by providing the function of offsetting the time position of the target audio event, the relative distance between audio events in time position can be changed, thereby realizing the change of rhythm in the development of compositional motifs. Furthermore, users can independently set the offset amount and offset probability, such as adjusting the offset amount or offset probability to make the rhythm of the melody more dense or the rhythm of the melody more relaxed, to meet the user's requirements for the melody in different listening scenarios or to meet the user's different listening preferences.
[0080] Based on the above embodiments and their preferred implementations, in another optional implementation, when arranging motivic melody fragments according to the musical structure to reproduce the musical motif, the similarity between multiple melody fragments can be determined according to the characteristics of the pitch represented by the audio event. Based on the similarity, a motivic melody fragment is determined among the multiple melody fragments, and a copy of the motivic melody fragment is generated. The motivic melody fragment and its copy are arranged according to the determined musical structure to obtain a melody fragment sequence.
[0081] The computer device can determine the similarity between two melodic segments based on the degree of approximation of their pitch characteristics. For example, it can determine the similarity between two melodic segments based on the similarity of their pitch name features, the similarity of their pitch start time features, or the similarity of their pitch duration. The method of determining the similarity is not limited; for example, it can be calculated using the Pearson correlation coefficient formula, or based on Euclidean distance, Manhattan distance, or other similar methods.
[0082] Each melody segment consists of multiple audio events. At this point, the melody segments generated by the network model are filtered based on the previously generated chord information and the key information derived from it, leaving only usable segments. For example, if the generated melody segment corresponds to a C chord, then notes belonging only to the C chord will be filtered out. Afterwards, similarity calculations are performed on the filtered melody segments.
[0083] Specifically, the similarity between any two melody segments can be calculated based on arbitrary features of the pitch of the melody segments. In a preferred embodiment, features such as the pitch name (pitch) and start time (start_sample) of the audio event can be extracted to calculate the similarity. For each pair of melody segments, a first similarity is determined between the pitch name features of one melody segment and the pitch name features of the other melody segment, and a second similarity is determined between the pitch start time features of one melody segment and the pitch start time features of the other melody segment. The first similarity and the second similarity are then weighted and summed, and the resulting weighted sum is used as the similarity between the two melody segments.
[0084] In this embodiment, audio events represent the pitch characteristics of a melody segment. Each audio event corresponds to a pitch, including pitch characteristics such as the pitch name, duration, start time, timbre, and end time. Figure 8 As shown, the audio event (note_event) corresponding to a note in a musical score can be described as having various information such as pitch name, timbre, end marker donesample, duration remainsample, and start marker startsample.
[0085] Therefore, based on the various pieces of information mentioned above in the audio event, such as... Figure 9 The two melodic fragments shown can be represented by numbers, where the pitch names of fragment 1 and fragment 2 can be replaced with numbers. Specifically, the pitch names C through B can be replaced according to the following rules: C=1, D=2, E=3, F=4, G=5, A=6, B=7. Furthermore, the start time is standardized to the length of a sixteenth note (calculated from bpm, i.e., 16note length = (60 / bpm) / 4), meaning the start time is converted into n sixteenth notes (n can be a floating-point number). Therefore, as described above, the pitch name characteristics of fragment 1 are [1,2,3], and the start time characteristics are [2,6,14]; the pitch name characteristics of fragment 2 are [1,2,3,3], and the start time characteristics are [2,8,12,14]. The similarity between each pair of melodic fragments can be calculated based on these quantified characteristics.
[0086] Of course, the quantization method for the pitch features of a melody fragment is not limited to the examples mentioned above. Other quantization methods can also be used to quantize the pitch features into other values. Furthermore, while pitch feature quantization facilitates the calculation of similarity between melody fragments, it is also possible to calculate the similarity between melody fragments directly based on the pitch features without quantizing the pitch features. This is not a limitation here.
[0087] After obtaining the pitch features of each pair of melodic segments, the similarity between the segments can be calculated based on these pitch features. For example, in the example above, the pitch name feature of segment 2 is [1,2,3,3]. Since note repetition is a way of developing musical motifs, the pitch array can be deduplicated in its pitch name feature, but the order of the pitches cannot be changed. Therefore, the final pitch name feature of segment 2 is [1,2,3]. For pitch start time features of inconsistent lengths, padding can be performed according to the longest feature size. For example, the padding operation can be to randomly select a sample value in the segment and repeat it, but the order of the pitches corresponding to the start time feature cannot be changed. Therefore, the pitch start time feature of segment 1 can be expanded to [2,2,6,14].
[0088] Using the Pearson correlation coefficient formula, the correlation between the pitch feature of segments 1 and 2 is calculated to be 1, and the correlation between the start time feature is 0.87. Different weights can be assigned to the two features, for example, 1:1, then the final correlation coefficient between segments 1 and 2 is 1*0.5 + 0.87*0.5 = 0.935.
[0089] By analogy, the similarity of different groups of melodic fragments can be calculated, and one or more groups of melodic fragments can be selected based on the similarity as several motif fragments to be repeated in the subsequent music generation.
[0090] Specifically, when determining motivic melody segments from multiple melodic fragments based on similarity, the melodic fragments with the highest similarity can be identified as motivic melody segments. For example, if each pair of melodic fragments is grouped together, the similarity between the two melodic fragments in each group can be calculated. The computer device can identify the group or groups of melodic fragments with the highest similarity among multiple groups of melodic fragments and generate copies of the group or groups of melodic fragments with the highest similarity. The multiple melodic fragments with the highest similarity and their copies are then arranged according to the aforementioned determined musical structure to obtain a sequence of melodic fragments.
[0091] For example, suppose that the melodic fragments with the highest similarity among multiple sets of melodic fragments are fragment 1 and fragment 2, and the musical structure is determined to be ABAB. Then, generate copies of fragment 1 and fragment 2, and arrange fragment 1, fragment 2 and their copies according to the above musical structure. The resulting melodic fragment sequence is fragment 1-fragment 2-fragment 1-fragment 2-...fragment 1-fragment 2.
[0092] Furthermore, assuming that fragment 1 and fragment 2 are the two melodic fragments with the highest similarity, fragment 1 can be designated as motif 1 and fragment 2 as motif 2, as referenced. Figure 4 As shown, by repeating and reproducing motif 1 and motif 2 according to the predetermined musical structure ABAB, the corresponding melodic fragment sequence can be obtained. Of course, fragment 2 can also be set as motif 1, and fragment 1 as motif 2, and repeated and reproduced based on this musical structure; this is not limited here.
[0093] Alternatively, multiple melodic fragments can be grouped together. The similarity between any two melodic fragments in each group can be calculated. The top n groups of melodic fragments with the highest similarity (n≥2) can be identified as motivic melodic fragments. The motivic melodic fragments and their copies can then be arranged to obtain a sequence of melodic fragments.
[0094] Alternatively, multiple melodic segments whose similarity meets a preset similarity threshold can be identified as motivic melodic segments. This preset similarity threshold can be an empirical value obtained by analyzing whether multiple melodic segments together constitute a musical motif based on their musical characteristics and summarizing the similarity between multiple sets of melodic segments that may constitute a musical motif, or it can be an empirical value obtained based on a large number of experiments. No limitation is made here.
[0095] In this embodiment, the arrangement of music can be flexibly changed according to the user's listening scenario and needs, satisfying the user's personalized customization of music. Therefore, the user only needs to set the music scenario, without changing the music arrangement template, and without needing to understand the professional terminology and music theory of music arrangement. This simplifies the operation process of generating music audio, reduces the time spent on music generation, improves the efficiency of music generation, and thus enhances the user's listening experience.
[0096] The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 10 One embodiment of the computer device in this application includes:
[0097] The computer device 1000 may include one or more central processing units (CPUs) 1001 and a memory 1005, wherein the memory 1005 stores one or more applications or data.
[0098] The memory 1005 can be volatile or persistent storage. The program stored in the memory 1005 can include one or more modules, each module including a series of instruction operations on the computer device. Furthermore, the central processing unit 1001 can be configured to communicate with the memory 1005 and execute the series of instruction operations in the memory 1005 on the computer device 1000.
[0099] Computer device 1000 may also include one or more power supplies 1002, one or more wired or wireless network interfaces 1003, one or more input / output interfaces 1004, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0100] The central processing unit 1001 can perform the aforementioned... Figures 2 to 3 The specific operations performed by the computer device in the illustrated embodiment will not be described in detail here.
[0101] This application also provides a computer storage medium, one embodiment of which includes: the computer storage medium storing instructions, which, when executed on a computer, cause the computer to perform the aforementioned... Figures 2 to 3 The operations performed by the computer device in the illustrated embodiment.
[0102] This application also provides a computer program product, one embodiment of which includes: when the computer program product is run on a computer device, causing the computer device to perform the aforementioned... Figures 2 to 3 The operations performed by the computer device in the illustrated embodiment.
[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for generating musical audio, characterized in that, The method includes: Generate chord progressions and musical structure based on the currently defined musical scene; Multiple melody fragments are generated. Based on the musical characteristics represented by each melody fragment, a motivic melody fragment is determined from each melody fragment. The motivic melody fragment serves as a musical motif corresponding to the musical scene. The motivic melody fragments are arranged according to the musical structure to reproduce the musical motivator, resulting in a sequence of melody fragments. Based on the chord progression, at least one interpolated melody segment is added to the sequence of melody segments to develop the musical motif, resulting in an audio file that matches the musical scene.
2. The method according to claim 1, characterized in that, The addition of at least one interpolated melody segment to the melody segment sequence based on the chord progression to develop the musical motif and obtain a musical audio file matching the musical scene includes: Based on the chord progression, at least one interpolated chord is added to the sequence of melodic fragments, and the interpolated melodic fragment corresponding to the interpolated chord is generated; The sequence of melodic fragments and its multiple interpolated melodic fragments are output as an audio file of music that matches the musical scene.
3. The method according to claim 2, characterized in that, Each of the melodic segments includes multiple audio events, which are used to represent the pitch characteristics of the melodic segment; The step of adding at least one interpolated chord to the melodic fragment sequence based on the chord progression and generating the interpolated melodic fragment corresponding to the interpolated chord includes: For each pair of adjacent melodic segments in the melodic segment sequence, the interpolated chord corresponding to the adjacent melodic segments is determined during the chord progression; The interpolated melody segment between the two melody segments of the adjacent melody segments is generated based on the plurality of audio events of each of the adjacent melody segments and the interpolated chords.
4. The method according to claim 3, characterized in that, The step of generating the interpolated melody segment between two adjacent melody segments based on the plurality of audio events of each of the adjacent melody segments and the interpolated chords includes: For each interpolated chord corresponding to the adjacent melody segments, based on the multiple audio events of the melody segment corresponding to the interpolated chord and the pitch of the interpolated chord, determine the multiple target audio events to be filled corresponding to the interpolated chord; Obtain the audio density corresponding to each time position between adjacent melody segments in the time axis where the melody segment sequence is located; For each target audio event, the time position to fill the target audio event is determined according to the audio density of each time position in the time axis, and the target audio event is filled at the time position to fill the target audio event; Among them, all the target audio events that complete the filling at the corresponding time position constitute the interpolated melody segment corresponding to the interpolated chord.
5. The method according to claim 4, characterized in that, The step of determining the time position for filling the target audio event based on the audio density at each time position in the time axis, and filling the target audio event at the time position for filling the target audio event, includes: In the timeline, starting from the end time of the previous melody segment in the adjacent melody segments, the process proceeds sequentially with a preset duration as the step size. At each step, it is determined whether the random number generated based on the current time position is greater than the audio density corresponding to the current time position. If so, the target audio event to be filled is filled at the current time position until all the target audio events are filled at the corresponding time positions in the timeline, thereby obtaining the interpolated melody segment corresponding to the interpolated chord.
6. The method according to claim 4, characterized in that, The interpolated chords include at least one interpolated chord corresponding to each of the adjacent melodic segments; the plurality of interpolated chords are arranged according to their order in the chord progression; The step of determining multiple target audio events to be filled corresponding to the interpolated chord based on the multiple audio events of the melody segment corresponding to the interpolated chord and the pitch of the interpolated chord includes: The target audio event is identified as the audio event in which the pitch of the interpolated chord is the same as the pitch of the interpolated chord among the multiple audio events of the melody segment corresponding to the interpolated chord. For each audio event in the multiple audio events of the melody segment corresponding to the interpolated chord whose pitch is different from the pitch of the interpolated chord, determine whether the pitch of the audio event has been replaced with the pitch of the interpolated chord. If yes, replace the pitch of the audio event with the pitch of the interpolated chord and determine the audio event whose pitch has been replaced with the pitch of the interpolated chord as the target audio event; if no, determine the audio event as the target audio event.
7. The method according to claim 4, characterized in that, The step of obtaining the audio density corresponding to each time position between adjacent melody segments in the timeline of the melody segment sequence includes: Get the currently set audio density range; For each time position in the time axis, the audio density corresponding to the time position is determined within the audio density range based on the distance between the time position and the two preceding and following melodic segments in the adjacent melodic segment.
8. The method according to claim 4, characterized in that, After all the target audio events are filled in at their corresponding time positions on the timeline, the method further includes: Obtain the offset amount of each target audio event shifted forward or backward on the time axis, and obtain the offset probability of each target audio event; For each target audio event that has been filled at the corresponding time position in the time axis, determine whether to move the target audio event at the time position in the time axis based on the offset probability of the target audio event. If yes, move the target audio event at the time position in the time axis based on the offset of the target audio event; otherwise, maintain the time position of the target audio event on the time axis.
9. The method according to claim 8, characterized in that, The step of determining whether to move the target audio event's time position on the time axis based on the target audio event's offset probability includes: The random probability value corresponding to the target audio event is determined based on the distance between the current time position of the target audio event and the time positions of its adjacent audio events; If the random probability value is greater than the offset probability of the target audio event, then it is determined that the target audio event will be moved to the time position on the time axis based on the offset of the target audio event; If the random probability value is less than the offset probability of the target audio event, then the time position of the target audio event on the time axis is determined to be maintained.
10. The method according to claim 1, characterized in that, Each of the melodic segments includes multiple audio events, which are used to represent the pitch characteristics of the melodic segment; The step of arranging the motivic melody fragments according to the musical structure to reproduce the musical motif and obtain a sequence of melody fragments includes: Based on the pitch characteristics represented by the audio events, the similarity between each pair of the multiple melody segments is determined; Based on the similarity, the motivic melody segment is determined from the multiple melody segments, and a copy of the motivic melody segment is generated. The motivic melody segment and its copy are arranged according to the musical structure to obtain a melody segment sequence.
11. The method according to claim 10, characterized in that, The step of determining the pairwise similarity between the plurality of melody segments based on the pitch characteristics represented by the audio events includes: For every two melodic segments, a first similarity is determined between the pitch name feature of one melodic segment and the pitch name feature of the other melodic segment, and a second similarity is determined between the pitch start time feature of one melodic segment and the pitch start time feature of the other melodic segment. The first similarity and the second similarity are then weighted and summed, and the resulting weighted sum is taken as the similarity between the two melodic segments.
12. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 11.
13. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed on the computer, cause the computer to perform the method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Music accompaniment generation method and device, equipment, storage medium and program product
CN117765902A
Chord sequence generation method and device, terminal and storage medium
CN118072698A